Why public tenders are hard to find

Searching public tenders is hard.

We’ve all been there; you search for the exact kind of contract you would bid for and nothing new comes back. A week later you get an email ‘why aren’t we bidding on this?’ if you’re lucky you’re still in the frame, if you’re unlucky the deadline is too close to respond in time.

We’re working every day to close the gap on this problem, as part of that work we decided to try to understand how big the gap was and what we found surprised us.

We took 2,405 open UK Find-a-Tender notices and 20 real customer searches, we then had every notice graded for relevance, and measured where the difficulty actually comes from. Buyers often describe a contract in words a searcher would not think to use, and standard keyword search does nothing to help the searcher bridge that gap. The pattern underneath is straightforward to state: the words a buyer writes and the words a searcher types are often different words for the same contract. The encouraging part is that bidders can work to close the gap, but it takes effort. We explain at the end what that takes and what Open Opportunities is doing to make it easier.

Throughout, the thing we care about is recall. Recall means the share of the tenders that really are relevant to you that the search actually puts in front of you. Ranking the results and filtering out the noise both come afterwards. If a relevant tender never appears at all, none of that can help.

A worked example

Here is the complete description of a real, open Find-a-Tender notice from Flagship Group.

“Flagship Group is setting up a DPS for vehicles and associated supplies and services. Mini competitions will be issued throughout the life of the DPS. All the bidders registered on the DPS who have passed the Selection Stage will be invited to participate in the mini-competitions. Access to the Selection Stage will be open throughout the life of the DPS. The DPS may be used for any of Flagship Group’s subsidiary companies.”

That copy is 72 words long, but only one word in that notice that has any value to the bidder: “vehicles”. The words around it, such as “associated”, “supplies” and “services”, appear in many thousands of notices, so they tell a search engine nothing about what this particular contract is for. Now think about how a real person searches. If you searched for “car hire”, this notice would never come back, because it does not contain the word “car” or the word “hire”. If you sensibly asked the search to return only notices that contain both “vehicle” and “hire”, it would still not come back, because the word “hire” is not there either. A contract you would happily bid for stays invisible, and the reason is simply that its one useful word was not among the words you chose. This is what a low signal-to-noise ratio means in practice: the real subject of a notice is often only one or two words, surrounded by filler.

How often the words do not line up

To measure this across the whole set, we did something deliberately simple and repeatable. For every notice that a judge had marked as genuinely relevant to a particular search, we checked whether any of the user’s own search terms appeared anywhere in that notice’s title or description. We allowed for singular and plural, so that “vehicle” also matches “vehicles”, but nothing more than that. We looked at the title and description only. Later, when we measure how many relevant tenders a search actually returns and in what order, we use BM25, the standard keyword-ranking method used across the industry. BM25 scores a notice by how many of the search words it contains and how distinctive those words are; on its own it does not match a word to a synonym.

How often the words you searched for are missing from the notice.

For 36% of relevant tenders, none of the user’s search terms appeared in the notice at all, even allowing for singular and plural. That figure splits in two by how relevant the tender is. For the tenders that were squarely about what the person had searched for, the words were missing 19% of the time. For related, nearby contracts that a bidder might also want to see, they were missing 45% of the time. So there is a real gap for users even when they’re using highly relevant search terms, when the subject of a tender is relevant but not a core topic that gap expands more. At Open Opportunities we call this “the ships in the night problem” effectively we have a language gap between the search terms and the buyer’s description. Take a notice for “Workforce Skills Framework” once you look at the underlying documents this opportunity includes apprenticeships, but the word “apprenticeships” never appears, the useful parts of the description are all pitched at the higher level and in broad terms, the specifics are absent.

Widening the search with more words

Some of the 20 searches we tested were just two words, and others were long lists of related terms that the customer had typed, such as “custom software”, “staff augmentation” or “legacy modernisation”. We ran each of the long-list searches two ways over the same topic and the same set of relevant tenders: once with only its first phrase, and once with the whole list.

Each search is described by two numbers. The match is how often the search words appear in a relevant notice at all. The recall is how often the tender is actually returned in the results which were capped at 200. A tender can only be returned if one of its words matched, so recall is never higher than the match, and where it is lower, the words were in the notice but the notice did not rank highly enough to appear in the list of 200 notices.

Adding related words improves match, recall and precision together.

With a single phrase, the search words appeared in only 27% of the relevant notices, and the search returned 26% of them. Almost everything that matched was returned, because so little matched that it all fit comfortably in the results. When we looked at the whole family of related words, the search words appeared in 83% of the relevant notices, but the search only returned 52% of them. So even when the words were present in the tender notice they were ranked at the bottom and outside the 200 results the search returned. For the tenders squarely on target the effect was stronger still: their words matched 92% of the time, but 65% were returned. This is plain keyword search, so the gain is not down to any clever technology. A person who uses more of the words a buyer might have used finds far more of the tenders they want.

The wider search also stayed clean. A search that returns more results might be expected to return more junk, but the opposite happened. Of the top twenty results, the share that were actually relevant rose from 37% with a single phrase to 48% with the full list, because the tenders that matched several of the specific terms were the ones that rose to the top.

The returns diminish

More words is not a lever you can pull forever. Adding the related terms one at a time, across the nine list searches, the early ones did most of the work and each later one added less.

In our searches, more words gave diminishing returns.

In this sample, recall climbed quickly over the first several words and then levelled off in the low fifties, even as the words kept appearing in more and more notices, up to 83% of them. The gap between those two lines is the point. Past a certain richness, extra words keep turning up in relevant notices but no longer lift those notices into the top results. We would not read an exact turning point off nine searches, and we are not claiming one. What the sample shows is diminishing returns: the first handful of well-chosen words buys most of the recall that words can buy, and beyond that, adding more is not where the next gain comes from.

But widening the wrong way does add noise

That last result comes with an important condition, and it is the reason widening a search is not as simple as it sounds. The gain came from adding specific related terms, not from making any single term vaguer. Broadening a search by reaching for one loose word does the reverse. Go back to the vehicle example. If, in order to be safe, you searched for “vehicle” on its own rather than “vehicle hire” or “car hire”, you would indeed catch the Flagship notice, but you would also catch every other use of the word: e.g. an “investment vehicle”, a “contracting vehicle” and every bus, van, boat and train in the corpus. The skill, and the risk, lies in choosing terms that are both related and specific, rather than simply loosening the search until everything comes back.

Half the opportunities are still missing

About half is as far as words alone take us, and it is not good enough. Even with a carefully chosen family of related terms, the full search returned 52% of the relevant tenders in the top 200. For a supplier that is one contract in two that they could have bid for, still out of sight after they have worded the search as well as anyone reasonably could.

Returning and reading more of the results

If more words is not the next lever, looking further into the results is. The 52% counts only the tenders that reached the top 200, and many relevant tenders did match a word but sat lower down. The obvious move is to return more of them and look deeper.

Reading deeper finds more, but buries you in noise.

It works, to a point. Reading down to the 500th result lifts recall to 72%, to the 1000th to 81%, and to the end of the list to 83%. The cost is that the results get steadily dirtier. By the bottom, only about seven in a hundred are relevant, and no person can sift that. So reading further only helps when something reorders the list first, pulling the relevant tenders back up where they can be seen. The practical form of “look further” is to have the search retrieve a much larger set and then rerank it, not to ask anyone to read to the bottom.

Why the search tops out at 83%

The search’s own list runs out at 83% recall, and the missing sixth are not buried somewhere below the bottom of it. They are not on the list at all. A keyword search ranks only the notices that share at least one of the search words; the rest score nothing and never appear, however far down you read. They sit in the corpus untouched, invisible to this particular search though not to another.

The tenders that share no word tend to be the thinnest and most awkwardly worded. Comparing the ones a full search finds against the ones it misses, the missed notices are shorter, with a median length of about 600 characters compared with about 930, and many are related contracts described in terms no reasonable searcher would have listed. The kind of notice matters too.

Standing arrangements are harder to find than one-off contracts.

The search words are missing 42% of the time for a framework or a Dynamic Purchasing System (DPS), compared with 29% of the time for a one-off contract. This is not because these notices are shorter. Framework and DPS descriptions are actually a little longer on average than one-off notices, and they are far less likely to be almost empty. The words go missing because these notices are often more concerned with explaining how to join the contract, when they do describe requirements they use very few words and list only the broad categories it covers, on the whole they rarely mention the specifics. As more public buying moves onto these arrangements, this part of the gap grows.

Of course, a search built from the words the buyer actually used would return these notices, the trouble is that a supplier cannot count on guessing those words every time. They are ordinary descriptions, just not the ones the searcher considered relevant to their needs. The broadness of the terms also means that these are the very terms that bring the noise back.

Dilemmas and asymmetry

So here’s the dilemma: refine your search to bring relevance at the expense of recall or broaden your search to secure recall but invite noise into your feed. Which also reveals the asymmetry of interests, those publishing notices are not really incentivised to write precise notices, instead they’re incentivised to comply with the legal framework that mandates publishing. So the publisher has no motivation to make their descriptions receptive to a BM-25 algorithm. However, suppliers looking for notices have a profound incentive to find every relevant notice they can, both to avoid losses and to ensure completeness.

What Open Opportunities are doing about it

Our customers are the suppliers, those that are highly motivated to find all of the relevant notices across hundreds of different sources. What this research shows (and what we’d always known) is that aggregating data alone is not enough. If you’re gathering 100% of the relevant documents and revealing just 50% of them you’re not solving your customers problems.

Our own search already outperforms these tests, Open Opportunities users can search using boolean operators, can exclude certain types of opportunity and add additional filters. We enrich every notice with language, country and value tags, and are multi-currency, to help narrow the search and reduce the noise. Finally we give every notice five CPV codes so that our users can only look at relevant notices.

We’re also building a search algorithm that uses vectors, a list of numbers that represent a meaning, rather than using a specific keyword we use meanings. So when the publisher uses just one viable word in their notice, such as “vehicle” our vector knows that the word “car” actually isn’t that far away and matches it.

We’re rolling out a test version to some existing customers this month and will be showing everyone our test results soon.

How we conducted our analysis

2,405 open UK Find-a-Tender notices, each with a deadline still in the future. 20 real customer searches. A language model graded every notice for relevance on a scale of 0 to 3, and we checked its grades against 222 human judgements, which it matched closely, with a quadratic-weighted kappa of 0.75. A grade of 2 or 3 counts as relevant. The searchable text is the title and description only, with the CPV code left out. We test with BM25, the standard keyword-search method used across the industry, and we wrote it out in full so that anyone can reproduce these figures exactly. Three measures appear above: a match means one of the search words appears in the notice; recall is the share of known-relevant tenders that the search returns in its first 200 results; and precision is the share of the top results that are actually relevant. The comparison figures are averaged across the searches. One limitation works against our own headline: this test matches words, not meanings, so it cannot see genuine synonyms and therefore slightly overstates how often a word is truly missing. The real gap is a little smaller than the figures suggest, and generating the related words well recovers more than a plain word-match test can show.

The method and the data

How these numbers were produced — the corpus, the validated judge, BM25, and the three measures defined once — is written up in full: Why public tenders are hard to find: the measurement. Every figure is available as a downloadable data file: tender-search-results.json.