How we measured it

The whole test, in plain words, and every file behind it to download.

What we did

  1. We wrote down 100 things suppliers search for, as a two-word term each, the way you would type it: "commercial cleaning", "civil engineering", "fire detection".
  2. On each of nine tender portals we typed each term into the portal's own search box and into ours, on the same day, and kept the first 100 results from both.
  3. We listed every tender the portal showed as open that day and matched it to our database, so both searches were judged on the same tenders. Tenders published in the two days before the test and tenders closing on the day were left out of both sides.
  4. A judge read every open tender on the portal, for every one of the 100 terms, and marked the ones that were for that kind of work. The judge is a fixed language model that never saw which search had found a tender.
  5. Native speakers checked the judge: on every portal a person re-graded 300 of its decisions without seeing them.
  6. Then we counted. For each search, which relevant tenders did each side show, and where in the list did they sit.

The portals

PortalTest dayOpen on the portalCounted in the testRelevant in either first 100Portal found / missedWe found / missedJudge: right / found, in 100
Find a Tender2026-09-162,3652,339709251 / 458672 / 3778 / 67
World Bank2026-09-15576532372136 / 236359 / 1389 / 86
TenderNed2026-09-171,172836340141 / 199325 / 1587 / 78
Doffin2026-09-161,087659417202 / 215388 / 2990 / 83
SAM.gov2026-09-1651,6682,51243293 / 339406 / 2682 / 69
SECOP II2026-09-162,5491,98058896 / 492563 / 2588 / 93
BOAMP2026-09-1615,6746,7581,404165 / 1,2391,372 / 3294 / 82
bund.de2026-09-1612,6028,1291,660207 / 1,4531,572 / 8885 / 83
TED2026-09-1746,96541,1314,354904 / 3,4503,753 / 60186 / 86

"Relevant in either first 100" is the count both sides are measured against: the relevant open tenders that at least one of the two searches showed in its first 100 results, over the 100 searches. "Judge right" is the share of what the judge called relevant that a native speaker agreed with; "judge found" is the share of what the person would call relevant that the judge found.

Words we use

A search One two-word term, typed into both search boxes on the same day.

Relevant A tender for that kind of work, in the judge's reading of the tender itself: its own subject is the work, or it is a framework or list set up for that work. Closely related work does not count.

The first page The first 20 results. Five pages is 100.

Found and missed Of the relevant tenders either search showed in its first 100, the ones this search showed, and the ones it did not.

Counted in the test The share of the portal's open notices that were in the comparison set on the test day: the notice types both searches handle. The rest are in our database as other document types and were not scored on either side.

Download the data

Everything on the results page can be recomputed from these files. Open any of them in a spreadsheet.

  • summary_by_portal.csv: One row per portal: what was open, what was counted in the test, found and missed on each side at 20 and 100 results, the judge's marks.
  • results_by_search.csv: 900 rows, one per search per portal: the term, how many relevant tenders existed, what each side showed in its first 20 and 100.
  • ranking_bands.csv: Where each side's relevant results sit: ranks 1 to 10, 11 to 20, 21 to 50, 51 to 100.
  • README.md: What every file and column means.

The full result lists for every search on every portal (both sides, with rank, notice id and grade), and the 300 notices per portal graded by a native speaker, are held in the data pack and sent on request.

Grades run 0 to 3: 3 the tender's own subject is this kind of work (frameworks and dynamic purchasing systems count when set up for it); 2 closely related work, or a broad notice that could include it; 1 loosely connected; 0 a different field. Relevant means 3. Every grade-3 was asked a second, narrower question, whether the tender is for the work itself rather than something broader or neighbouring, and kept only if it passed; the wording of that second question was chosen by testing three versions against the native speakers' grades on two portals before it was used.

The judge is Gemini 3.6 Flash at temperature zero with a fixed prompt, given the tender's title, buyer, category codes and description, and the two-word term. About 1.4 million tender-by-term readings were made across the nine portals. On the five smaller portals every open tender was read for every term. On BOAMP, bund.de, SAM.gov and TED the two result lists were read in full plus 800 tenders per term drawn at random from the rest, and the relevant tenders outside the lists were estimated from that draw; those estimates are not used in the found and missed figures, which count only tenders one of the searches showed.

Two rules follow from how our search index is built. It is rebuilt twice a day, so a new notice is searchable within 12 hours at most. To keep both searches on the same notices, the test left out tenders published in the two days before the runs. It also drops a tender on its closing day, while the portal still lists it; tenders closing on the test day are left out of both sides. Both counts are in the files.

Portal searches were driven through each portal's own website, the way a supplier uses it, at the site's own paging and rate limits, with one exception: SECOP II's search sits behind a reCAPTCHA, so its official open-data search stood in. Find a Tender was asked in its all-words form, the words joined with a plus sign, which is that site's syntax for requiring every word. Our search was asked the same words as plain text, restricted to the same set of tenders.