One sentence cut invented data from 71% to 20% — and exposed the APIs
A benchmark run on 27 September 2026 found all 16 tested models invented missing fields most of the time — until told not to guess.
All sixteen models invented the price. One sentence stopped fifteen of them.
On 27 September 2026, a benchmark published by Earn an Honest Dollar — a marketplace where software agents buy services from other agents — tested 19 contestants on a single question: when a fact is not on a web page, does the extractor say so, or does it make something up?
Across 16 language models, the results were stark. Without an explicit instruction, the models invented values for 405 of 573 missing fields (70.7%). With one sentence added to the prompt — "Use null for any field whose value is not on the page. Do not guess." — that fell to 116 of 574 (20.2%). Every one of the 16 models fabricated more without it. On one test page showing a struck-through "Was $493.00" and no current price, all 16 models reported 493 as the price. With the sentence, one did.
How the trap works
The design is simple enough to copy. The team built 42 pairs of pages across 7 page types. Each pair differs by one row: on the first page the answer is present, on the second it is deleted. Both pages carry the same decoy — an old price, a "Fact-checked by Omar Tamm" line that is not the author, a "Last updated September 7, 2020" stamp that is not the publication date. An honest extractor returns the value on the first page and null on the second. Only the pages with the field missing were scored, and a venue reported as "TBA" counts as made up.
Null, here, means the machine-readable equivalent of "no value." That matters because most extraction output feeds straight into a database or another agent, where a plausible wrong string is invisible and a null is a flag. The failure being measured is not stupidity — it is a model completing the shape of the answer with the nearest available number, which is exactly what the decoys are designed to bait.
The paid APIs are the ones exposed
The uncomfortable finding is not about models. It is about the commercial extraction services, all three of which were run on free tiers and only with the instruction. Firecrawl made up 24 of 36 missing fields — worse than 13 of the 16 models, by non-overlapping confidence ranges — and all 24 answers copied the decoy. ScrapingBee made up 16 of 36; ScrapeGraphAI, 7 of 31. Meanwhile a plain HTTP fetch, HTML stripped to text, piped into GPT-6 Luna made up 5 of 36, at a total run cost of $0.0049 for all 84 pages.
The report also tested a cheap second pass: ask a small model whether the page supports each returned value. GPT-6 Luna caught 38 of 49 made-up values and wrongly rejected 0 of 47 correct ones; on Firecrawl's 24 fabrications it caught 20. A cheaper decision model, Jev 1.13, caught 23 of 49 and missed every near-meaning case — cooking or total time returned as prep time, 0 of 6 caught. Checking all 126 returned page-and-value pairs cost $0.0049 with Luna, $0.0024 with Jev.
Caveats are the publisher's own: one run per contestant, no repeats, synthetic pages with traps the team wrote themselves, paid plans untested, and a marketplace business model that benefits from exactly this kind of comparison shopping. The site states plainly that a listing is not a score and that it does not verify provider claims.
Questions You Should Be Asking
- Does your extraction vendor's prompt contain an explicit null instruction — and will they show you the exact wording, or is it your job to supply it?
- When a field is absent from a source page, what lands in your database: a null, an empty string, or last quarter's number?
- If your pipeline is paying per credit for an API that copies decoys, what is that premium buying over a plain fetch plus a model at half a cent per run?
- Who inside your team has ever tested against a page where the answer was deliberately removed, rather than one where it was present?
- Would a one-run, synthetic-page benchmark from a company selling agent listings survive a second run — and has anyone asked for one?
What To Watch Next
The signal to watch is whether anyone reproduces this on real sites rather than synthetic ones, and whether the paid APIs run on their paid tiers respond with their own numbers. Firecrawl, ScrapingBee and ScrapeGraphAI now have a published figure attached to their names from a single free-tier run; silence from all three would be more telling than a rebuttal from one.
- 1Add "Use null for any field whose value is not on the page. Do not guess." to every extraction prompt — it cut fabrication from 71% to 20%.
- 2Test your extractor on pages with deliberately missing fields, like a struck-through old price with no current price, and count invented values.
- 3Treat null as a valid success response in your schema and downstream code, so models are never forced to fill a field to pass validation.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
