Nobody Measured the 2% Where the Model Rewrites Your Invoice Number
A viral dev post argues half of production AI agents are decision trees with a GPU bill. The scenarios are illustrative — the practitioner agreement is the real data.
The 2% nobody logs
Here is the scenario at the heart of a DEV.to post published this week by Dimitris Kyrkos, a market analyst at Cyclopt in Thessaloniki. A team needs to pull invoice numbers out of incoming emails. The numbers follow a fixed format: INV- plus eight digits. The team sends every email to a large language model with a prompt asking it to extract the number.
It works 98% of the time. In the other 2%, Kyrkos writes, the model "helpfully" reformats the number, or grabs a purchase order number instead. Every call costs money and adds a few hundred milliseconds. A four-line regular expression does the same job deterministically, in microseconds, for effectively nothing.
Be clear about what this is: Kyrkos states plainly that his scenarios are illustrative, not case studies. The 98% is a stand-in, not a measurement. What makes the post worth reading is not the number — it is what happened in the comments.
Three category errors, explained
The argument rests on a distinction that gets lost in AI procurement conversations: the difference between questions about facts and questions about meaning.
A regular expression is a pattern-matching rule. Given "INV- then eight digits," it either finds a match or it does not. There is no probability involved, no temperature setting, no run where it decides to be creative. That is the trade: no flexibility, total predictability.
Vector search works the opposite way. Text is converted into a list of numbers — an embedding — that positions it in a space where similar meanings sit close together. You then retrieve the nearest handful of results, the "top-k." This is genuinely good at "find tickets that resemble this complaint." It is the wrong tool for "show me all unpaid orders from customer 4417 in the last 30 days," which Kyrkos answers with a four-line SQL query. Igor Eduardo, an AI engineer who works in clinical AI and pharmacovigilance, put the problem sharply in the comments: that is not a retrieval-quality issue, it is a category error. A filter query has a complete, correct answer. A vector hit list is a probabilistic shortlist that was never asked to be complete.
The third case is the routing example: enterprise plan plus billing goes to account management, bugs open a ticket, everything else gets the FAQ link. Four branches. Build that as an autonomous agent — a model with tool access, a planning loop and a memory store — and the routing becomes probabilistic, occasionally loops, and nobody can reconstruct why a given ticket went to the wrong team.
What is actually evidenced here
No benchmark, no spend data, no named company. What the thread does contain is unusually specific practitioner agreement. Eduardo says he keeps seeing failure mode 2 inside "knowledge" bots. Brian, who runs an AI news account, says teams treat 98% extraction as good enough and never measure the 2% that quietly rewrites the number — and adds a concrete fix: log the misses as a labeled set, and after a month you will know whether the residue is genuinely messy or just a second pattern you never wrote.
The thread also contains its own counterweight, which is the part most summaries will skip. Michael Hairetis, a full-stack engineer with 25 years in production, says his if-statements became "ai if statements" and calls it a useful trick. Kyrkos concedes the point: when the condition itself is fuzzy, an LLM soft router is a solid pattern — provided that step stays isolated so the rest of the system remains predictable.
Pierre-Laurent Medori, head of backend engineering at GoodBarber, offered the most portable framing. He compares it to Python's stance on classes versus Java's, and cites Jack Diederich's 2012 PyCon talk "Stop Writing Classes": if a class has two methods and one is init, it is a function. Transposed: if your agent has one tool and the planner picks it every time, it is a function call with a token bill. His real point is architectural — keep the step behind a plain function signature and upgrading from if-statement to model call stays a local change. Make the agent framework the paradigm on day one and every caller inherits the dependency.
Questions You Should Be Asking
- For each class of question your system answers, is it exact, a filter, or semantic? If it is exact or filter and the path still runs through embeddings and an agent loop, what is the GPU spend buying you?
- Are you logging the cases where the model's output disagrees with a deterministic check — and has anyone looked at that set?
- When a routing decision is wrong, can someone reconstruct why within an hour, or does the answer involve the phrase "the planner decided"?
- If the failure mode is silent bad data rather than a visible error, why is a nondeterministic component in that path at all?
- Who upgrades, understands and debugs each framework in the stack twelve months from now, and have they agreed to it?
What To Watch Next
Watch whether anyone publishes the measurement this thread is missing: a production team putting real numbers on the residue — how many inputs actually fall through a deterministic path, what the model recovers, what it silently corrupts. Brian's suggestion is the cheapest version of that experiment, and it takes a month. Until someone runs it and shows the data, "half the agents in production are if-statements" remains a well-argued hunch with strong practitioner backing — which is more than most AI architecture advice has, and still less than a decision should rest on.
- 1Validate every LLM extraction against the known format (e.g., regex ^INV-\d{8}$) and route failures to a fallback or human review.
- 2If the target data has a fixed, deterministic pattern, run regex first and call the model only on the strings it fails to match.
- 3Log the disagreement rate between your model output and a rule-based check, so the silent 2% shows up as a metric instead of a customer complaint.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
