Insurers say AI coding cost them $1B. Nobody got sicker.
Hospitals bought AI to find billable diagnoses. Insurers bought AI to deny them. Both can show positive ROI while the system gets more expensive.
A billion dollars moved and the patients were unchanged
The Blue Cross Blue Shield Association says insurers paid out nearly $1 billion more than usual across 2024–2025 because of AI-assisted hospital documentation. Hospitals deployed software to scan patient records and surface secondary diagnoses that human coders had missed. More diagnoses on a claim means a higher reimbursement. What did not move, according to BCBSA, was treatment intensity. Its SVP of data science put it plainly: the AI is “identifying more billable conditions, not sicker patients.” Insurers have responded by deploying their own AI to deny those claims.
Two caveats before you build a worldview on this. BCBSA is a trade association for the organisations writing the cheques, so its framing of hospital billing is an interested party's account, not a neutral audit. And the figure reached wide circulation via a widely-shared post on X from the account Hedgie on 27 September, rather than a published methodology you can pick apart. Treat the number as a claim with a motive attached. The mechanism underneath it, though, is real and worth understanding.
How software finds a diagnosis a doctor didn't bill for
Most of what happens in a hospital is written down as free text — admission notes, nursing observations, discharge summaries, lab commentary. Billing, by contrast, runs on structured codes. Human coders read a sample of that text and translate it. They cannot read all of it, and they were never expected to.
A hospital bill is not a flat fee. It is priced partly on how sick the patient is documented to be, so a secondary condition recorded alongside the main reason for admission — a metabolic problem, a chronic condition, a complication — can move the same episode of care into a higher-paying severity tier. The clinical work does not change. The invoice does.
Language models are good at exactly the task standing between those two facts: read everything, flag anything that could be coded. This is the rare enterprise AI deployment with a clean, immediate, measurable payback, which is precisely why it spread. On the insurer side, the mirror-image tool reads incoming claims and flags those whose documented severity looks unsupported. Same technology, opposite sign.
The signal: AI's first mass-market ROI case is adversarial
Most corporate AI spending is still justified with soft numbers — time saved, productivity uplift, tickets deflected. Revenue-cycle AI is different. The return lands in the general ledger within a quarter, which makes it one of the few genuinely proven business cases in the market. That is the uncomfortable part. The clearest ROI in enterprise AI so far comes from moving money between two parties, not from creating anything.
Both sides can be rationally correct and the system still gets worse. The hospital's tool pays for itself. The insurer's tool pays for itself. The vendors are paid by both. The combined cost, per the argument in the source, lands on premiums. This is the pattern to watch for wherever two counterparties share a contract and a dispute process: procurement, freight claims, litigation discovery, insurance of every kind. Once one side automates, the other has no choice.
Note also what BCBSA's argument implies about detection. Its evidence is not that the diagnoses were fabricated — it is that documented severity rose while treatment did not. That divergence test is easy to run at scale, and it is the template any regulator or auditor will reach for first.
Questions You Should Be Asking
- If a vendor's ROI comes entirely from your counterparty's pocket, what happens to that ROI when the counterparty automates too — and is that scenario in the business case you signed off?
- Can your documentation AI show that flagged conditions correlate with a change in treatment, or only with a change in reimbursement?
- Who bears the legal exposure when an AI-surfaced diagnosis is later ruled unsupported: the vendor, the coder who accepted the suggestion, or the clinician whose notes it was drawn from?
- On the denial side, what is the appeal-overturn rate for automated denials, and is anyone measuring the ones patients never appeal?
- Is BCBSA publishing a methodology for the $1bn figure, or asking you to take a number from an interested party at face value?
What To Watch Next
Watch whether the divergence test — documented severity rising while treatment intensity stays flat — shows up in a regulatory or payer audit standard rather than a trade-association talking point. The moment that comparison becomes a formal trigger for review, every hospital running documentation AI needs an answer for it, and the market for these tools reprices overnight.
- 1Track coding-intensity metrics (risk scores, CC/MCC capture rates) against treatment-intensity metrics like length of stay and drug spend to spot billing-only shifts.
- 2Require that any AI-surfaced secondary diagnosis be tied to documented clinical evidence and a treatment or monitoring plan before it hits a claim.
- 3Treat trade-association loss figures as advocacy, not audit: ask for methodology, baselines, and independent data before citing the $1B number.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
