bytevyte
bytevyte
Language
ai-beats —

AI Medical Coding Tied to $942M in Added Insurer Costs, Blue Cross Study Finds

AI medical coding

Blue Cross Blue Shield Association has linked hospital adoption of AI medical coding tools to about $942 million in additional expenses for its member health plans, measured across 2024 and 2025 against a 2023 baseline. The association's study, released this week, attributes $653 million of that total to secondary diagnoses surfaced across more than 55,000 inpatient cases, roughly $11,000 per excess complex case. The tool categories it names are ambient clinical scribes and automated scanners that comb patient records for additional billable conditions.

BCBSA covers more than 100 million people through 31 independent insurance companies, which gives the estimate reach beyond any single plan's claims data. The association's core claim is a mismatch: documented diagnoses rose far faster than the treatments actually performed. Insurers read that gap as software finding more lucrative codes rather than sicker patients. Hospital groups reject the reading and argue the tools improve documentation accuracy.

Why the AI medical coding figure is a counterfactual

The figure is a comparison, not a ledger total. BCBSA priced 2024 and 2025 claims as if the 2023 coding mix had persisted and labeled the difference new cost. That design isolates coding intensity, which is what the study set out to test, but it also assumes the 2023 baseline was complete. If charts were routinely underdocumented before scribes and scanners arrived, part of the $942 million is revenue that was always owed rather than money the software created.

Both sides are internally consistent, and that is the difficulty. No party in the dispute can observe the counterfactual: what the same charts would have recorded without an AI tool reviewing them. The $942 million is a defensible measure of a documentation shift and, simultaneously, an unverifiable claim about clinical intent.

The baseline year carries the argument. BCBSA treats 2023 as the pre-deployment reference, then prices two later years against it. Shift that reference forward by twelve months, into a period when scribes and scanners were already running, and the measured effect compresses, because the tools sit on both sides of the comparison. The $942 million is sensitive to that single design choice, which is why the disagreement has settled on methodology rather than arithmetic.

The DRG staircase and the two-sided arms race

The payment mechanism explains why coding intensity moves money. Inpatient reimbursement in the United States runs through diagnosis-related groups, and each documented secondary condition can push a case into a higher-paying tier. Secondary diagnoses alone account for $653 million of the total, which fits that pattern. The remainder reflects broader increases in coding intensity across inpatient stays.

Payers run their own automation. Insurers deploy machine learning for payment review and prior authorization, and hospital coding gains are partly clawed back downstream through denials and payment downgrades. That cuts both ways for the headline number: $942 million overstates net leakage to plans, since some of it returns through review, while understating the administrative cost both sides now carry to contest the same claim.

MeasureFigure
Additional member-plan expense, 2024-2025 vs 2023$942 million
Attributable to secondary diagnoses$653 million
Inpatient cases with excess complexityMore than 55,000
Added cost per excess complex caseAbout $11,000
People covered by member plansMore than 100 million

Scribes and scanners reach the bill through different doors

An ambient scribe drafts the encounter note as care is delivered, so it shapes what enters the chart in the first place. A record scanner reviews charts after the fact, hunting for conditions that were treated but never coded. Retrospective scanning is where the sharper billing question sits, because re-reading a chart months later cannot change what happened to the patient, only how completely the episode is described to a payer.

That distinction shapes how each tool gets defended. Scribes can be justified as cutting physician documentation burden, a benefit clinicians feel directly. Scanners have a thinner clinical story and a clearer revenue story, which is why the secondary-diagnosis line carries most of the controversy. Payers use the word upcoding for that pattern; hospitals use documentation improvement. The two terms describe the same added codes and imply opposite intent.

What the return-on-investment question leaves out

For hospital finance leaders the tools have already paid for themselves in billed revenue, which is why adoption spread through inpatient departments. That return is real for the buyer and a cost for the payer, so the transaction is closer to a transfer between two balance sheets than a net gain in care delivered. Vendors selling scribes and chart-review software now carry a second problem: their product categories are named in a document that quantifies payer harm, and that document will appear in contract negotiations and regulatory files.

Federal regulators are weighing new rules on AI use in coverage decisions, which puts scrutiny on the payer side of the same contest. Disclosure requirements aimed at either side would run through the same coding logic that produced the $942 million estimate.

The trade-offs differ by actor:

  • Hospitals can defend complete documentation when the clinical record supports each added code, and lose the argument when blanket deployment produces codes the chart cannot justify.
  • Insurers can cut payments with counter-automation, at the cost of wrongful-denial exposure and member friction.
  • Vendors would benefit most from accuracy benchmarks tying each added code to a documented finding, because that is the evidence the current dispute lacks.

What hospital finance and vendor teams should watch

For hospital finance leaders, the practical exposure is contract language. Payer agreements that add audit rights or clawback provisions for AI medical coding convert a revenue gain into a contingent liability, and the $942 million estimate gives payers a number to negotiate against. A hospital that cannot produce the clinical finding behind each added secondary diagnosis is negotiating from weakness, whether or not the code was correct.

For vendors, the study is a sales-cycle problem. Buyers evaluating AI medical coding and scribe tools will ask for accuracy evidence tied to the clinical record rather than coding volume, and the categories named in the BCBSA study now carry a documented cost figure that payers will cite. Vendors that can show an added code matched a documented condition have an answer; those that cannot are selling into a dispute they did not choose.

What would settle the argument

Three things would move the dispute from assertion to evidence. BCBSA could publish case-level methodology so outside researchers can test the 2023 baseline for systematic undercoding. Hospital groups could release a matched clinical severity analysis showing the excess cases genuinely involved sicker patients. Regulators could require either side to disclose the logic behind automated coding and payment decisions. None of that exists yet, which leaves the honest reading of the $942 million as a measurement of documentation behavior rather than of patient health.

The stakes differ in kind. Hospitals face a compliance question with financial penalties attached. Insurers face a member-experience and litigation question if counter-automation denies care the chart supports. Vendors face a credibility question that a benchmark could answer. That asymmetry explains why the loudest objection to the study came from providers, who have the most to lose if the coding-intensity interpretation becomes the accepted one.

Why this matters

Plan sponsors and employers will absorb the $942 million through premiums long before anyone verifies whether the extra codes reflected extra illness. The dispute also sets a template for other sectors where one party's automation generates a bill another party pays: the return is real to the buyer, the counterfactual stays unpublished, and the argument ends in regulation rather than evidence. The markers to watch are whether BCBSA's methodology survives independent review and whether federal rules on AI in coverage decisions force disclosure from hospitals or insurers.

AI-generated image.

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.