bytevyte
bytevyte
Language

Google AI Flu Forecasting Claims Top Spot in CDC's 2025-26 FluSight Season

Google AI flu forecasting produced the top model in the CDC's 2025-26 FluSight evaluation, as ERA wrote a winner that beat 38 rival entries.

Google AI flu forecasting
Photo by Alban on Unsplash

Google's AI-generated forecasting model finished first among 39 eligible entries in the CDC's end-of-season FluSight evaluation, published on September 30, 2026, for the 2025-26 respiratory season. The winning submission, Google_SAI-FluEns, came out of Empirical Research Assistance (ERA), an internal Google tool that writes optimization algorithms, rather than being assembled by hand. The result is the clearest evidence so far that Google AI flu forecasting holds up when it is scored against real hospital admissions instead of benchmark data.

FluSight runs from October through May. Every week the CDC collects forecasts from government, academic and industry teams, each predicting U.S. influenza hospital admissions for the current week and the three weeks that follow. The agency blends the submissions into a single ensemble and uses that pooled figure to describe expected state-level demand for medical services. The season-end analysis released this week found Google's entry tracked observed admissions more closely than any other eligible model.

How ERA Powers Google AI Flu Forecasting

ERA works as a search system rather than a standalone forecasting model: a large language model proposes candidate algorithms, and a tree-search routine scores and refines them, keeping the variants that perform better on the target problem. Applied to influenza, the machine wrote the statistical machinery, selected the features and tuned the parameters, while Google researchers defined the problem and supplied the data.

That division of labour is the strategic point. Most FluSight teams have spent years hand-tuning model families such as mechanistic compartment models, autoregressive statistical models and gradient-boosted trees. The evaluation record shows that accumulated seasons of experience correlate with better scores. ERA compresses that iteration into an automated loop. Google has published the underlying research in Nature, and the technology behind ERA is now open to trusted testers through its experimental science tools.

Why the Win Comes With Qualifiers

First place among individual submissions is not the same as beating the ensemble. In the 2024-25 FluSight evaluation, the CDC's own pooled model outperformed every submitted entry on the primary metric, average relative weighted interval score. A blend has a structural edge, because no single team's bad week drags it as far off course. Google's model winning as a standalone entry, in a season where the ensemble also competed, indicates the automated approach surfaced information the pooled forecast did not already hold.

A second limit is the forecast horizon. FluSight entries reach three weeks ahead, and the CDC notes that ensemble forecasts have struggled when disease trends turn sharply. A model that wins on average interval score across a season is rewarded for calibration over many weeks, not for calling a turning point. Hospital administrators making staffing and supply decisions care most about exactly those weeks.

The scoring history cuts both ways. A decade-long retrospective of the challenge found that in influenza-like-illness seasons, statistical models generally beat mechanistic and machine-learning entries, and that consistent differences became hard to detect once the target shifted to hospital admissions. Machine-learning submissions have not held a structural edge in this contest. ERA's first place is a result about generated algorithms, not a blanket verdict on AI forecasting.

An underrated factor is the data regime. FluSight hospital forecasts have been based on the CDC's National Healthcare Safety Network since the 2021-22 season, replacing the earlier FluSurv-NET basis. A steadier weekly reporting stream gives any algorithm a more stable signal to fit, and it gives an automated search loop more consistent feedback. Part of ERA's gain may come from the target being more tractable than the influenza-like-illness estimates Google Flu Trends chased.

Google has run a version of this experiment before, and the earlier one failed in public. Google Flu Trends launched in 2008 and mined search queries to estimate influenza-like illness. During the 2012-13 season it projected roughly double the level the CDC recorded from laboratory surveillance. Gary King and co-authors at Harvard used the episode to argue that big-data signals need constant revalidation against the ground truth they claim to replace.

Google Flu Trends (2008 onward)ERA / Google_SAI-FluEns (2025-26)
MethodFixed linear regression over search queriesLLM writes algorithms; tree search refines them
TaskNear-real-time estimate of influenza-like illnessOne- to three-week forecast of hospital admissions
Reference dataCDC laboratory surveillanceCDC NHSN reported admissions
Outcome2012-13 season: roughly 2x the CDC levelRanked first of 39 eligible models

The architecture differs in a way that matters. Flu Trends was a fixed linear regression over search terms, presented as an estimate of current activity. ERA generates and re-tests algorithms, and the FluSight task is a forecast scored after the fact against reported admissions. One is a static correlation; the other is an automated search loop judged on out-of-sample error. ERA is not thereby immune to the same failure, but the failure would now surface in a scorecard rather than in a retrospective.

The Trade-offs for Buyers and Rivals

For hospital systems and state health departments, an automated model-writer shifts the build-versus-buy calculation. A team that currently spends a season maintaining a forecasting pipeline could point ERA at its own data and get a tuned algorithm back. The cost moves from modelling labour to validation labour, because someone still has to confirm the generated algorithm is not fitting a data artefact. The gap between a model a human can explain and one a search routine produced is real and will not close on the strength of a single season.

Competing forecasters face a different squeeze. ERA's edge lies in the ability to run thousands of design iterations that a small team cannot, rather than in epidemiological knowledge. That advantage narrows for groups holding proprietary data, such as regional admission feeds Google does not have, and widens for anyone competing on the same public NHSN series. Academic groups have moved in a similar direction: Johns Hopkins' PandemicLLM, built on language models, predicts hospitalization trends one to three weeks out.

For Google, the finish converts a research claim into a validated capability. The company now holds peer-reviewed research in Nature, a first-place result in a government-run evaluation, and a trusted-tester program it can turn into partnerships with health systems, insurers or public agencies. The sequence is deliberate: publish, compete, restrict access, then sell.

Open questions remain. Google has not said whether it will enter FluSight next season, and a single season is a small sample; a decade-long retrospective of the challenge found that the number of seasons a team has entered correlates positively with its accuracy. Nor has Google disclosed how much human input shaped the winning entry, which determines whether ERA is a tool for specialists or a substitute for them. Pricing and access terms for the trusted-tester preview are also undefined.

The Verdict

The evidence supports using AI-generated algorithms as a candidate generator inside an established forecasting pipeline, with the ensemble and human review kept in place. The marker to watch is the 2026-27 season: a second Google entry would test whether ERA's edge is repeatable, and a published account of how much of the model humans touched would settle whether the win belongs to the tool or to the researchers operating it.

Why this matters

The FluSight result is a proof point for one specific claim: an AI system can generate a scientific model that outperforms models people built by hand, in a contest with real consequences and a government referee. If that holds across seasons and across diseases, the bottleneck in operational forecasting shifts from modelling skill toward data access and validation capacity. Google's own history with Flu Trends is the reason to test the claim season by season rather than treating one first place as settled.

Sources

Google's AI ranks #1 for predicting flu hospitalizations.

Photo by Alban on Unsplash

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.