Stephan Rasp, Boris Babenko, Dominic Masters, Andrew El-Kadi, Samier Merchant, Guy Shalev, Ilan Price, Fred Zyda
Operational, benchmark-leading global AI weather model from a field-shaping group that demonstrates a compelling direction (raw observation ingestion), though reproducibility is limited and novelty is integration-heavy.
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.
WeatherNext 3 (WN3) tackles two structural limitations of state-of-the-art AI weather models: (a) their exclusive reliance on reanalysis/analysis data for both initialization and training, which caps resolution and inherits analysis biases, and (b) their coarser spatial/temporal granularity relative to the best physics-based NWP systems. WN3's central move is to make an operational global AI ensemble model *multimodal over raw observations* — ingesting low-latency geostationary satellite mosaics to enable hourly (rather than 6-hourly) re-initialization, and directly predicting satellite-derived precipitation (PARDIG/IMERG), tropical-cyclone attributes, and sparse in-situ station observations. A notable architectural contribution is the continuous-query station head that decodes 2m temperature/dewpoint at arbitrary locations and times from an interpolated 0.1° latent conditioned on local metadata (elevation, land/sea mask), effectively fusing data assimilation, forecasting, and post-processing (MOS-style station calibration) into a single trained model. This is a conceptually clean instantiation of the "end-to-end observations-to-forecast" vision that several groups are pursuing.
The evaluation is unusually thorough for a system paper. WN3 is benchmarked against the strongest available baselines per task (WN2, ECMWF ENS, AIFS ENS v2), against multiple independent ground truths (HRES-fc0, held-out METAR/Mesonet stations, MRMS, rain gauges, IBTrACS), and with careful attention to confounds (interpolation/regridding penalties, ground-truth source discrepancies, latency-adjusted "operational lead time" analysis). The held-out-station protocol convincingly demonstrates spatial generalization (only a few percent overfit). The authors are commendably candid about weaknesses: degraded 6–12h analysis skill for some variables, mesh-scale hexagonal artifacts, per-member global temperature bias in the station head, cyclone under-dispersion likely from larger model capacity, and a small (6-week) real-time evaluation sample. The main soft spot is that the flagship real-time comparison against AIFS ENS v2 rests on limited data and involves ground-truth/interpolation asymmetries that the authors themselves flag as partially favoring WN3.
High and largely applied. This is a deployed operational system with a public access path, from a group whose prior models (GraphCast, GenCast, WeatherNext-Cyclones) have already reshaped the field and been adopted operationally. Demonstrating that direct observation ingestion improves skill — a ~30–40% CRPS reduction for near-surface temperature versus analysis-based forecasts and up to 60% for precipitation at short leads — provides a strong empirical argument that will steer the subfield away from analysis-only pipelines. The station head's ability to eliminate a separate post-processing stage has direct economic value for downstream forecasting products (energy, agriculture, aviation, flood/disaster response).
Squarely on the current frontier. The integration of raw observations into global AI models is precisely the emerging bottleneck the community is converging on (Aardvark, Huracan, AIFS-DOP, WeatherMesh). WN3 is among the first to demonstrate this operationally at NWP-competitive resolution rather than as a research prototype.
Strengths: operational deployment, comprehensive multi-ground-truth benchmarking, genuine architectural novelty in continuous sparse decoding and native multi-resolution handling, honest reporting of artifacts. Limitations for scientific impact assessment: reproducibility is severely constrained — PARDIG is a proprietary derived product ("further details in an upcoming paper"), the satellite mosaic and training pipeline require massive TPU infrastructure (~13.7 chip-years across TPUv4/TPU7x), and no code/weights for training are released. Independent verification by academic groups is essentially impossible; the community must take the benchmarks largely on trust. The novelty is incremental at the framework level (building directly on WN2/FGN), with the innovation concentrated in integration and data engineering rather than a new modeling paradigm. Artifacts (hexagonal mesh patterns, temporal discontinuities at 6h boundaries, per-member bias) indicate the marginal-CRPS objective still compromises joint structure.
Additional observations: The paper's main conceptual insight — that augmenting analysis-trained models with observations in both input and output "combines the best of both worlds" — is a load-bearing framing likely to be cited as motivation across the emerging observation-based-forecasting literature. The latency-adjusted evaluation methodology is itself a useful contribution for fairly comparing hourly vs. 6-hourly systems.
Overall, this is a high-impact, deployment-grade advance that will strongly influence the direction of AI weather forecasting, tempered by limited reproducibility and incremental (rather than paradigm-shifting) methodological novelty.
Generated Sep 4, 2026
Operational, benchmark-leading global AI weather model from a field-shaping group that demonstrates a compelling direction (raw observation ingestion), though reproducibility is limited and novelty is integration-heavy.