A June 10 study reports that FuXi-CNOP, an AI and dynamics-based ensemble forecasting system, reduced tropical-cyclone track error by as much as 32.33% against ECMWF's IFS ensemble in retrospective tests. The comparison covered 91 forecasts for 62 storms from 2018 through 2023; performance was similar during the first 12 hours, while the reported advantage emerged at longer lead times.
A peer-reviewed study published June 10 reports that an AI and dynamics-based ensemble system called FuXi-CNOP reduced tropical-cyclone track error by as much as 32.33% compared with the European Centre for Medium-Range Weather Forecasts' Integrated Forecasting System Ensemble Prediction System, or IFS-EPS.
The result comes from retrospective experiments, not a live operational deployment. Researchers evaluated 91 forecasts covering 62 tropical cyclones from 2018 through 2023 in the western North Pacific and North Atlantic. Those years were outside FuXi's stated 1979-2015 training period.
How the ensemble is built
FuXi-CNOP combines Fudan University's FuXi weather model with orthogonal conditional nonlinear optimal perturbations, or O-CNOPs. Instead of adding arbitrary noise to an initial forecast, the method searches for physically constrained perturbations that are expected to grow rapidly under the model's nonlinear dynamics.
The system used 31 members: one control forecast and 30 forecasts created from 15 positive-negative perturbation pairs. The researchers compared it with the 51-member IFS-EPS, which represents initial and model uncertainty using a different combination of perturbation methods.
That design matters because ensemble forecasts are intended to describe a range of plausible storm paths, not produce a single deterministic line. An ensemble can be useful only if its spread reflects the uncertainty seen in actual forecast errors.
What the tests found
The paper reports comparable track performance during the first 12 hours. Beyond 18 hours, FuXi-CNOP's ensemble-mean track errors declined relative to IFS-EPS, with the difference becoming statistically significant at the 90% confidence level from 36 through 120 hours. The largest reported reduction was 32.33%; that figure is a maximum across lead times, not an average improvement for every forecast.
For probabilistic track assessment, FuXi-CNOP reduced the continuous ranked probability score by as much as 29.2%. Its Brier scores for strike probabilities were comparable with IFS-EPS, while its receiver-operating-characteristic performance was marginally worse. The mixed result suggests stronger overall track-distribution scoring without a clean win on every probability metric.
What remains unproven
The comparison covered two ocean basins and a bounded historical period. It did not establish operational reliability across every basin, storm regime, data pipeline or future season. The systems also used different ensemble sizes and methods: FuXi-CNOP represented initial-condition uncertainty, while IFS-EPS also included model uncertainty.
For weather and ML teams, the practical contribution is the perturbation strategy. It shows how automatic differentiation in a learned forecast model can be paired with physical constraints to generate ensemble members that are both computationally tractable and dynamically meaningful. Operational value will depend on prospective evaluation, latency, data availability and performance under conditions not represented in the study.
Key Points #
- 1FuXi-CNOP reduced ensemble-mean tropical-cyclone track error by as much as 32.33% against IFS-EPS across 91 retrospective forecasts for 62 storms.
- 2The system uses physically constrained nonlinear perturbations to create 30 ensemble members around one control forecast rather than relying on arbitrary initial noise.
- 3The reported advantage emerged at longer lead times, while strike-probability discrimination was marginally worse and operational performance remains unproven.
Scoring Rationale #
The study presents a concrete physics-informed method for constructing AI weather ensembles and benchmarks it on 91 out-of-training-period forecasts against a leading operational ensemble. Its methodological and uncertainty-quantification value is notable, while the retrospective scope, mixed probability metrics and lack of prospective operational validation limit immediate deployment claims.
Sources #
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.