Winner sets are unstable
In 12-month rolling evaluations, winner-set switch rates range from 17.9% to 38.5% across target-horizon combinations, with a mean switch rate of 29.1%.
A model can post the lowest average forecast error without establishing a statistically distinguishable or persistent advantage. This paper evaluates those properties separately in a common data-rich macroeconomic forecasting experiment.
The analysis uses the June 2026 FRED-MD vintage and evaluates twelve forecasting approaches across CPI inflation, PCE inflation, industrial production growth, and the monthly change in the unemployment rate at 1-, 3-, 6-, and 12-month horizons. The central question is not only which model ranks first on average, but how much confidence can be placed in that ranking.
Four headline statistics summarize the gap between numerical ranking, formal inference, and temporal stability in the reported pseudo-out-of-sample exercise.
Four U.S. monthly targets are evaluated with the June 2026 FRED-MD vintage under common predictor screening, lagging, and missing-data rules.
Direct forecasts use 360-month rolling estimation windows and up to 90 target months from December 2018 through May 2026, before target-specific exclusions.
Nine individual models span persistence, autoregression, regularized linear models, PCA factors, and tree learners; mean, median, and inverse-RMSE combinations are added.
Numerical leadership is much stronger than the evidence of statistical separation.
XGBoost's average relative RMSE of 0.821 corresponds to a 17.9% average reduction relative to the persistence benchmark across the sixteen target-horizon combinations.
None of the 176 squared-error Diebold-Mariano comparisons is significant in favor of an alternative after Holm adjustment. Five absolute-error comparisons reject, all at one-month industrial production.
The 90% Model Confidence Set reaches the same conclusion from a multi-model perspective: all twelve approaches survive in 25 of 32 panels, and no panel retains fewer than ten. Survival does not prove equal accuracy; it indicates that the available loss data provide limited evidence for eliminating most candidates.
The identity of the leading approach changes over time, and a small number of difficult months account for much of the measured advantage and deterioration.
In 12-month rolling evaluations, winner-set switch rates range from 17.9% to 38.5% across target-horizon combinations, with a mean switch rate of 29.1%.
Removing only the largest persistence-benchmark squared-error date within each target-horizon combination raises the best average relative RMSE from 0.821 to 0.916. After twelve such exclusions, the persistence benchmark leads on average.
Historical values come from one revised June 2026 FRED-MD vintage. The exercise does not reconstruct real-time data vintages, release lags, or historical revisions.
The primary evaluation contains at most 90 target months and includes the COVID-19 collapse and reopening, limiting power and the number of distinct macroeconomic environments.
The conclusions apply to the stated linear, factor, tree-based, and combination methods. Neural networks and time-series foundation models are outside this model set.
Rolling winner-set turnover and loss concentration are descriptive diagnostics. They complement formal inference but do not by themselves establish a general instability regime.
The public replication repository contains forecasting and statistical-analysis code, frozen outputs, paper table and figure data exports, the locked R environment, data-acquisition guidance, and release validation tools. The source FRED-MD CSV is not redistributed.
Research code, reproducibility documentation, frozen outputs, and paper exports.
Open GitHub ↗Frozen public software and replication snapshot for the reported research design and outputs.
Open release ↗Persistent archival record for the versioned replication release.
Open DOI ↗The interactive forecasting dashboard provides a reader-facing interface to explore model performance and related forecasting evidence. It is a research companion to the versioned paper and replication materials, not a substitute for the paper's fixed empirical record.