Academic Research

Beyond Average Accuracy

Statistical Distinguishability and Temporal Concentration in Data-Rich Macroeconomic Forecasting
Sungkyu LeeAI Computing Program, Graduate School of Computing, Yonsei UniversityAcademic Working PaperSSRN 7164118August 2026

Overview

A model can post the lowest average forecast error without establishing a statistically distinguishable or persistent advantage. This paper evaluates those properties separately in a common data-rich macroeconomic forecasting experiment.

The analysis uses the June 2026 FRED-MD vintage and evaluates twelve forecasting approaches across CPI inflation, PCE inflation, industrial production growth, and the monthly change in the unemployment rate at 1-, 3-, 6-, and 12-month horizons. The central question is not only which model ranks first on average, but how much confidence can be placed in that ranking.

Research Question

When machine-learning forecasts beat a simple benchmark on average, is that advantage statistically distinguishable and persistent through time, or concentrated in a small number of difficult months?

Key Findings

XGBoost records the lowest average relative RMSE, 0.821, across sixteen target-horizon combinations, with Random Forest, Boruta RF, and the forecast combinations also delivering sizable numerical gains.
After Holm adjustment, none of the 176 squared-error Diebold-Mariano comparisons rejects equal predictive accuracy in favor of an alternative. Five absolute-error comparisons reject, all at the one-month industrial-production horizon.
A 90% Model Confidence Set retains all twelve approaches in 25 of 32 target-horizon-loss panels and never retains fewer than ten, indicating broad model-selection uncertainty in the available sample.
The leading set changes frequently through time, and loss differences are highly concentrated. In the squared-loss diagnostic, the twelve largest favorable monthly loss reductions account for 85.6% of gross improvement on average, while the twelve largest unfavorable increases account for 89.1% of gross deterioration.

Selected Evidence

Four headline statistics summarize the gap between numerical ranking, formal inference, and temporal stability in the reported pseudo-out-of-sample exercise.

0.821
XGBoost mean relative RMSE across 16 target-horizon combinations
0 / 176
Holm-adjusted squared-error DM rejections in favor of alternatives
25 / 32
MCS panels retaining all 12 approaches
29.1%
Mean switch rate of the 12-month rolling RMSE winner set
Under the squared-loss concentration diagnostic, the twelve most favorable monthly loss reductions account for 85.6% of gross improvement on average, while the twelve most unfavorable monthly loss increases account for 89.1% of gross deterioration relative to the persistence benchmark.

Forecasting Design

01 · DATA

Common FRED-MD information set

Four U.S. monthly targets are evaluated with the June 2026 FRED-MD vintage under common predictor screening, lagging, and missing-data rules.

02 · WINDOWS

Fixed-length rolling estimation

Direct forecasts use 360-month rolling estimation windows and up to 90 target months from December 2018 through May 2026, before target-specific exclusions.

03 · MODELS

Twelve forecasting approaches

Nine individual models span persistence, autoregression, regularized linear models, PCA factors, and tree learners; mean, median, and inverse-RMSE combinations are added.

Inference & Model Uncertainty

Numerical leadership is much stronger than the evidence of statistical separation.

Average accuracy
17.9%

XGBoost's average relative RMSE of 0.821 corresponds to a 17.9% average reduction relative to the persistence benchmark across the sixteen target-horizon combinations.

Formal separation
0

None of the 176 squared-error Diebold-Mariano comparisons is significant in favor of an alternative after Holm adjustment. Five absolute-error comparisons reject, all at one-month industrial production.

The 90% Model Confidence Set reaches the same conclusion from a multi-model perspective: all twelve approaches survive in 25 of 32 panels, and no panel retains fewer than ten. Survival does not prove equal accuracy; it indicates that the available loss data provide limited evidence for eliminating most candidates.

Temporal Durability

The identity of the leading approach changes over time, and a small number of difficult months account for much of the measured advantage and deterioration.

Rolling rankings

Winner sets are unstable

In 12-month rolling evaluations, winner-set switch rates range from 17.9% to 38.5% across target-horizon combinations, with a mean switch rate of 29.1%.

Benchmark-error sensitivity

Average gains depend on influential dates

Removing only the largest persistence-benchmark squared-error date within each target-horizon combination raises the best average relative RMSE from 0.821 to 0.916. After twelve such exclusions, the persistence benchmark leads on average.

Scope & Limitations

Current-vintage pseudo-out-of-sample design

Historical values come from one revised June 2026 FRED-MD vintage. The exercise does not reconstruct real-time data vintages, release lags, or historical revisions.

Short, shock-heavy evaluation period

The primary evaluation contains at most 90 target months and includes the COVID-19 collapse and reopening, limiting power and the number of distinct macroeconomic environments.

Fixed comparison universe

The conclusions apply to the stated linear, factor, tree-based, and combination methods. Neural networks and time-series foundation models are outside this model set.

Diagnostic, not universal, instability claims

Rolling winner-set turnover and loss concentration are descriptive diagnostics. They complement formal inference but do not by themselves establish a general instability regime.

Reproducibility

The public replication repository contains forecasting and statistical-analysis code, frozen outputs, paper table and figure data exports, the locked R environment, data-acquisition guidance, and release validation tools. The source FRED-MD CSV is not redistributed.

Research Dashboard

The interactive forecasting dashboard provides a reader-facing interface to explore model performance and related forecasting evidence. It is a research companion to the versioned paper and replication materials, not a substitute for the paper's fixed empirical record.

Citation

Lee, S. (2026). Beyond Average Accuracy: Statistical Distinguishability and Temporal Concentration in Data-Rich Macroeconomic Forecasting. SSRN Working Paper 7164118. DOI: 10.2139/ssrn.7164118.