FUURAA AI Knowledge Library · forecasting evaluation

How to evaluate AI forecasting and time-series capability claims

Use six evidence gates to turn “forecast demand more accurately,” “lead on long-horizon forecasting” or “automate time-series analysis” into an exact target and horizon, data-availability time, leakage-free rolling backtests, strong baselines, point and probabilistic accuracy, complete failures, operating cost, regime change and a dated applicability boundary.

Published30 August 2026Evidence statusMethod synthesis grounded in primary forecasting competitions, archives and time-series-model researchScopePoint, probabilistic, hierarchical and long-horizon time-series research

Lower average error is not live forecasting capability; one curve is not transferable evidence

Forecasting capability must keep the target, data availability, backtest, baselines, uncertainty, complete denominator and regime change in one evidence chain.

Forecasting is not a fitting task under an ordinary random split. Evaluation must isolate the future at every forecast origin, use the data vintage actually available then, and measure point and probabilistic forecasts with horizons, losses and baselines aligned to the decision.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, dataset or leaderboard. It is not a guarantee of demand, market, financial, operational, scientific, climate or other outcomes.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the forecasting target, horizon and decision

Decision question
What exact quantity, scope, granularity, cut-off, lead time and decision must be forecast, and which errors matter?
Minimum evidence
Verbatim claim and date; exact system, model, features, training and serving versions; target definition, units, aggregation, frequency, horizon, users, decision, asymmetric loss, thresholds and abstention rules.
Stop condition
Stop when a broad “predicts demand” or “best forecasting model” statement replaces a named target, horizon, cadence, loss and decision.
02

Version series, availability, revisions and covariates

Decision question
What exact values and external variables were available at each forecast origin, and what changed later?
Minimum evidence
Series provenance, snapshot and availability timestamps; original and revised vintages; hierarchy and aggregation map; calendar, price, weather, promotion and policy covariates; missing, censored, anomaly, stock-out and intervention rules.
Stop condition
Stop when final revised values, future-known features or an undocumented cleaning step enter training or evaluation as if available in real time.
03

Isolate the future and reproduce rolling-origin backtests

Decision question
Can every training, validation and test window be rebuilt using only information available at its forecast origin?
Minimum evidence
Chronological splits and rolling origins; embargo or gap rules; feature-computation timestamps; hyperparameter and model-selection boundary; retraining cadence, lookback window and seeds; frozen evaluation code, baselines and full window denominator.
Stop condition
Stop when random splits, overlapping windows, test-set tuning, future aggregates or unavailable covariates make the backtest easier than deployment.
04

Compare strong baselines and evaluate uncertainty

Decision question
Does the system beat relevant naïve and statistical baselines on the metrics and uncertainty needed by the decision?
Minimum evidence
Seasonal-naïve, last-value, drift, moving-average and tuned statistical baselines; MAE, RMSE, MASE, sMAPE or decision loss with declared formulas; quantiles and prediction intervals; calibration, empirical coverage and width; hierarchy, intermittency, horizon and critical-slice results.
Stop condition
Stop when one average error, one benchmark rank or a narrow interval without coverage replaces baseline comparison, slice analysis and calibrated uncertainty.
05

Count every series, window, failure and operating cost

Decision question
What happened across the complete eligible population, repeated runs and tail conditions?
Minimum evidence
All series and forecast origins; empty, late, missing and failed forecasts; low-volume, intermittent, cold-start and regime slices; repeated-run variance; retraining failures, fallbacks, overrides and reconciliation; latency, compute, storage, review time and total cost.
Stop condition
Stop when selected series or mean accuracy hide unavailable forecasts, critical misses, unstable repeats, manual rescue, slow tails or cost.
06

Transfer through regime change and expire the claim

Decision question
Does the result remain useful after demand, policy, price, behaviour, sensors or operations change?
Minimum evidence
Shadow or staged operation; representative freshness and latency; intervention and structural-break tests; drift, calibration and coverage monitoring; human acceptance and downstream decision outcomes; retraining triggers, fallback, rollback, owner and expiry date.
Stop condition
Stop when a static backtest transfers directly to live inventory, staffing, market, financial, scientific or public decisions without target validation and rollback.

Minimum forecasting claim failure matrix

Check these conditions deliberately before transferring selected backtests, average error or benchmark rank into real forecasting capability.

  • 01
    Target or feature leakage lets future information enter training or evaluation

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 02
    Final revised data is used although only an earlier vintage existed at the forecast origin

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 03
    A benchmark horizon, cadence or aggregation differs from the operational decision

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 04
    Average accuracy hides weak rare, low-volume, intermittent or cold-start series

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 05
    A complex model does not reliably beat seasonal-naïve or another tuned baseline

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 06
    Prediction intervals are narrow but miss the realised value more often than claimed

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 07
    Hierarchy reconciliation improves a total while hiding critical item or location errors

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

  • 08
    A promotion, outage, policy shift or structural break invalidates historical evidence

    Record the affected target, series, horizon, data vintage, metric and decision; preserve the narrowest conclusion that remains and specify whether to resplit data, repair leakage, rerun, widen intervals, fall back to a baseline, restrict use or withdraw the claim.

Minimum forecasting capability evaluation record

Let the next reviewer reconstruct the conclusion with the same data vintage, forecast origins, baselines, system and complete windows.

  1. 01Claim, date, owner, target, scope, granularity, horizon, cadence, users and decision loss
  2. 02System, model, features, training, serving, packages and deployment versions
  3. 03Series provenance, snapshots, availability times, revisions, permissions and exclusions
  4. 04Hierarchy, covariates, business definitions, missing, anomaly, censoring and intervention rules
  5. 05Chronological splits, rolling origins, embargoes, lookback, retraining cadence and seeds
  6. 06Naïve and statistical baselines, metric formulas, decision thresholds and reference code
  7. 07Point errors, quantiles, interval calibration, coverage, width and critical slices
  8. 08All series and windows, missing forecasts, failures, repeats, fallbacks and corrections
  9. 09Latency, compute, storage, human review, downstream impact and total operating cost
  10. 10Target evidence, drift monitoring, change triggers, fallback, rollback, owner and expiry

Common evidence states

Bind conclusions to the exact pipeline, target, data vintage, backtest, complete denominator, target workflow and date.

Supported

The exact pipeline beats declared baselines and meets point, uncertainty, critical-slice, availability, latency, cost and decision thresholds on representative target series, with a current review date.

Conditional

Support holds only for named targets, series populations, horizons, data vintages, covariates, regimes or operating controls.

Mixed

Baseline lift, calibration, coverage, availability, critical errors, stability, latency or cost vary materially by horizon, series or regime.

Insufficient

Target, availability vintage, leakage-free backtest, baseline, uncertainty, complete denominator, target transfer or expiry is missing.

FUURAA analysisThe minimum decision unit for an AI forecasting and time-series capability claim is exact system, model, feature, training and serving version × forecast target, scope, granularity, lead time, frequency, user and decision × time-series provenance, availability time, revisions, hierarchy, covariates, missingness and anomalies × splits, rolling backtests, baselines, training, updating and post-processing × point forecasts, intervals, calibration, coverage, critical errors and business loss × all series, windows, failures, variance, human correction, latency and cost × target-workflow boundary and cut-off date. Lower offline average error is an observation, not transferable proof of live capability.

Primary sources and non-transfer boundaries

These sources constrain large competitions, hierarchical retail, cross-domain archives, long-sequence efficiency, decomposition and multi-period modelling.

Sources rechecked 30 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published October 2018M4 Competition — Results, Findings and the Way Forward

Evaluates point forecasts and prediction intervals across 100,000 series and shows the value of strong statistical baselines, combinations and hybrid methods.

BoundaryCompetition rankings over selected frequencies and horizons do not prove accuracy for another population, information set, loss function, regime or live decision.

Open primary source ↗
Published October 2022M5 Accuracy Competition — Retail and Hierarchical Forecasting

Tests 42,840 hierarchical retail sales series with intermittency, explanatory variables, grouped evaluation and strong machine-learning entries.

BoundaryOne retailer, competition data availability and a weighted metric do not establish transfer to another catalogue, hierarchy, revision process, cost function or operation.

Open primary source ↗
Published 14 May 2021Monash Time Series Forecasting Archive

Provides 20 datasets from varied domains, frequencies, lengths and missing-value conditions, plus standard baselines across eight error metrics.

BoundaryA broad public archive improves comparative coverage but does not reproduce a private target series, data latency, revisions, interventions or business loss.

Open primary source ↗
Published 14 December 2020Informer — Long Sequence Time-Series Forecasting

A Chinese-led study proposes sparse attention, distillation and a generative decoder for efficient long-sequence forecasting.

BoundaryResults on selected long-horizon datasets do not establish superiority over tuned simple baselines, leakage-free backtests, calibrated intervals or live reliability.

Open primary source ↗
Published 24 June 2021Autoformer — Decomposition and Auto-Correlation

A Chinese-led study integrates progressive decomposition and auto-correlation to model seasonal and trend structure in long-term forecasting.

BoundaryArchitecture gains on benchmark splits do not prove that the same decomposition, variables, horizons or errors remain suitable after a structural break.

Open primary source ↗
Published 5 October 2022TimesNet — Temporal 2D-Variation Modeling

A Chinese-led study maps multi-period temporal variation into two dimensions for forecasting and other general time-series tasks.

BoundaryGeneral benchmark performance does not establish production data provenance, future-data isolation, probabilistic calibration, operating cost or decision value.

Open primary source ↗

Continue checking

Move from forecasting into data analysis, causality, decision-making or the full library.

Enter the AI Knowledge LibraryEvaluate data analysisEvaluate factualityEvaluate information extractionEnter AI Evidence Atlas