Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New R package puts an end to data snooping in forecast comparisons

October 1, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
New R package puts an end to data snooping in forecast comparisons

New R package puts an end to data snooping in forecast comparisons

New R package puts an end to data snooping in forecast comparisons

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every forecaster faces a seductive trap. When dozens of competing models are lined up against a benchmark, at least one of them will almost always look brilliant by pure chance — and if you cherry-pick that lucky winner, you have fallen victim to what econometricians call data snooping bias. The problem, formalized by Halbert White in his landmark 2000 Reality Check paper, is that running many pairwise comparisons inflates the probability of falsely declaring a winner far beyond the nominal significance level. A new open-source software package called RCtest, described in the journal SoftwareX by Joanna Jędrzejewska and Krzysztof Drachal of the University of Warsaw, now gathers the entire arsenal of Reality Check methodology into a single, coherent R workflow, promising to make rigorous multi-model forecast evaluation accessible to anyone who can install a package from CRAN.

The core insight behind White’s Reality Check is deceptively simple. Instead of testing each competing model separately, the Reality Check tests a joint null hypothesis: that no competing model has lower expected loss than the chosen benchmark. The test statistic is the maximum across models of the sample mean loss differential, and because its distribution under the null is analytically intractable, p-values are obtained through bootstrap resampling. RCtest implements this using the Moving Block Bootstrap of Hans Künsch, which resamples overlapping blocks of consecutive observations to preserve the short-run autocorrelation structure of time series data — a prerequisite for valid inference when forecasts are serially dependent, as they almost always are in economics and finance.

White’s original test, however, is known to be conservative: it can fail to detect genuinely superior models, particularly in small samples. Peter Hansen’s 2005 Superior Predictive Ability test addresses this by studentizing each model’s mean loss differential with its own heteroskedasticity-and-autocorrelation-consistent standard deviation, sharpening the test’s power. RCtest reports two p-values for the SPA test: a consistent p-value, obtained by recentering the bootstrap statistics at each model’s sample mean loss differential, and a conservative p-value computed under the so-called Least Favourable Configuration, which is equivalent to White’s original bootstrap. The consistent p-value is the recommended output and is always no greater than its conservative counterpart — a property the package’s automated test suite explicitly verifies.

Beyond point forecasts, the package reaches into the harder territory of density forecast evaluation. The Kullback-Leibler Information Criterion test of Corradi and Swanson compares predictive distributions by evaluating negative log-likelihood scores at each realized outcome, under a Gaussian predictive density assumption; lower expected negative log score corresponds to lower Kullback-Leibler divergence from the true data-generating process. A companion ZP test examines tail probability accuracy by scoring whether realized outcomes fall below a user-specified threshold, typically the fifth percentile, against the probabilities implied by each model’s predictive distribution. Both use SPA-style studentization and Moving Block Bootstrap inference. The package also computes the Continuous Ranked Probability Score, a strictly proper scoring rule that jointly rewards calibration and sharpness, and applies a CDF-based Reality Check comparison to series of CRPS loss differences.

Conditional superiority is another dimension the package captures. The Conditional Predictive Ability test of Giacomini and White, extended to the multiple-model setting with Hansen-type studentization, weights each loss differential series by a conditioning instrument observable at the time the forecast is made. Recommended instruments include the absolute value of realizations as a volatility proxy, lagged realizations, or any external economic indicator. A constant instrument reduces the test to the unconditional Reality Check, which the authors suggest as a built-in sanity check. This matters because a model may dominate only in turbulent periods — during the 2020 pandemic shock, say, or the 2021–2022 commodity supercycle — and such state-dependent superiority can escape entirely from unconditional tests.

To demonstrate the machinery, the authors apply RCtest to a built-in dataset of 165 monthly observations of World Bank base metals price index forecasts from March 2011 to November 2024, spanning fourteen competing models including Bayesian Dynamic Mixture Models, Dynamic Model Averaging, Bayesian LASSO and RIDGE regressions, time-varying parameter regression, and a simple AR(1) benchmark. Using the historical average as the realized series and AR(1) as the benchmark, the joint tests delivered a strikingly consistent verdict: White’s Reality Check, the consistent SPA test, and the CPA test all rejected the null across MSE, MAE, and MASE loss functions, with p-values below 0.01 in most cases, while the conservative SPA variant failed to reject — exactly the conservatism it is designed to exhibit. The CRPS-based CDF comparison also rejected, whereas the KLIC and ZP distributional tests favored the benchmark.

The package’s credibility rests on an unusually thorough validation program. Monte Carlo experiments with 500 replications per configuration assessed finite-sample behavior across sample lengths of 100 to 250 observations, autoregressive dependence coefficients from 0 to 0.6, and varying cross-model correlation and heteroskedasticity. Under the baseline null, rejection rates hovered near nominal size — 7.2 percent for WRC, 8.4 percent for SPA, 7.6 percent for CPA — while under the baseline alternatives the tests detected inferior benchmarks with power reaching 84.4 percent for WRC and 85.8 percent for SPA, and effectively 100 percent for the KLIC test. Selected routines were also checked against independent references: Diebold-Mariano results matched the forecast package’s dm.test() decisions for all thirteen model comparisons, CRPS values converged to the closed-form Gaussian result from scoringRules, and the Kupiec likelihood ratio statistic agreed with ExactVaRTest to numerical precision.

Practical usability clearly drove the design. A single high-level function, run_comprehensive_erc_analysis(), orchestrates the entire pipeline from raw forecast matrices through nine test variants across three loss functions, returning structured hypothesis-test objects compatible with base R’s printing conventions. Companion functions flatten all results to spreadsheet-ready data frames, generate automated Markdown narrative reports, and produce publication-quality ggplot2 visualizations — including cumulative loss difference plots that directly label the best and worst performers, and scatter plots linking forecast accuracy to portfolio risk-contribution weights. Runtime benchmarks on an Apple M2 machine show the full WRC-SPA-CPA battery completing in roughly a tenth of a second for thirteen models with 999 bootstrap replications, and under a second even with thirty models and nearly five thousand replications. The code is released under the GPL-3 license, with automated testing across Ubuntu, macOS, and Windows, and an archived Zenodo release for reproducibility.

The significance of this release extends well beyond econometrics. The authors point to applications in financial risk backtesting, where the package’s Kupiec unconditional coverage test verifies whether Value-at-Risk forecasts breach at their nominal rate; in epidemiological forecast comparison, where model tournaments are now routine; and in climate model selection. In the illustrative metals application, no model violated the 5 percent VaR coverage null, while the per-model tail probability scores ranked the Bayesian models most favorably. Perhaps most tellingly, the cumulative loss plots revealed that no single model dominated uniformly: several Bayesian models built an early advantage that eroded during the high-volatility years of 2020 to 2022. That kind of honest, jointly tested nuance — rather than a cherry-picked champion — is precisely what the Reality Check tradition was invented to deliver, and what RCtest now makes routine for the R community.

Subject of Research: An R package for comprehensive multi-model forecast evaluation using Reality Check and predictive density tests

Article Title: RCtest: An R Package for comprehensive forecast evaluation via reality check and predictive density tests

Article References: Jędrzejewska, J., & Drachal, K. (2026). RCtest: An R Package for comprehensive forecast evaluation via reality check and predictive density tests. SoftwareX, 36, Article 103078. https://doi.org/10.1016/j.softx.2026.103078

Image Credits: AI Generated

DOI: 10.1016/j.softx.2026.103078

Keywords: forecast evaluation, data snooping, Reality Check test, R package, econometrics, bootstrap methods, predictive density, CRPS, Value-at-Risk, SPA test, Diebold-Mariano test, open-source software

Cite Scienmag News

Denise Maddox. (October 1, 2026). New R package puts an end to data snooping in forecast comparisons. Scienmag. https://scienmag.com/new-r-package-puts-an-end-to-data-snooping-in-forecast-comparisons/

Denise Maddox. "New R package puts an end to data snooping in forecast comparisons." Scienmag, 1 October 2026, https://scienmag.com/new-r-package-puts-an-end-to-data-snooping-in-forecast-comparisons/. Accessed 1 October 2026.

Denise Maddox. "New R package puts an end to data snooping in forecast comparisons." Scienmag. October 1, 2026. https://scienmag.com/new-r-package-puts-an-end-to-data-snooping-in-forecast-comparisons/

Tags: bootstrap methodsCRPSdata snoopingdata snooping correctionDiebold-Mariano testeconometricseconometrics model testingforecast comparison biasforecast evaluationjoint null hypothesis testing in forecastingmodel performance comparisonmulti-model forecast validationopen-source forecast assessment toolsopen-source softwarepredictive densitypreventing cherry-picking in forecastsR packageR package for forecast evaluationReality Check methodologyReality Check testsoftware for robust forecast evaluationSPA teststatistical significance in model selectionValue-at-Risk
Share26Tweet16
Previous Post

Moore Foundation Backs UC Riverside Physicist to Probe Quantum States in 2D Materials

Next Post

Surgical Residents Recover From Burnout as Training Progresses, Landmark Mexican Cohort Finds

Related Posts

AI Learns to Fill the Gaps: Language Models Supercharge Industrial Knowledge Graphs
Technology and Engineering

AI Learns to Fill the Gaps: Language Models Supercharge Industrial Knowledge Graphs

October 1, 2026
Topology-Aware Reformulation Supercharges Quantum Annealing for Planar Optimization
Technology and Engineering

Topology-Aware Reformulation Supercharges Quantum Annealing for Planar Optimization

October 1, 2026
New AI Model Reads the Hidden Rhythms of Evolving Knowledge Graphs
Technology and Engineering

New AI Model Reads the Hidden Rhythms of Evolving Knowledge Graphs

October 1, 2026
From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI
Technology and Engineering

From CLIP to Llama 4: A sweeping review maps the rise of pretrained multimodal AI

October 1, 2026
Liquid Metals Could Finally End the Trade-Off Between Stretchy Circuits and Real Performance
Technology and Engineering

Liquid Metals Could Finally End the Trade-Off Between Stretchy Circuits and Real Performance

October 1, 2026
Robots Learn When You Feel Unsafe: New Framework Tunes Speed and Distance in Real Time
Technology and Engineering

Robots Learn When You Feel Unsafe: New Framework Tunes Speed and Distance in Real Time

October 1, 2026
Next Post
Surgical Residents Recover From Burnout as Training Progresses, Landmark Mexican Cohort Finds

Surgical Residents Recover From Burnout as Training Progresses, Landmark Mexican Cohort Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • AI Learns to Fill the Gaps: Language Models Supercharge Industrial Knowledge Graphs
  • Surgical Residents Recover From Burnout as Training Progresses, Landmark Mexican Cohort Finds
  • New R package puts an end to data snooping in forecast comparisons
  • Moore Foundation Backs UC Riverside Physicist to Probe Quantum States in 2D Materials

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading