Bitcoin price models: AI loses to naive forecasts in academic tests

Share

A systematic review conducted by Carlos Baquero, a researcher at the University of Porto, concludes that no Bitcoin price prediction model, however sophisticated, has demonstrated durable superiority over naive forecasts. The analysis, covering 23 studies selected from hundreds of publications, challenges the real added value of machine learning, deep learning, and popular valuation frameworks used by the crypto community.

🔑 In brief

  • 23 scientific studies reviewed by Carlos Baquero (University of Porto, May 2026)
  • No durable model superiority over naive benchmarks on 1 to 6 month horizons
  • The Bitcoin market is non-stationary: its statistical structure evolves each cycle
  • Stock-to-flow, Metcalfe, and power law models fail in out-of-sample validation
  • Eugene Fama estimates the probability that Bitcoin becomes worthless by 2036 at nearly 100%

The non-stationarity problem of the Bitcoin market

The central finding of Baquero’s review is methodological: relationships between Bitcoin market variables do not remain stable over time. This property, called non-stationarity, means that a model calibrated on past data does not correctly describe future conditions.

Market structure has indeed changed dramatically between cycles. A model trained on the retail-driven cycle of 2017 faces a totally different derivatives structure in 2021. The arrival of spot Bitcoin ETFs in 2024 then created a new channel for institutional capital flows and price discovery, further altering the correlations on which earlier models relied.

Thus, the few independent market cycles Bitcoin has experienced do not compensate for intra-cycle data abundance: millions of one-minute bars repeat observations from the same 2018 bear markets or the same 2020 liquidity shocks, without providing structurally new information.

Academic studies confirming the superiority of naive forecasts

The most damaging study for sophisticated models is the one led by Francesco Puoti, Fabrizio Pittorino, and Manuel Roveri. It compares twelve statistical, machine learning, and deep learning approaches on five major crypto assets over 1-day, 7-day, and 30-day horizons.

ModelCategoryOut-of-sample performance
Naive forecast (persistence)BaselineBest at 1, 7, and 30 days
ARIMAClassical statisticsBelow naive
Prophet (Meta)Time seriesBelow naive
Random forestsMachine learningBelow naive
XGBoostGradient boostingBelow naive
LSTM networksDeep learningBelow naive
N-BEATSDeep learningBelow naive
Synthetic comparison from the Puoti, Pittorino, and Roveri study on five major crypto assets

Results show that simple naive models systematically outperformed all tested approaches. For an asset as persistent as Bitcoin, predicting that tomorrow’s price will be close to today’s price is a remarkably hard benchmark to beat.

Alexander Shelton published a peer-reviewed study in 2024 on Bitcoin return prediction. He found that stock-to-flow (S2F) and Metcalfe variables helped explain in-sample returns, but offered limited to no out-of-sample predictive power. Once time effects were introduced into the S2F regression, its statistical strength disappeared.

The explanation is almost tautological: Bitcoin’s supply ratio rises on a predetermined schedule, and its price has also risen for much of its history. Two time-trending series appear correlated without any causal link existing.

Savva Shanaev and co-authors pushed the analysis further using instrumental variables across six proof-of-work assets. Once autocorrelation and the bidirectional relationship between network activity and price were accounted for, the positive effects attributed to hashrate and transaction count vanished.

The power law model, which describes a stable mathematical relationship between price and time, has not yet completed the formal tests required according to Baquero: sensitivity to start date and performance on future observations remain to be validated.

The overfitting trap and misleading backtests

The study by Ritwik Dubey and David Enke, published in Machine Learning with Applications in June 2025, used on-chain data with the Boruta feature selection algorithm combined with a CNN-LSTM model. The reported results are spectacular: 82.03% accuracy on next-day direction, annualized return of 1,682.7%, and a Sharpe ratio of 6.47.

But sources stress that a claimed 99% accuracy may simply measure how closely a predicted price tracks the real price, an easy task for a persistent series. A model predicting $100,500 when Bitcoin moves from $100,000 to $99,500 shows low price error but still executes the wrong trade.

« There are always incentives for people to corrupt the blockchain. »

Eugene Fama, Nobel laureate in economics, Capitalisn’t podcast

The backtest overfitting phenomenon was formalized by David Bailey and co-authors. Testing more model variations mechanically increases the chance of finding an excellent historical result by pure chance. Among the papers reviewed by Baquero, none evaluated the same approach across multiple non-overlapping validation windows covering different market regimes.

Toward a rigorous forecasting standard

An honest forecasting standard requires several cumulative conditions. First, publish the naive benchmark alongside the proposed model to demonstrate a real gain rather than a statistical illusion. Second, report performance regime by regime: a model that shines during the 2020-2021 bull cycle may collapse in a prolonged bear market.

It also requires including trading costs in any evaluation, because high-rotation strategies do not survive real-world slippage. Finally, disclose the number of variations tried: this number determines whether the winning backtest is statistically meaningful or simply lucky.


Conclusion: caution as the best strategy

The message of Baquero’s review is not a rejection of machine learning applied to crypto, but a reminder of statistical rigor. Bitcoin’s abundant data masks a reality: few independent cycles, a lot of noise, and a market structure in perpetual evolution. For investors, the lesson is pragmatic: treat any promise of accuracy above 80% on Bitcoin direction with a healthy dose of skepticism, and never neglect the value of the naive benchmark as an honest comparison point.

Sources

This article is published for informational and educational purposes only. It does not constitute investment advice in any form. Do your own research (DYOR) before making any decision.

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles