Does AI trading actually work? What the independent evidence shows
On the independent published record, AI trading does not reliably beat a low-cost index after costs. The Eurekahedge AI Hedge Fund Index returned 9.8% annualised from December 2009 to July 2024 against 13.7% for the S&P 500. A peer-reviewed review of 27 machine-learning equity experiments found no conclusive evidence of machine-learning funds delivering returns at scale. AI does show measurable value elsewhere — in execution cost, surveillance and fraud detection.
Published 27 August 2026 · Updated 28 August 2026 · AI Trading Book Editorial · Reading time about 14 minutes
- Professional AI funds have underperformed a passive benchmark over fifteen years, and underperformed it again in the one documented stress month.
- The academic record is negative on excess returns and positive on the reasons why: selection bias, look-ahead bias, survivorship bias and omitted costs.
- No audited retail benchmark exists. Any "best-performing AI bot" ranking is built on vendor claims or backtests.
- Where AI does work is unglamorous: execution cost, fraud detection, AML, surveillance, research throughput.
- The honest benchmark is a low-cost index fund — out of sample, after fees, spread, slippage and tax.
The claim being tested
Vendor material generally makes one of three claims: that AI produces excess returns, that AI predicts direction with high accuracy, or that automation removes the behavioural errors that cost traders money. These are separate claims with separate evidence, and they are worth separating before looking at data. The first is a performance claim, testable against benchmarks. The second is a statistical claim, testable against out-of-sample accuracy. The third is a behavioural claim, and it is the only one where automation has an unambiguous mechanical advantage — though not the one usually advertised.
Throughout this page, "works" means one thing: produces a net result, after all costs, that a plainly available alternative does not. The alternative used is a low-cost index fund, because that is what a retail participant can buy instead in about four minutes.
The professional fund record
The best-known index of AI-driven hedge funds has trailed the S&P 500 across every window for which figures were published. This is the strongest single piece of evidence available, because it measures live capital managed by well-resourced professionals rather than simulations.
| Window | AI index | Benchmark | Source |
|---|---|---|---|
| Dec 2009 – Jul 2024, annualised | 9.8% | 13.7% (S&P 500) | Eurekahedge, via IG |
| Jan 2011 – Jan 2020, cumulative | 115% | 210% (S&P 500); 133% (MSCI World) | Buczynski et al. 2021 |
| January 2022, monthly | −3.38% | −1.11% (S&P 500) | Eurekahedge, via Fortune |
Three caveats belong with this table, and publishing them is part of taking the evidence seriously. Index construction for hedge-fund indices involves self-selection and survivorship effects that generally flatter the index rather than penalise it. The constituent set is not disclosed at the level needed to reproduce the calculation. And a fifteen-year window that contains one of the strongest equity bull runs in history is a demanding benchmark for any strategy that holds cash or hedges. None of these caveats reverses the direction of the finding; all three should temper how confidently it is stated.
The academic record
Peer-reviewed work has repeatedly failed to find evidence of machine-learning strategies delivering large returns at scale, and has produced a well-developed explanation of why apparent evidence keeps appearing anyway.
The 27-experiment review
Buczynski, Cuzzolin and Sahakian reviewed 27 machine-learning equity experiments in the International Journal of Data Science and Analytics in April 2021 and concluded there was "no conclusive evidence of any ML-driven investment funds delivering spectacular returns at scale". The paper's value is not the verdict alone but its documentation of methodological heterogeneity: the studies differ in universe, horizon, cost assumptions and validation design so substantially that they cannot be pooled into a single estimate.
The replication problem in factor research
Harvey, Liu and Zhu argued in the Review of Financial Studies in January 2016 that "most claimed research findings in financial economics are likely false", and proposed that a newly claimed factor should clear a t-statistic above 3.0 rather than the conventional 2.0 — because hundreds of factors have been tested and published, and the conventional threshold does not account for that search. If this is true of academic factor research conducted under peer review, the implication for unreviewed vendor backtests is straightforward.
Backtest overfitting as a formal problem
Bailey, Borwein, López de Prado and Zhu developed the Probability of Backtest Overfitting and the Deflated Sharpe Ratio precisely because analysts "backtest millions (if not billions) of alternative strategies". In Notices of the AMS in 2014 the same authors wrote that they "suspect that a large proportion of backtests published in academic journals may be misleading". The mechanism is selection, not fraud: given enough trials, an impressive Sharpe ratio arises by construction, and reporting the winner without correcting for the number of attempts is a statistical error rather than a marketing choice. Detailed treatment on the overfitting page.
Where the research is positive
The literature is not uniformly negative. Finance-specific language-model work, including FinLlama and LLaMA-based sentiment and return-prediction studies, demonstrates that models can extract usable signal from text. Sirignano and Cont reported in Quantitative Finance in 2019 that deep LSTM models achieved roughly 65–75% directional accuracy against 57–67% for linear models on a high-frequency event task across 500 NASDAQ stocks — a genuine and substantial improvement on a well-specified problem. What none of this establishes is durable net live returns after costs, which is a different claim requiring different evidence.
| Study | Task | Reported accuracy | Status |
|---|---|---|---|
| Yu and Yan | S&P 500 next-day direction | 58.07% | Published study |
| Ding et al. | Equity direction | 64.21% | Published study |
| ScienceDirect deep-learning and econometrics study, 2025 | Transformer architecture, directional | approximately 69.1% | Peer-reviewed journal |
| Sirignano and Cont, Quantitative Finance, 2019 | High-frequency event task, 500 NASDAQ stocks | approximately 65–75% (deep LSTM) vs 57–67% (linear VAR) | Peer-reviewed journal |
| Methods review, arXiv 2407.09831 | Critique of studies reporting 80–90%+ | Attributes such figures to look-ahead bias and shuffled splits | Preprint, not peer-reviewed |
Prediction accuracy: the realistic range
Careful directional-accuracy figures cluster between roughly 55% and 65%. Published anchors include Yu and Yan at 58.07% for next-day S&P 500 direction, Ding et al. at 64.21%, a 2025 deep-learning and econometrics study reporting approximately 69.1% for a transformer architecture, and the Sirignano and Cont high-frequency range above.
A 2024 methods review states that studies claiming "80–90% or even 90+% accuracies… may suffer from lookahead bias… especially when shuffling the datasets before splitting, which will lead to artificially inflated accuracy results". That source is a working paper rather than a peer-reviewed article and is attributed as such — but the mechanism it describes is uncontroversial and reproducible: shuffle a time series before splitting it and the model sees the future.
The practical point is that a 58% directional model is not a licence to trade. Whether it makes money depends on payoff asymmetry, turnover, capacity and cost per trade — a 58% win rate with negative expectancy loses money reliably. Detail on the accuracy page.
Why no retail ranking can be trusted
No independent, cost-adjusted, like-for-like live return series was found for any set of retail AI trading vendors. Vendors publish subscription tiers, feature quotas and user counts; they do not publish audited net performance on comparable capital over comparable periods with disclosed methodology. Where user numbers appear at all — a platform describing itself as used by 100 million traders, for instance — they are unaudited vendor claims rather than independent active-user counts.
This absence has a direct consequence: any article ranking "the best AI trading bots by returns" is constructed from something other than measured returns. Usually it is a blend of vendor marketing, backtested figures supplied by the vendor, and affiliate economics. We do not publish such a ranking. The platform page compares what can be verified — price, quotas, asset coverage, control, custody model, regional availability — and states plainly that it contains no performance ranking.
Where AI does measurably work
The evidence is much stronger for applications where the objective is measurable and the counterfactual is clear. These are rarely the applications sold to retail traders.
| Application | Evidence strength | Why |
|---|---|---|
| Execution cost reduction (smart order routing, VWAP/TWAP scheduling) | Strong | Implementation shortfall is directly measurable against a benchmark price; improvement is attributable |
| Fraud detection and AML | Strong | Labelled outcomes exist; cited by 33% of respondents in the Bank of England and FCA survey |
| Cybersecurity | Strong | Named by 37% of respondents; detection performance is measurable |
| Market surveillance | Moderate to strong | Supervisory literature documents benefits; ground truth is partially observable |
| Internal process and research throughput | Moderate | Named by 41% of respondents; benefit is time saved rather than return generated |
| Text and sentiment feature extraction | Moderate | Research demonstrates usable signal; net live return remains unestablished |
| Directional price prediction | Weak | Accuracy marginally above chance; costs consume most of the gross edge |
| Autonomous retail profit generation | Not supported | No audited evidence; published fund record below passive benchmark |
The pattern is consistent. AI performs where there is a stable objective, plentiful labelled data and a measurable counterfactual. Directional price prediction has none of those properties: the target is close to a martingale, the data are non-stationary, and the counterfactual moves as soon as others trade the same signal.
The behavioural claim, examined
Automation genuinely does apply a rule without discretionary hesitation, and that is not a trivial benefit: it removes panic exits, revenge trading and the fatigue that degrades discretionary decisions late in a session. But three qualifications matter.
First, it executes a bad or stale rule with identical consistency. Second, it does not remove the human behaviours that surround the system — the operator still switches it off during a drawdown, retunes it after a losing week, and adds leverage after a winning one. Third, and most importantly, it does not reduce leverage, which is the dominant driver of documented retail losses. The Bank of England has separately warned that wider AI trading could produce correlated positions that amplify shocks, which is a behavioural risk introduced by automation rather than removed by it.
The loss statistics, in context
The most reliable retail outcome data in Australia, New Zealand and the United Kingdom concern contracts for difference. ASIC REP 828, published 20 January 2026, found 68% of Australian retail CFD investors lost money in FY2024. The FCA states that "approximately 80% of customers lose money when investing in CFDs", with its CP16/40 account sample finding 82%. Individual firm disclosures published under the quarterly recalculation requirement show similar magnitudes: IG reports 70% of retail accounts losing on spread bets and CFDs, and Interactive Brokers UK has been reported at 69%.
These figures are about leveraged CFDs, not about AI tools, and it would be dishonest to present them as a measure of AI performance. They are included because they establish the base rate in the product category where most retail AI-branded automation is actually marketed and because they quantify what "the dominant risk is leverage" means numerically. Australian intervention data also show the counterfactual: after the 2021 product intervention order, aggregate retail CFD net losses fell 91% and the average number of loss-making accounts per quarter fell 51%. Detail on the loss statistics page.
The verdict, stated precisely
Four statements are supportable on the evidence reviewed here.
- Professional AI funds have not beaten a passive equity benchmark over the fifteen-year window for which index figures were published, nor in the one documented stress month.
- No audited evidence supports retail AI bot performance claims, and the structural reason is that no independent like-for-like live return series exists.
- AI does add measurable value in execution, surveillance, fraud detection and research throughput — roles with measurable objectives.
- The binding constraint for retail participants is cost, leverage and fraud, not model quality. Improving the model does not relax that constraint.
What would change this assessment: a multi-year, independently audited, cost-inclusive live return series for a defined set of AI strategies, benchmarked against a low-cost index and disclosing the number of strategies attempted. That is a high bar, but it is exactly the bar the claim requires. Until it is met, the rational default for a retail participant is the index fund — and the burden of proof sits with anyone selling the alternative.
Frequently asked questions
Does AI trading work?
Not as a route to reliable excess returns on the independent record: the published AI fund index trails the S&P 500 over fifteen years and peer-reviewed review found no conclusive evidence of machine-learning funds delivering returns at scale. It does work in execution cost reduction, surveillance and fraud detection.
Do AI hedge funds beat the S&P 500?
Historically no: 9.8% versus 13.7% annualised between December 2009 and July 2024, and 115% versus 210% cumulative between January 2011 and January 2020.
Which is the best-performing AI trading bot?
Unanswerable from evidence. No audited, cost-adjusted, like-for-like live return series exists for retail vendors, so any ranking by performance is built from vendor claims or backtests.
Why do backtests beat live results so consistently?
Selection bias across many tested variants, look-ahead bias, survivorship bias and omitted costs. The Deflated Sharpe Ratio and the Probability of Backtest Overfitting exist to correct the first.
Does automation improve returns by removing emotion?
It removes impulsive overrides, which is real, but it applies a bad rule just as consistently as a good one, does not remove operator behaviour around the system, and does not reduce leverage.
What proportion of retail traders lose money?
For CFDs specifically: 68% in Australia in FY2024 per ASIC REP 828, and approximately 80% in the United Kingdom per the FCA, with an 82% figure in its CP16/40 sample.
How accurate can AI price prediction get?
Roughly 55–65% directional in careful work, with up to approximately 65–75% reported on a narrow high-frequency task. Above 90% almost always indicates look-ahead bias.
Can language models generate usable trading signals from news?
They can extract signal — finance-specific LLM research demonstrates this — but no study establishes durable net live returns, and sentiment remains one noisy feature rather than a system.
Did AI funds hold up better in a downturn?
Not in January 2022, when the AI index fell 3.38%, its worst month since inception, against 1.11% for the S&P 500.
What return should I expect from an AI strategy?
No credible universal figure exists, and a specific promised return is itself a warning sign. Benchmark against a low-cost index fund out of sample and after all costs.
Is this evidence against AI, or against retail AI?
Against the claim of easy excess returns at every level. The index evidence concerns professionals and is negative relative to passive; retail participants face the same statistics plus higher costs, leverage and fraud exposure.
Does any of this mean AI trading is a scam?
No. The techniques are real and legitimately used by regulated institutions. What is fraudulent is the specific marketing pattern of guaranteed or fixed high returns from AI automation — a pattern regulators are actively removing at scale.
About this page
Compiled by the AI Trading Book editorial desk. Every performance figure here is reproduced with its measurement window, benchmark and reporting source; where a figure reaches us through a secondary report of an index, we say so rather than implying direct access to the index. Caveats that weaken our own conclusion are published alongside it. Published 27 August 2026 · Updated 28 August 2026. Next scheduled review: November 2026. Editorial policy · Report an error.
Sources
- Eurekahedge AI Hedge Fund Index, as reported by IG, "Is the Impact of AI on Hedge Funds Overhyped?" — 21 November 2024 — 9.8% versus 13.7% annualised, December 2009 to July 2024.
- Eurekahedge, as reported by Fortune — January 2022 monthly return of −3.38%, worst month since 2010 inception, against −1.11% for the S&P 500.
- Buczynski, Cuzzolin and Sahakian, International Journal of Data Science and Analytics 11(3) — April 2021 — review of 27 machine-learning equity experiments; cumulative 2011–2020 series as summarised by Alpha Architect, 18 October 2024.
- Harvey, Liu and Zhu, "…and the Cross-Section of Expected Returns", Review of Financial Studies 29(1) — January 2016 — t-statistic above 3.0.
- Bailey and López de Prado, "The Deflated Sharpe Ratio" — 2014; Bailey, Borwein, López de Prado and Zhu, "Pseudo-Mathematics and Financial Charlatanism", Notices of the AMS — 2014; "The Probability of Backtest Overfitting" — revised 2015.
- Sirignano and Cont, Quantitative Finance — 2019 — deep LSTM directional accuracy on 500 NASDAQ stocks.
- Yu and Yan; Ding et al.; ScienceDirect deep-learning and econometrics study, 2025 — directional accuracy anchors.
- arXiv 2407.09831 (preprint, attributed as such) — 2024 — look-ahead bias and shuffled splits inflating reported accuracy.
- FinLlama, Proceedings of the 5th ACM International Conference on AI in Finance; and finance-specific large language model sentiment and return-prediction research, ScienceDirect.
- ASIC, REP 828 — 20 January 2026 — 68% of retail CFD investors lost money in FY2024.
- ASIC, REP 724 — aggregate retail CFD net-loss reduction of 91% and 51% fewer loss-making accounts per quarter following the product intervention order.
- FCA — "approximately 80% of customers lose money when investing in CFDs"; CP16/40 sample 82%; firm-level disclosures under COBS 22.5.
- Bank of England and FCA, "Artificial intelligence in UK financial services – 2024" — 21 November 2024, n=118 — use-case mix including cybersecurity 37%, fraud detection 33%, internal process optimisation 41%.
- Bank of England, "Financial Stability in Focus: Artificial intelligence in the financial system" — 9 April 2025 — correlated positions amplifying shocks.
Informational research only. Nothing on this page is personal financial, legal, tax or investment advice, or a recommendation to trade any instrument. Past performance, whether of an index, a fund or a strategy, does not indicate future results.