In 1988, The Wall Street Journal began throwing darts at the stock pages on purpose. Burton Malkiel had proposed the test fifteen years earlier in A Random Walk Down Wall Street: “a blindfolded monkey throwing darts at a newspaper’s financial pages,” he wrote, “could select a portfolio that would do just as well as one carefully selected by experts.”[1] The Journal swapped the monkey for staff members and ran the contest for fourteen years. Professional stock pickers against random darts, round after round, in public.
The pros won more rounds than they lost. Sixty-one of the first hundred, and on the raw averages they looked comfortably ahead, 10.8 percent a round against the darts’ 4.5.[2] That deserves saying plainly, because the honest numbers are more interesting than the legend.
Malkiel pointed out the two things the raw numbers hide: the pros picked riskier stocks, and their picks got a bounce just from being printed in the Journal. Adjust for both, he argued, and the advantage effectively disappears. Against literal randomness, some of the best-paid stock pickers in the world produced an edge that a risk adjustment could erase.
The scoreboard has only grown since. The SPIVA scorecard, which has compared active funds to their index for more than two decades, shows roughly nine in ten actively managed large-cap funds trailing the plain S&P 500 over fifteen years.[3]
And the belief in the pickers did not die. Chart readers still find heads and shoulders in the price history. Forecast shops still sell models. Enormous sums still move on the conviction that someone can see the pattern.
Nassim Taleb wrote Fooled by Randomness a quarter century ago to explain how that survives. Run enough coin flippers and a few will flip heads ten times straight. Those few will sincerely believe, and sincerely sell, their technique. The fees collect on assets, not accuracy. The backtest can always be revised. And the sellers were not the ones bearing the losses.[4]
I bring up fifty years of market history because the same belief is reassembling around AI, and this time there is no scoreboard at all.
The chartist and the prompt-reader
A large language model is a prediction machine. It predicts language: given everything so far, the most plausible next stretch of text. That machinery turns out to be spectacularly useful, and I use it daily.
But watch what happens in the reading. The model produces a fluent strategy memo, and the reader sees insight: analysis, judgment, a mind at work. The document reads as if someone reasoned their way to it, because it was assembled from millions of documents written by people who did.
The chartist looks at a random walk and sees heads and shoulders. The prompt-reader looks at a statistically assembled stream of text and sees strategy. It is the same act. The pattern detector fires, and the insight is contributed by the reader.
I published my own caution about this last year: “AI is mean-regressing, trend-following, and reflective of the common point of view.”[5] I chose those words deliberately. Trend-following is what technical analysis is. The model hands you the consensus of everything it has read, rendered confidently, and the consensus of everything ever written about strategy sounds a great deal like insight.
There is a law hiding in why unverified output lands so hard. The output impresses you in proportion to the gap between it and what you could have produced yourself. Your ability to check it runs in the opposite direction.
The person most dazzled by an AI market analysis is, by construction, the person least equipped to tell whether it is right. Where you know the domain cold, the model reads like a competent draft. Where you know nothing, it reads like a genius. The genius impression peaks exactly where your ability to verify hits zero.
That is why the low-stakes uses feel magical and stay harmless. Ask for a birthday poem or a first draft of a memo and the output really is better than what most of us would produce. No one is paying for it to be right, and no one gets hurt when it is not.
The trouble starts when the same feeling of being impressed is carried into decisions that someone pays for, in a setting where nobody tracks which recommendation came from where. A wrong number in a fluent report can sail on for years without anyone tracing the failure back to its source.
The question was never whether patterns exist
Now the concession. Markets do contain structure. Momentum exists. Factor premiums exist. The anomaly literature is real, and Malkiel never claimed otherwise. His argument was that the extractable part, out of sample and after costs, is far smaller than the industry’s marketing, and fifty years of fund scoreboards have backed him.
The same concession applies on the other side: the regularities a language model captures are real, which is why it is so useful for language. There are even published backtests claiming that a language model’s reading of the news predicts next-day stock moves. The same study reports its strategy’s returns shrinking as more traders adopt the models, which is exactly what you would expect real structure to do once it is traded.[6]
So the question in both domains is the same, and it is narrower than the sales pitch. Not “is there structure?” but “is it extractable, out of sample, after costs, by you, with accountability for the misses?”
Markets answer that question whether you like it or not. Every trade gets scored, in dollars, on a date. Most AI use never faces anything like it. I wrote recently about the prompt that never appears in the replace-the-consultants genre, the one that checks the work.[7] This is the same missing layer viewed from the other end: no attribution, no closure, and a failure mode that can run for years on charm.
Which suggests the move. Take the AI-insight claim to the one arena where the scoreboard is mandatory.
The dartboard, rebuilt
If language models carry genuine insight about business and markets, there is a clean way to find out, and Malkiel designed the protocol in 1973. Have the model pick stocks. Trade the picks with real money. Publish everything.
One rule matters more than all the others: no backtests. The training corpus contains the history. A model asked to “predict” 2023 from inside a system trained on 2024 may simply be reading the answer key. Careful researchers work around it by testing only on news the model has never seen. The cleaner protocol needs no workaround at all: live picks, forward only, timestamped before the market moves.
The design has to respect how hard this measurement is. Returns are noisy enough that telling skill from luck in a single portfolio takes years, which is Taleb’s whole point. The way to buy statistical power is breadth, not patience: many small, independent, timestamped decisions rather than one heroic account.
Long-short construction, so a rising market cannot flatter every method at once. Parallel sleeves, because “AI” is not one thing: the bare model reasoning from its training, the model with tools and live data, the model as a research assistant under a fixed human risk overlay. And controls on both ends: a plain index fund, and, in Malkiel’s honor, actual darts.
The stakes should be real and can be small, because skin in the game is binary. Either the losses are felt or the exercise collapses back into content. Publish the prompts, the picks, and the equity curves as they happen, not after selection has done its work.
To be clear about what this is: an experiment design, not a strategy. There is no ticker in this piece and no prediction about any price. The bet being tested is not mine on a stock. It is the discourse’s bet that fluency carries insight.
And the scope should be stated as plainly as the design. Markets are the hardest test available, adversarial and picked over by professionals. Scoreboards settle real capability quickly when it exists. Chess engines proved themselves on the board, and the debate closed. The market’s scoreboard has run for fifty years without producing that kind of closure for the pattern-readers. Failing there falsifies the grand claim, the one where the machine sees what you cannot. It does not make the tool useless, any more than index funds made reading annual reports useless. Medicine, law, and corporate strategy have slower, noisier scoreboards. They inherit the diagnosis, not the verdict.
Say it before the data arrives
Here is the part I want on record in advance, because Taleb’s book was never really a prediction about markets. It was a prediction about discourse.
Let thousands of people run AI portfolios through noisy markets and some will post spectacular results, guaranteed, whether or not any model has an edge. The winners will screenshot. The losers will quietly delete. The evidence will assemble itself either way, and all of it will point the same direction.
So even a careful null result will be “refuted” within a week by somebody’s chart going up and to the right. The anecdote is not the unit of evidence here. The distribution is. Every sleeve that started, the whole curve, the denominator.
That rule travels well beyond trading. Any time you are shown an AI success story, in any domain, three questions tell you most of what you need. How many attempts started? Where did the failures go? Who bore the loss?
The index-fund move
Fifty years of mandatory scorekeeping never killed the forecasting business. It produced something more useful than a verdict: an exit. The index fund let anyone use the market without buying the prediction story. Take the structure that is really there, decline the insight premium, keep the difference.
The AI era needs the same product, and you can assemble it yourself. Use the model for what a prediction machine demonstrably does: drafts, summaries, recall, translation between formats, work you can check. Decline the insight premium, the part of the pitch where the machine has supposedly seen something about your market or your strategy that you cannot verify and are asked to act on.
The market spent half a century teaching us how this story goes. Keep score, or keep believing.
Reference Sources
- Malkiel, Burton G. A Random Walk Down Wall Street. W. W. Norton, 1973. Source of the systematic-outperformance argument and the test itself: “a blindfolded monkey throwing darts at a newspaper’s financial pages could select a portfolio that would do just as well as one carefully selected by experts.”
- “The Wall Street Journal Dartboard Contest.” Investor Home, compiled record of the Journal’s contests, October 4, 1988 through 2002. Accessed July 29, 2026. Through the 100th contest (October 1998), professionals won 61 of 100 against the darts, averaging 10.8 percent a round against the darts’ 4.5 and the Dow’s 6.8, and Malkiel’s assessment that the advantage “effectively disappears” adjusted for riskier picks and the announcement effect.
- S&P Dow Jones Indices. “SPIVA U.S. Scorecard, Mid-Year 2025.” Report 1a. Accessed July 29, 2026. 88.29 percent of all large-cap funds underperformed the S&P 500 over the 15 years ending June 2025, and 91.03 percent over 20 years.
- Taleb, Nassim Nicholas. Fooled by Randomness. Texere, 2001. Survivorship and luck mistaken for skill.
- Shivamber, Leon. “What you should know about my AI Use.” LinkedIn, August 2025. Accessed July 25, 2026. “AI is mean-regressing, trend-following, and reflective of the common point of view.”
- Lopez-Lira, Alejandro, and Yuehua Tang. “Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models.” arXiv, April 2023. Accessed July 29, 2026. ChatGPT sentiment scores of news headlines predicted next-day returns in a post-training-cutoff sample, and the authors report strategy returns declining as model adoption rises.
- Shivamber, Leon. “The Missing Prompt.” shivamber.com, August 2026. The verification layer missing from the replace-the-consultants genre.