Frontier AI models outperform analysts on earnings prediction Samaya reported that the latest frontier AI models running on its finance-specific prediction harness outperformed human expert analysts at predicting earnings surprises, the first time it has observed models beat experts on the task. Samaya built the harness to supply models with point-in-time financial information at parity with human experts and asked them to predict revenue, gross margin, operating income and adjusted EPS one week before earnings for companies in a global stock market worth more than $150 trillion. The company evaluates models on prediction error and surprise correlation, using the consensus estimate of analyst predictions as the expert baseline. Finance is arguably the largest and hardest area of knowledge work. Accurate predictions of the global financial market drive trillions of dollars of value, and take the best experts many years to hone, and even then imperfectly. A couple of months back, we created a research effort at Samaya to study AI’s predictive capabilities in finance. Making accurate financial predictions requires access to a large set of high quality, real-time financial information — finance’s “open-world” equivalent of a codebase. We built a finance-specific prediction harness and environment to give AI models comprehensive, point-in-time financial information at parity with human experts so that we could push their capabilities to the limit. Our results show that we have reached a critical inflection point. For the first time, we see that the latest frontier AI models working with Samaya’s finance harness outperform human experts in financial prediction. Specifically, we find that the latest AI models outperform expert analysts in predicting earnings surprises. Earnings predictions and surprises Global stock markets are worth more than $150 trillion and are the most closely watched asset class in the world, for professionals and individual investors alike. More than ten thousand public companies make up the investable universe across global markets, and most of them report earnings results every quarter, sharing metrics such as revenue , gross margin , operating income and adjusted EPS , also referred to as actuals . These metrics are the foundation for investment decisions into these companies and so an enormous amount of analyst time is spent on modelling, predicting and publishing these metrics ahead of earnings. The average of these predictions is called the consensus estimate . Consensus estimates form a market baseline for the expectation of a company’s performance. When the company reports, the actual is either a “beat” above consensus or a “miss” below consensus with the gap being the earnings surprise . Because predicting earnings is extremely challenging, and even the best consensus estimates miss, the market can react strongly to earnings surprises. So consensus estimates provide a strong “feasible” expert baseline to evaluate AI’s ability to predict earnings and earnings surprises. Building the environment and harness Besides being a very important task for investing, earnings prediction is also a great task for AI. Not only do we have ground-truth actuals and a consensus human baseline, we also have surprise drivers revealed by the company management which can help us understand if the models’ reasoning process was correct. The earnings prediction task advances financial reasoning: it tests the ability to understand company fundamentals, do deep search and retrieval on competitors, supply chain and macro factors, identify key drivers, make the right assumptions and account for them appropriately. 1. Task environment In the earnings prediction task, we run the models under Samaya’s harness and make predictions one week before earnings. Through our harness, we provide access to all financial sources available until that point in time. We ask the models to predict four headline metrics: revenue, gross margin, operating income and adjusted EPS. These metrics track the flow of money through the income statement and capture essential aspects important for financial analysis described in Appendix B appendix-b . We call each such prediction task, predicting all four metrics for one company ahead of one earnings release, an instance . To compare the AI models, we calculate three performance metrics precise definitions in Appendix C appendix-c : - Prediction error measures how far the predicted numbers are from the actuals. - Surprise correlation measures if the surprise i.e. actual − consensus is correlated with the predicted surprise prediction − consensus . This is an overall metric that measures the ability of models to predict big beats and misses correctly. - Hit rate measures if the prediction and actual are on the same side of consensus. Normalizing for volatility: Because some companies’ financials vary more than others, we normalize by each company’s historical surprise volatility to make the predictions comparable across companies. Building a harder expert baseline: We found the consensus baseline relatively easy for AI models to outperform. This is because analysts systematically lower their estimates ahead of earnings,