LLMs can't trade & higher reasoning doesn't help. we ran SOTA models for a 2y period. TL;DR: they suck & reasoning doesn't help.
-
no model comes close to simple static baseline
-
more reasoning ≠ better trading
-
when losing money Sol trades less instead of better details 👇
-
leaderboard on avg Sharpe and profit (10 rollouts) is very misleading bec Gemini is supposedly best. but looking at detailed analysis later, it's actually just gemini performance being ultra high variance (v bad). any another 10 runs would give completely different ranking.this is why using markets as proxies to test models is so cool, trading requires some (transferable) skills and we can reason about it from first principles to understand model behaviour and weaknesses. for more detailed write-up, check out our blog:
-
did you also check for long term investing vs trading? say 1-5 year time period. do you expect all retail investors to use LLMs for investing recommendations. it must already be happening in pockets?