The Frontier AEO Tracker: What Astra Chooses (and every other frontier model, and what you can do about it) Latent Space released the Frontier AEO tracker at aeo.latent.space, which evaluates how 7 frontier AI models answer 6 prompt variations across 161 categories, with answer extraction performed by Astra and scored using a proprietary AEO score that weights first choices, alternatives, mentions, and negative anti-recommendations. The tracker, inspired by AmplifyingAI's 'What Claude Code Actually Chooses,' also lists top cited sources and an analysis of top failures, and the article notes that naive autoresearch investment in AEO has yielded impressive ROI. Naive autoresearch investment https://x.com/swyx/status/2078244735794413786 in our AEO have yielded impressive ROI, and so naturally it was time to take it seriously. We were inspired by What Claude Code Actually Chooses https://amplifying.ai/research/claude-code-picks , and decided to extend/adjust it to our tastes. After a https://www.latent.space/p/astra few billion tokens https://www.latent.space/p/astra of prototyping, aligning, and scaling https://www.latent.space/p/astra pipelines, here’s the Latent Space Frontier AEO tracker https://aeo.latent.space/ . Our methodology https://aeo.latent.space/methodology extends AmplifyingAI’s https://amplifying.ai/research/claude-code-picks/report to run 6 prompt variations over 7 models https://aeo.latent.space/models 1 footnote-1 search on in 161 categories https://aeo.latent.space/categories , from coding agents https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=coding-interactive to AI podcasts https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=ai-podcasts to AI Sandboxes https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=ai-sandbox to Managed Databases https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=managed-databases to ASR models https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=model-transcription to even oddball categories like Angel investors https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=angel-investors and Corporate spend https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=corporate-spend and Payroll software https://aeo.latent.space/categories?marketCategory=coding-delegated&productGroup=overview&highlights=%5B%22entity%3Ac44a039a-732e-510b-b48f-3993c799bf73%22%2C%22entity%3A88663828-da00-4703-b087-16e6404e8ba7%22%5D&cat=startup-payroll . Answer extraction was done by Astra, and scored for a proprietary AEO score that gives weight to first choices, alternative choices , mentions , but also negative weights to mild and strong anti-recommendations which are rare, but do happen . Because we know you’ll want it, we also extracted the top cited sources https://aeo.latent.space/sources/?population=matched&measure=conditional&topic=all&question=all&source=all&search=&include-grok=false&sort=mean&direction=desc which influence Agent recommendations, as well as an analysis of top failures https://aeo.latent.space/sources/?population=matched&measure=conditional&topic=all&question=all&source=all&search=&include-grok=false&sort=mean&direction=desc failure-report . Basic Results Here are the most dominant products in their categories in the world: There are some familiar names in there — opening up the natural question of contamination, which we have checked. Since we have nothing to hide, every prompt and answer pair https://aeo.latent.space/categories?cat=ai-newsletters is inspectable. However, bias does exist - when models are asked for coding agent recommendations, Fable/Opus like Claude Code https://aeo.latent.space/entities?entity=Claude+Code&entityId=platform%3Aa625159d18da57180b7215e5&cat=coding-interactive and Sol/Astra like Codex https://aeo.latent.space/entities?entity=OpenAI+Codex&entityId=entity%3A2fa20db7-646b-4db3-ba91-1937e00416ba&cat=coding-interactive and Grok loves Cursor https://aeo.latent.space/entities?entity=Cursor&entityId=platform%3A46a4eebd20d881ecfc0ecb54&cat=coding-interactive and Muse loves Muse Code https://aeo.latent.space/entities?entity=Muse+Code&entityId=entity%3Afe0616b6-d14b-4fcc-9f72-80229b0d2353&cat=coding-interactive and SWE-1.7 loves Devin https://aeo.latent.space/entities?entity=Devin&entityId=entity%3A88663828-da00-4703-b087-16e6404e8ba7&cat=coding-interactive and so on. I wonder why. You can see other “ soft biases https://aeo.latent.space/entities?q=lovable&entity=Lovable&entityId=entity%3Aeeabf3f2-663b-597a-9f8c-917b8f6a732d&cat=ai-app-builders ” emerge too… That said there are notable examples of GPT models recommending Claude, a laudable nonbias: There are 28 categories https://aeo.latent.space/ agreement-title out of our total 161 which have a universally dominant primary choice - among all surveyed frontier models. There are a lot more “close contests” and “always the vibesmaid, never the vibe” https://aeo.latent.space/home competitive categories which should be key AEO battlegrounds. Sol vs Astra, Opus vs Fable New pretrains for new model classes represents a new opportunity to check in on what the labs are moving towards in their data and RL priorities, and to check in on whether startups’ investments in AEO are paying off. We prepared special reports analyzing our rankings, observing VERY consequential flips in model choices https://aeo.latent.space/ disagreement between model generations from the same lab. We have separate Opus→Fable https://aeo.latent.space/models?mode=compare&from=claude-opus-5&to=claude-fable-5-1 and Sol→Astra https://aeo.latent.space/models?mode=compare&from=gpt-5.6-sol&to=gpt-6-astra summary pages. For some flips, we highlighted a neutral analysis of what competitors did better in each scenario. Efficiency vs Confidence, and Recommendation Sourcing One of our most surprising findings between Sol→Astra and Opus→Fable is that Anthropic seems to be biasing their models to searching more sources Sol median of 9 sources, vs Astra median of 5, vs Opus median of 11 sources, vs Fable of 15 . Astra seems to be just generally a lot more “ confident ”, or “ efficient ”, depending how you look at it - Astra is FAR less likely to change its mind when you lightly paraphrase your question. This makes the value of AEO itself rise as choice randomness declines. Sources analysis https://aeo.latent.space/sources/ also somewhat strongly predicts what the labs do prioritize vs don’t. However the sample size is small here and only represents what we can scrape from attempted toolcalls, not the pretrain dataset. What we CAN validate is that AEO practices measured by Ora and Vercel https://is-agentic.com/ , like markdown content-negotiation , are real and failures discourage models from reading your content. Just for fun Here are the top Angels https://aeo.latent.space/models?mode=compare&from=claude-opus-5&to=claude-fable-5-1&topic=angel-investors in the world according to LLMs some dedupes left to do… . See more We also made a little family feud type game https://aeo.latent.space/play where you can see if your priors align with the data. Fun We are open to further suggestions and business enquiries to develop this if it is of interest. Ping @latentspacepod http://x.com/latentspacepod/ or email \ email protected\ http://business@latent.space we have a business manager now woo 1 footnote-anchor-1 As we note in our methodology post, we did try VERY hard to include Gemini/Antigravity, GLM/Zcode, and DeepSeek/DeepCode, but errors and rate limits made them untenable to include in this first run analysis. Please let us know how to raise limits if you represent these companies.