No title In a blind evaluation on the How I AI podcast, host Claire Vo tested Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol, and other newly launched models across emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing, scoring outputs without knowing which model produced them. Vo said Astra won her heart while Opus 5.5 won her week, with Sol still dividing her on clear writing, readable PRDs, and price, and an LLM judge disagreed with her rankings. The Barbie Bench 3D fashion-game test showed the models remain far from AGI, Vo said, adding that "the hands are tragic. Listen or watch on YouTube https://youtu.be/LMT-bknLmNo , Spotify https://open.spotify.com/show/4aRP2XSavdtrLG5FZoonOK , or Apple Podcasts https://podcasts.apple.com/us/podcast/how-i-ai/id1809663079 I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived. What you’ll learn: 1. How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it 2. Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend 3. Where Sol still wins me over on clear writing, readable PRDs, and price 4. The character SVG results that completely overturned my prediction about Anthropic 5. What happened when I asked these models to edit video, and why I think skills explain part of the disappointment 6. Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t In this episode, we cover: 00:00 https://www.youtube.com/watch?v=LMT-bknLmNo&pp=0gcJCWMAwfN6Pr3D LIVE setup and new model launches 01:30 https://www.youtube.com/watch?v=LMT-bknLmNo&t=90s What’s new in Opus 5.5, Sol, and Luna 04:11 https://www.youtube.com/watch?v=LMT-bknLmNo&t=251s Guardrails, personality, and speed 09:00 https://www.youtube.com/watch?v=LMT-bknLmNo&t=540s The How I AI bench and blind evaluation process 11:31 https://www.youtube.com/watch?v=LMT-bknLmNo&t=691s&pp=0gcJCWMAwfN6Pr3D Email and personal-productivity results 13:50 https://www.youtube.com/watch?v=LMT-bknLmNo&t=830s Frontend prototype vibe checks 24:10 https://www.youtube.com/watch?v=LMT-bknLmNo&t=1450s&pp=0gcJCWMAwfN6Pr3D Backend, agent personality, and long-running tasks 28:25 https://www.youtube.com/watch?v=LMT-bknLmNo&t=1705s SVG illustration test 29:48 https://www.youtube.com/watch?v=LMT-bknLmNo&t=1788s AI video-editing results 30:43 https://www.youtube.com/watch?v=LMT-bknLmNo&t=1843s Predictions before the reveal 31:20 https://www.youtube.com/watch?v=LMT-bknLmNo&t=1880s Barbie Bench: the 3D fashion-game test 34:17 https://www.youtube.com/watch?v=LMT-bknLmNo&t=2057s Results: Astra, Sol, and Opus 5.5 35:04 https://www.youtube.com/watch?v=LMT-bknLmNo&t=2104s Writing clarity and creative surprises 36:51 https://www.youtube.com/watch?v=LMT-bknLmNo&t=2211s Why the LLM judge disagreed with me 37:24 https://www.youtube.com/watch?v=LMT-bknLmNo&t=2244s What each model is actually best for Tools referenced: • Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5 https://www.anthropic.com/claude-opus-5-5 • GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/ https://openai.com/index/introducing-gpt-6-sol-and-luna/ • Codex OpenAI : https://openai.com/codex https://openai.com/codex Where to find Claire Vo: ChatPRD: https://www.chatprd.ai/ https://www.chatprd.ai/ Website: https://clairevo.com/ https://clairevo.com/ LinkedIn: https://www.linkedin.com/in/clairevo/ https://www.linkedin.com/in/clairevo/ Production and marketing by https://penname.co/ https://penname.co/ . For inquiries about sponsoring the podcast, email \ email protected\ https://www.lennysnewsletter.com/cdn-cgi/l/email-protection .