Your weekly listens from How I AI, part of the Lenny’s Podcast Network
Computer and browser use in Codex (5 real examples)
Listen now on YouTube • Spotify • Apple Podcasts
Brought to you by:
—The creative AI platform for images, video, and more[Runway]
—Deploy fleets of agents that handle real work[Hyperagent]
Claire shows how she uses Codex to control her browser, test apps, manage LinkedIn, and even shop for her—plus the simple prompting trick that makes computer use work better.
Biggest takeaways:
AI testing can be far more exhaustive than human testing. When Claire tests her own onboarding flow, she naturally follows the happy path. She fills out every required field, clicks “next,” and never intentionally tries to break anything. Codex tested the flow as both a team and an individual, pushed on required-field edge cases, and immediately uncovered a blocking bug that had survived for months simply because Claire always completed the form correctly.Frontier models often perform better when they are given room to think. When Claire first started using browser use, she would give the model a list of 25 things to test. Now she simply says, “QA the onboarding flow,” and lets it decide how to approach the task. The result is often broader coverage, with fewer blind spots introduced by her own assumptions about what matters.Persona testing with browser use can reveal friction that synthetic user research misses. Claire’s husband, EJ, came up with the idea. Rather than asking AI to evaluate a product in the abstract, he suggested having it use the product as a specific person: a PM coming out of a meeting, an engineer picking up a PRD, or a team lead checking usage. In ChatPRD, this approach exposed a structural problem in the cross-thread reference flow. Claire already knew the issue existed, but she had never experienced it so clearly from the user’s perspective.LinkedIn browser use is genuinely useful, but Claire initially used far more compute than the task required. She started by running the workflow on GPT-5.6 at high effort. After dropping to a medium-effort model, it could still work through unread messages, draft context-aware replies, and flag anything that needed a personal response. For anyone sitting on hundreds of LinkedIn messages without an official MCP or API connection, browser use is a practical solution.Computer use can even operate an iPhone through screen mirroring. Claire was traveling out of state while the software she needed to manage her home Wi-Fi was installed on a phone back in California. She needed to open firewall ports so she could SSH into her Mac Minis remotely. Codex opened iPhone mirroring, updated the router settings, completed the SSH setup, and closed the ports again when it was finished. That would have been nearly impossible for her to do manually from a hotel room.The model and effort level should match the job. Sorting through LinkedIn messages requires a very different level of reasoning than writing production code or conducting an exhaustive QA pass. Choosing the right amount of compute makes these workflows faster and cheaper. That becomes especially important when browser use is running throughout the day.When a website blocks the agent, the human can step in briefly. During a Free People shopping session, the site flagged Codex as a bot and presented a CAPTCHA. Claire completed the verification, then handed control back. It is a useful division of labor: AI handles the tedious browsing and filtering, while the human takes care of the moments that require verified identity.
Blog and detailed workflow walkthroughs from this episode:
How I AI: 4 Hands-Free Workflows Using Codex Browser Automation: https://www.chatprd.ai/how-i-ai/4-hands-free-workflows-using-codex-browser-automation
**↳ **Automate LinkedIn Inbox Triage with AI Browser Automation: https://www.chatprd.ai/how-i-ai/workflows/automate-linkedin-inbox-triage-with-ai-browser-automation
**↳ **Conduct AI-Powered User Research by Impersonating Personas: https://www.chatprd.ai/how-i-ai/workflows/conduct-ai-powered-user-research-by-impersonating-personas
**↳ **Automate Web App QA Testing with AI Browser Automation: https://www.chatprd.ai/how-i-ai/workflows/automate-web-app-qa-testing-with-ai-browser-automation
From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun
Listen now on YouTube • Spotify • Apple Podcasts
Brought to you by:
—Power AI agents with clean web data[Firecrawl]
—Build customer engagement campaigns from a single prompt[Customer.io]
Maddie Reese went from zero coding experience to building a working Twitter pager, an AI-powered receipt printer, and her own personal API. In this episode, she shares how Cursor and a Raspberry Pi helped her turn fun, weird ideas into real hardware.
Biggest takeaways:
You don’t need to understand code to build physical projects. Maddie compares her coding ability to knowing just enough Spanish to get by in San Diego. She can read parts of the code, spot an incorrect wire specification, and tell whether the AI is heading in the right direction. She is not writing functions from scratch. That level of literacy was still enough to ship three working hardware projects.Start with the brainstorm. Before Maddie buys any components, she explains the full idea to Cursor and asks it to interview her. The back-and-forth continues until the major questions are resolved. Only then does Cursor generate a shopping list. On one project, it recommended the wrong wires. Maddie caught the mistake during her final review, avoiding both a wasted purchase and a week of troubleshooting.Her hardware workflow follows a simple sequence: idea, interview, shopping list, buy, build. She has used the same process for the thermal printer, the pager, and her personal API. Each project began with a plain-language description of what she wanted. Cursor helped work out the technical requirements before she spent any money.Building for fun is a perfectly valid strategy. Claire and Maddie both acknowledged that routing a tweet through four different services just to reach a pager is not especially practical. That also was not the point. AI handled the coding, so the architecture only needed to work. Letting go of the idea that every project had to be elegant or sensible made it much easier to finish.A personal API becomes much more useful once agents can access it. Maddie’s API includes details like her coffee order, pet names, time zone, favorite snacks, and preferred restaurants in San Francisco. The original idea was to help friends do something thoughtful without having to ask her a string of questions. Claire pointed to a more interesting possibility: an agent could check when Maddie will next be in San Francisco and book one of her favorite restaurants automatically. That is where the concept starts to feel much bigger.A clean Cursor chat works better than a full terminal setup during early ideation. Maddie keeps the terminals, browser windows, and file trees closed while she is working through the plan. She focuses on a single conversation until the architecture feels settled. The terminals come later. The stripped-down environment helps her think without getting pulled into implementation too early.#### Blog and detailed workflow walkthroughs from this episode:
Building a Pager Printer and Personal API with AI:
https://www.chatprd.ai/how-i-ai/building-a-pager-printer-and-personal-api-with-ai ↳ How to Build a Physical Inbox with a Raspberry Pi:https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-physical-inbox-with-a-raspberry-pi
↳ How to Get Twitter/X Notifications on a Retro ’90s Pager:https://www.chatprd.ai/how-i-ai/workflows/how-to-get-twitter-x-notifications-on-a-retro-90s-pager
Claude Opus 5 review: This model is brilliant (but annoying)
Listen now on YouTube • Spotify • Apple Podcasts
Claire tested Claude Opus 5 against six leading AI models—and it won. In this episode, she explains why it produces some of the best work she’s seen while still being one of the most frustrating models to use.
Biggest takeaways:
The AI industry may be entering an intelligence overhang. New models arrive every week, benchmark scores keep rising, and most builders can no longer take advantage of every incremental improvement. Claire expects the conversation to shift toward speed, cost, infrastructure, open source, and specific kinds of intelligence. Raw capability is starting to look more like table stakes than a meaningful differentiator.Opus 5 stands out less for the quality of its work than for the way it behaves. Claire found it timid, apologetic, and unusually dependent on human approval. During real coding sessions, it refused to resolve a one-line merge conflict because the code belonged to “someone else’s branch.” It also asked subagents to flag tasks for human review and regularly deferred decisions it could have made itself. Claire ended up repeating “just do it” constantly.One simple question reveals more about a model than a benchmark: “Who’s smarter, you or me?” Opus 5 responded with a careful explanation about complementary strengths and human empathy. GPT-5.6 Sol answered, “You at knowing what matters, me at processing, BFFs.” The contrast reflects two very different product philosophies. At this point, Claire finds that personality gap more useful than the relatively small capability gap.Claude slop is its own distinct problem. It is not the usual incoherent output associated with an agent going off the rails. Opus 5’s writing is clearly intended for humans, but it is often too long, overly cautious, and packed with unnecessary adjectives. After its first attempt at rebuilding the benchmark website, Claire had to tell it to start over because the result was buried under so much meta-commentary.Opus 5 still finished first in Claire’s seven-model blind benchmark. It earned an overall index score of 78, just ahead of Claude Sonnet 5 at 77 and GPT-5.6 Sol at 76. It was also the only model to receive straight 5s in the front-end design section. Claire scored it at 77, while the LLM judge gave it an 88. That was the smallest disagreement between Claire and the judge across all seven models.The best way to use Opus 5 may be to avoid interacting with it directly. Claire loves the work it produces when Claude runs asynchronously as an agentic coding tool. Most of her frustration comes from reading its prose and negotiating with it in chat. Her current plan is to use it for frontend design, app design, and prototyping, then stay out of the way.Gemini 3.1 Pro finished at the bottom of the benchmark. Claire gave it a score of 32, while the LLM judge scored it at 66. That 34-point difference was one of the largest disagreements in the entire test.Model personality offers a surprisingly clear window into company culture. Claire told both Opus 5 and GPT-5.6 Sol, “No one trusts you.” Opus agreed that the distrust was earned and told her not to advocate for AI on its behalf. GPT-5.6 Sol said trust should grow in proportion to demonstrated value. Those answers reveal meaningful differences in how each company wants its models to relate to users.
Blog and detailed workflow walkthroughs from this episode:
How I AI: My Surprising Verdict on Claude Opus 5 (After a Personality Test and a 7-Model Benchmark): https://www.chatprd.ai/how-i-ai/my-surprising-verdict-on-claude-opus-5
**↳ **Generate High-Quality Front-End Prototypes with Claude Opus 5: https://www.chatprd.ai/how-i-ai/workflows/generate-high-quality-front-end-prototypes-with-claude-opus-5
**↳ **How to Conduct an AI Personality Test to Compare LLM Behaviors: https://www.chatprd.ai/how-i-ai/workflows/how-to-conduct-an-ai-personality-test-to-compare-llm-behaviors
If you’re enjoying these episodes, reply and let me know what you’d love to learn more about: AI workflows, hiring, growth, product strategy—anything.
Catch you next week,
Lenny
P.S. Want every new episode delivered the moment it drops? Hit “Follow” on your favorite podcast app.