Advanced evals: How to find (and fix) hidden AI failures in your product Hamel Husain and Shreya Shankar, drawing on work with over 50 AI companies, published an advanced follow-up on building eval systems for AI products, arguing that most teams skip the critical step before writing metrics and end up measuring the wrong things. The post cites Shopify's eval-guided AI workflow builder running 2.2 times faster and 68% cheaper than the frontier-model system it replaced, Cursor cutting costs 41% with its Auto Balance routing evals, Ramp lifting automatic receipt collection accuracy from 35% to 83%, and Harvey nearly doubling its AI contract reviewer's internal quality score. Nearly half of the 25 PM job openings Lenny Rachitsky shared last week ask for eval-writing experience, and Anthropic Labs head Mike Krieger is quoted saying "writing evals is now probably the most important thing" to teach product people. 👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs https://www.lennysjobs.com/ | Lennybot https://www.lennybot.com/ | Become an AI-Native Builder https://maven.com/tech-for-product/become-an-ai-native-builder and other favorite AI/PM courses https://maven.com/lenny P.S. Get a full free year of Cursor, Notion, Lovable, Replit, Wispr Flow, Linear, Factory, ElevenLabs, PostHog, Granola, Brain.fm, Waking Up, and more, by becoming an Insider subscriber while supplies last . Learn more https://www.lennysproductpass.com/ . Evals have been coming up more and more in my conversations with podcast guests and PMs. And nearly half of the 25 awesome PM https://x.com/lennysan/status/2100657334771237106?s=20 job openings I shared on socials last week ask for experience writing evals. This skill is only becoming more valuable. So I asked Hamel https://www.linkedin.com/in/hamelhusain/ and Shreya https://www.linkedin.com/in/shrshnk/ to write an advanced sequel to their very popular “Building eval systems that improve your AI product” https://www.lennysnewsletter.com/p/building-eval-systems-that-improve post from last year. Drawing from their work with over 50 AI companies, they’ve noticed that most teams jump straight to writing metrics—and end up measuring the wrong things. Below, they share the critical part of the process most teams skip, which steps you can and cannot automate, and a free plugin that lets a coding agent do most of the heavy lifting for you. Enjoy To go deeper, join their upcoming AI Evals for Engineers & PMs course https://maven.com/parlance-labs/evals?promoCode=LENNYSLIST and use the discount code LENNYSLIST at checkout to get 25% off. By now, you’ve probably heard that evals are a defining skill for AI PMs. Mike Krieger, Anthropic’s former CPO and now head of Labs, has said that “if there’s one thing we can teach product people, it’s that writing evals is now probably the most important thing.” Garry Tan, the CEO of Y Combinator, shared that “evals are emerging as the real moat for AI startups.” A number of guests on Lenny’s Podcast have argued that “evals are the new PRDs,” and increasingly, leading companies have been talking about how investing in evals has paid off: - Shopify https://shopify.engineering/fine-tuning-agent-shopify-flow used evals to guide development of an AI workflow builder that was 2.2 times faster and 68% cheaper than the frontier-model system it replaced. - Cursor https://cursor.com/blog/how-cursor-router-works developed its Auto Balance routing performance with evals, resulting in much higher user satisfaction while reducing costs by 41%. - Ramp https://engineering.ramp.com/post/apple-intelligence-receipt-matching increased its automatic receipt collection product’s accuracy from 35% to 83% after investing in evals. - Harvey https://www.harvey.ai/blog/rebuilding-playbook-review-as-a-multi-agent-system rebuilt its AI contract reviewer with evals, nearly doubling the product’s internal quality score. Rippling https://www.rippling.com/blog/building-mcp-server , Glean https://www.glean.com/blog/glean-skills-launch-2026 , Abridge https://tech.abridge.com/blog/offline-evaluations-to-improve-our-systems , ElevenLabs https://elevenlabs.io/blog/eleven-v3-is-now-generally-available , and Robinhood https://robinhood.com/us/en/careers/blog/loop-stays-dumb-so-the-model-can-be-smart/ have also shared how they’ve been using evals to systematically make their AI products better. AI products are easy to change but hard to predict. A prompt, model, or code change can improve one behavior while breaking another. Evals turn your judgment about what “good” looks like into repeatable tests your team can run before shipping. Production errors flagged by evals can become additional test cases that improve your AI, creating an advantage that compounds over time. And now that AI can produce changes faster than people can review them, evals help teams ship quickly by automatically checking if a product still works as intended. In our prior post https://www.lennysnewsletter.com/p/building-eval-systems-that-improve , we laid out the full process for building evals: discover and analyze errors, create customized metrics, and set up a continuous improvement loop. Unfortunately, we’ve found that most teams skip the first stage of error discovery and jump straight to writing metrics. It is easy to see why. Looking through lengthy user session records to find failures feels slow and hard to scale, whereas metrics are concrete and easy to automate. But if you write metrics too early, you end up making too many assumptions about what’s important—and potentially measuring the wrong thing, or the right thing poorly. This is why error discovery is the eval equivalent of product discovery. Just as product discovery shows which problems are worth solving, error discovery reveals which AI failures are worth measuring. Without it, teams risk building dashboards around generic metrics that waste time and steer the product toward the wrong outcomes. We believe error discovery is so important that if you only have time for one part of the eval process, you should prioritize it. In this post, we’ll show you the three steps to running effective error discovery with a coding agent like Codex or Claude. We’ve used this process with more than 50 companies, and each time the tools uncovered major product flaws that were hurting the customer experience. This entire workflow takes only about 30 minutes to complete once you learn the basics. Note: Error discovery has changed a lot since our last post. Our prior post called this process “error analysis.” We now call it “error discovery,” because the goal is to identify failures that are worth measuring. Keep reading to learn about the new approach. Find the errors that matter to your product When you’re building an AI product, you need evals to understand where it makes mistakes. But maintaining evals costs time and money; it doesn’t make sense to measure everything. Good error discovery identifies which failures are worth measuring and tracking over time. Even when you know you need the error discovery stage, it can be tempting to start by handing an agent a folder of traces—complete records of user sessions with your AI product—and asking it to find problems. Agents are often faster than humans at spotting obvious issues and can find patterns we might miss. But they are far less reliable when a failure depends on your definition of a good product experience. You can explain those standards to the agent, but you often only discover them in the first place by reviewing the data. This process, where reviewing examples changes your definition of good, is called criteria drift https://arxiv.org/abs/2404.12272 . For example, below is an interaction from Nurture Boss https://nurtureboss.com/ , an AI leasing assistant we worked with that helps property managers handle conversations with prospective tenants: Prospect: “This is out of my budget. Thank you for your business.” Leasing assistant: “You’re welcome If your situation changes or if you have any other questions in the future, feel free to reach out. Have a great day ” The full conversation from this interaction is provided below in the discussion on traces. To most agents, this looks like a success; they wouldn’t identify an error in this trace. But the product goal is to facilitate sales, which includes finding the right property matches for prospects’ different needs. In this situation, the agent should have explored cheaper units or other properties owned by the same company and offered the prospect alternatives. If we’d prompted the agent to look for this “objection handling” failure up front when we gave it traces to check, the agent would have caught the error automatically. But we’d only know to add “objection handling” to our criteria after seeing this trace ourselves. This is a classic case of criteria drift—and why it’s critical to take a step back to review failures and define success before dispatching an agent to find errors in a stack of traces. In a broader study https://parlance-labs.com/blog/posts/auto-evals/index.html , we ran automated eval tools and coding agents against 100 production traces from this same apartment-leasing assistant. We found the following: 1. Agents missed issues requiring product judgment and context outside the trace such as Markdown formatting in text messages and missed human handoffs in addition to objection handling . 2. Agents are good at catching failures that are obvious inside a trace, like answers that are contradicted by tool output. 3. Agents also find issues that humans miss , but they introduce noise by flagging good responses as failures. Clearly, automated approaches are still useful for finding some types of errors, and they work especially well in conjunction with human judgment. So how do you benefit from an agent’s automation while keeping a human in the loop? The answer is a process that draws on active learning