Claude’s new auto eval tool
Anthropic released new eval tooling for Claude Code, adding `build_eval` and `hill-climb` commands to its `claude-api` plugin that help developers build evals, check graders, and improve applications …
Anthropic released new eval tooling for Claude Code, adding `build_eval` and `hill-climb` commands to its `claude-api` plugin that help developers build evals, check graders, and improve applications …
Hamel Husain and Shreya Shankar published an AI Evals FAQ distilling the most common questions from teaching 700+ engineers and product managers about AI evaluation. The guide distinguishes model benc…
Hamel Husain, a software engineer, published a blog post arguing that developers should inspect the final prompts that LLM abstraction tools like DSPy, guidance, and instructor send to language models…
Machine learning engineer Hamel Husain argues that automated evaluations (evals) for AI systems are flawed and often misleading, based on his 20+ years of experience and work at Airbnb and GitHub. He …
AI product builders often claim their product is hard to evaluate, but this objection signals a design flaw: if it's hard for developers to verify, it's likely hard for users too. Three case studies s…