{"slug": "using-evals-to-cut-ai-errors-7x", "title": "Using evals to cut AI errors 7x", "summary": "Hex used 1,800 eval cases and an automated hill-climbing loop to cut the wrong-edit rate of its Quick Edits feature from 21% to 3%, the company reported in a blog post. The Quick Edits feature uses a small model (GPT-6 Luna or Claude Haiku 4.5) to restyle charts 20X faster and cheaper than the full Hex agent, and Hex adopted GPT-6 Luna over GPT-5.6 Luna despite a roughly 1 percentage point lower overall pass rate because it made 20% fewer wrong edits at half the cost. Hex said most of the gains came from the harness around the model rather than the prompt, and pointed readers to Anthropic's guide to eval design and hill-climbing.", "body_md": "###### [Blog](https://hex.tech/blog/)\n\n# Hill climbing to glory: using evals to cut AI errors 7x\n\nHow 1,800 evals showed us AI feature quality improvements go beyond prompting\n\nHex's Generative Apps let you build fully custom data apps by chatting with the Hex agent. Last week we launched [Quick Edits](https://learn.hex.tech/docs/share-insights/apps/generative-apps#quick-edits), which uses a small model to let you restyle charts 20X faster and cheaper than with the full agent.\n\nWe've [used evals](https://hex.tech/blog/evaluate-data-agents/) to build AI features at Hex for a while. They're the only reliable way to tell whether a change actually made a feature better when the model's output is non-deterministic. We took a different approach with Quick Edits: we used evals to drive its development from the beginning, not just as a QA step at the end.\n\nInitially, we wrote a few hundred eval cases, and went on to create 1,800 total cases as the feature evolved. Coding agents then tested hundreds of changes against those evals and kept only the ones that scored higher. That process, called hill-climbing, took our wrong-edit rate from 21% to 3%. Surprisingly, most of the gains came from the harness around the model, not the prompt.\n\nIf you're interested in taking a similar approach, Anthropic's [guide to eval design and hill-climbing](https://claude.dev/blog/automating-eval-design-and-hillclimbing/) is a good starting point. Here's what worked for us beyond what they shared:\n\n## Start from the product experience\n\n### How the feature works\n\nQuick edits use a small model (GPT-6 Luna/Claude Haiku 4.5) to read a user's request – \"make the Enterprise line green\" – and propose edits to chart properties in one shot.\n\nIf the user's request is too complicated – \"why didn't the numbers go up??? make them go up!\" – it hands off to the Hex agent, which can make more complex changes to the app.\n\nOur evals focused on making sure the small model gets these edits and handoffs right.\n\n### Picking the right metric to optimize\n\nThere are two ways our feature, and therefore our evals, can fail:\n\n- Wrong edit: making an edit the user didn’t ask for.\n- Incorrect handoff: handing a request to the Hex agent that the small model should have been able to handle.\n\nIt’s tempting to want to minimize both kinds of failures, and maximize overall pass rate. But for this feature, it was much more important for us to minimize incorrect edits than incorrect handoffs.\n\nA wrong edit would change the user's chart in a way they didn't ask for, which is jarring and erodes trust. A wrong handoff just results in the main agent making the right edit slower, which is less bad.\n\nFor example, we adopted GPT-6 Luna even though its overall pass rate was about 1 percentage point lower than GPT-5.6 Luna’s, because it made 20% fewer wrong edits at half the cost.\n\nWhen we began hill climbing, our “wrong edit” rate was 21%, by the end, it was 3%. In real-world usage for our customers we expect the wrong edit rate to be much, much lower than that – we intentionally made our evals as challenging as possible.\n\n## Write evals from real usage, then extend\n\n### Rich eval cases\n\nWe built a suite of more than 1,800 eval cases, starting with a few hundred drawn from real chart-edit requests from internal users. We wanted examples of the kinds of edits people ask for, and the specific words (and languages) they use.\n\nWe extended those core cases to cover all the different types of chart edits (and hand-offs) we wanted to handle.\n\nI read and gave feedback on the first 150-ish synthesized cases, which was enough for LLMs to do a great job synthesizing the rest.\n\nI had GPT-6 Astra generate new cases and used Claude Fable 5.1 to fill in any gaps. Using multiple models gives us confidence the examples are diverse and comprehensive.\n\nWe made sure that many of the test cases are difficult and ambiguous, which gave us more runway to hill-climb against. We wanted to test whether the model could tell which requests it should handle and which it should hand off. A few examples:\n\n- “make it pop” → hand off (nobody knows what this means, including the model)\n- “make it weekly” / “hazlo semanal” → hand off (that’s a data change dressed up as a style change)\n- “hide the legend and only show 2025” → hand off, even though half of it is doable (partial edits are worse than no edit)\n- “make all the lines gray except Canada, hide the legend and make the lines thicker” → apply all three (and don’t touch Canada’s color)\n- “format the y axis as dollars” on a horizontal bar chart → hand off (because the Y-axis is the category axis there)\n\n### Handling fixtures\n\nEach eval case includes the \"user request\" and a \"fixture\" – context about the chart that we send to the model along with the user request.\n\nFor this feature, fixtures represent the starting state of a chart in JSON, including previous edits, to handle prompts like \"no, I liked the old color better\".\n\nFixtures like these can be surprisingly complex and detailed, and are harder to make. We used ~50 base charts, with many variants and edit histories, once again inspired by real usage and extended by Astra/Fable.\n\n## The harness mattered more than the prompt\n\n### Validator > prompt\n\nThe validator is code that checks and cleans up the model's output before we use it to edit the chart.\n\nFor example, the model might return the right hex color but leave out the quotation marks our code expects. We were rejecting those responses even though the model had understood the request. We changed the validator to accept them.\n\nSurprisingly, more than 50% of our hill-climb improvements came from expanding the validator to handle more almost-right outputs, such as:\n\n- Fixing small formatting mistakes, like missing quotation marks, or values returned in the wrong type\n- Rounding numbers to the nearest allowed values\n\nThe validator also enforces rules the model might get wrong: min below max, ISO dates, whether a property can actually take effect on this chart…\n\nValidator improvements increased pass rates by 7 percentage points on Luna and 10 on Haiku. They cut Haiku’s hard errors (outputs our code couldn’t accept) by 99%!\n\nAn unintuitive lesson we learned was that prompt changes weren’t nearly as impactful as we expected they would be. The *vast* majority of pure prompt tweaks failed our evals. In fact, we ended up taking things *out* of the prompt and putting them into the validator, or into the chart property descriptions and labels the model reads.\n\nThe model regularly confused the axes on horizontal charts, for example. Adding “(x axis)” to the relevant property labels and improving their descriptions helped it choose the right setting. We hadn’t been able to fix that with prompt changes.\n\n### Handling different models\n\nWe evaluated both Luna and Haiku to make sure any tweaks work well with multiple models across multiple providers, ensuring this feature worked across our customer base no matter their model setup.\n\nFor example, we initially had a \"no change\" option for the model to say the chart already matched the request. Luna handled it well, but Haiku overused it, with 30X more wrong \"no change\" responses than correct ones. This is another example of behavior we couldn’t fix with prompting; we removed the option and pushed that logic into the validator instead.\n\n## Make sure improvements are real\n\n### \"Holdout\" cases to avoid overfitting\n\nIn our experience, it’s critical to keep a meaningful number of eval cases hidden from the agent doing the hill climbing so that it doesn’t overfit (over-optimize for your evals). Without these holdouts, it’s like a student memorizing the answers to a practice test. They might ace the test, but struggle when the questions change.\n\nOur holdout cases were almost 40% of the eval suite. We only accepted proposed changes when the holdout results improved.\n\n### Tuning noise and attempts\n\nYou also need to consider what it will take to ensure quality eval data. In our case, we found that running only 3 attempts per case was too noisy. Re-running the same suite differed by approximately 8 wrong edits. 10 attempts per case was the sweet spot for us, bringing holdout noise down to 0.2 percentage points.\n\nIt's worth asking the agent to rerun the same [eval set](https://hex.tech/blog/evals/) a few times to see how much results vary before changing anything.\n\n## Fast evals mean faster product iteration\n\nEven at 18,000 attempts per full run (1,800 × 10), a run costs less than $10 on Luna and takes a few minutes at high concurrency.\n\nThat meant the evals weren’t just a final step. We ran them after every change to every part of the feature: the prompt, the validator, the chart properties, the handoff rules, the UX.\n\nWe changed the charts’ underlying properties three times in ten days while the hill-climb was running. Each time, we reran the suite and could immediately see if anything had broken or needed improvement.\n\nIt worked in the other direction too. When the model kept getting something wrong, the fix was often not in the prompt. We renamed a property, added a description, or changed a handoff rule, then reran to check it helped.\n\nWe iterated on the AI and the rest of the feature at the same time, against the same evals.\n\n## Product experience should drive your evals\n\nWe rooted our eval cases, fixtures, metrics to hill-climb, and everything else in the ways we see people using Hex today, and the product experience we wanted to enable for them.\n\nSo if you're building evals, my overall advice is to think carefully about that product experience and design your evals from there.\n\n(Also, feel free to copy this post and paste it into your agent, so it can make use of what we learned, too.)\n\nIf this is is interesting, click below to get started, or to check out opportunities to join our team.", "url": "https://wpnews.pro/news/using-evals-to-cut-ai-errors-7x", "canonical_source": "https://hex.tech/blog/we-used-evals-to-improve-ai-feature/", "published_at": "2026-10-09 00:35:34+00:00", "updated_at": "2026-10-09 00:47:12.333712+00:00", "lang": "en", "topics": ["ai-tools", "ai-products", "large-language-models", "ai-agents", "mlops"], "entities": ["Hex", "Quick Edits", "GPT-6 Luna", "GPT-5.6 Luna", "Claude Haiku 4.5", "GPT-6 Astra", "Claude Fable 5.1", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/using-evals-to-cut-ai-errors-7x", "markdown": "https://wpnews.pro/news/using-evals-to-cut-ai-errors-7x.md", "text": "https://wpnews.pro/news/using-evals-to-cut-ai-errors-7x.txt", "jsonld": "https://wpnews.pro/news/using-evals-to-cut-ai-errors-7x.jsonld"}}