{"slug": "show-hn-a-practical-ai-evaluation-pattern", "title": "Show HN: A practical AI Evaluation pattern", "summary": "Public AI benchmarks should be used as capability screens to build a shortlist rather than as deployment decisions, according to a practical evaluation framework published by deepsense.ai. The framework cites OpenAI's July 2026 estimate that roughly 30% of SWE-Bench Pro tasks were broken, GDPval's one-shot evaluation across 44 occupations, SWE-Lancer's more than 1,400 freelance software tasks with $1 million in real payouts, and Harness-Bench's finding of substantial performance differences across model-harness pairings on a shared task environment. The piece argues teams should evaluate the whole system — model, harness, tools, memory, execution path, reliability, cost, and business outcome — under production-like conditions, because once a public benchmark becomes influential it can become a target for optimization and its score narrows in meaning.", "body_md": "Table of contents\n\n**Should AI leaders still follow model benchmarks? Yes, but mainly to understand **what is worth testing next**, not to assume that the top-performing model will be the best fit for the particular use case.**\n\nIn this issue:\n\n1. We’ll explain **where public benchmarks are useful, where they stop being meaningful** , and how to use them to build a credible shortlist rather than make a deployment decision.\n2. We’ll show **how to evaluate AI systems as a whole** — model, harness, tools, memory, execution path, reliability, cost, and business outcome — under production-like conditions.\n3. And we’ll give you a **practical framework** for comparing model candidates, diagnosing failure modes, and deciding what is actually ready to ship.\n\nOff the leaderboard, into the real world. First, a quick look at where public evaluation stands today.\n\n## Public Benchmarks vs. Production AI Evaluation\n\nIn 2026, public AI evaluation is moving closer to real work. **Several benchmarks have moved beyond assessing narrow academic tasks** toward end-to-end professional workflows.\n\nThese benchmarks test not only whether models can produce correct outputs, but also whether they can use tools, complete complex tasks, and do so reliably across repeated attempts. [SWE-Lancer](https://openai.com/index/swe-lancer/?utm_source=chatgpt.com) covers more than 1,400 freelance software tasks, with $1 million in real payouts. GDPval evaluates professional deliverables across 44 occupations. τ-bench checks whether tool-using agents reach the correct state in a simulated business system and whether they can repeat that result across multiple runs.\n\nThen there is our own [EDA benchmark](https://deepsense.ai/tech-expertise/llms-rag/custom-synthetic-datasets-for-llm-vlm-evaluation-and-training/?utm_source=LinkedIn&utm_medium=Newsletter_26_08_2026_1&utm_campaign=AI_Eval), designed to test whether AI agents can solve structured analytical problems accurately and repeatedly across diverse domains.\n\nThat is progress. **But it does not close the gap between a benchmark and production.**\n\n### Benchmarks Are Getting Better. The Gap Still Remains\n\nGDPval’s first version, for example, is still a one-shot evaluation. It does not capture iterative workflows in which a system gathers context, uses several tools, revises its work, and responds to feedback. In July 2026, OpenAI also estimated that [roughly 30% of SWE-Bench Pro tasks were broken](https://openai.com/index/separating-signal-from-noise-coding-evaluations/), mainly because they required overly strict tests or underspecified prompting. And [Harness-Bench](https://arxiv.org/abs/2605.27922?utm_source=chatgpt.com) found substantial performance differences across model-harness pairings, even when the underlying task environment was shared.\n\n*The lesson is not that public benchmarks have nothing to offer. They are useful* *capability screens**that help teams build a shortlist and compare models under common conditions.*\n\nStandardized evaluations also improve transparency and reproducibility. **But a benchmark score is not a deployment decision.**\n\n### When the Benchmark Becomes the Target\n\nThere is another limitation. **Once a public benchmark becomes influential, it can become a target for optimization.**\n\nProviders may tune training data, post-training, prompts, inference strategies, or agent harnesses to perform well on a particular evaluation. Public benchmark tasks may also find their way into training data. Scores can therefore improve without equivalent gains on new tasks or production workloads.\n\nThis does not necessarily mean deliberate gaming. It is a predictable consequence of **repeatedly optimizing against a visible metric**. A system can become very good at a benchmark’s task distribution, format, and grader without becoming equally capable on different users, data, tools, and constraints.\n\n*The benchmark may still be useful, but the meaning of its score narrows: it indicates that a particular system* *performed well under a specific evaluation protocol**. It does not prove that the system will generalize to your production environment.*\n\nThat is why **benchmarks should help determine what to evaluate next** – not what to deploy.\n\n### Use Benchmarks to Shortlist AI Models, Not to Ship Them\n\nA public benchmark asks: Can this model solve this type of task under this test setup? An enterprise evaluation must ask:\n\n*Can our system repeatedly deliver a useful business outcome**, using the right evidence and actions, at an acceptable cost, speed, and level of risk?*\n\nThose are different questions.\n\nIn production, the model is only one part of the result. The rest comes from system prompts, tool definitions, MCP servers, context management, memory, permissions, retries, validation, output contracts, and orchestration.\n\nThat is the boundary we need to draw clearly: **public benchmarks** measure specified configurations under standardized test conditions. **Production evaluation** measures the system you intend to deploy under representative operating conditions.\n\nOnce we started evaluating AI systems this way, the same lessons kept resurfacing. **They can be distilled into five principles for making evaluation a credible basis for deployment decisions.**\n\n## 1. A Correct Answer Can Still Be a Failed Run\n\nA final answer can look correct while hiding a broken workflow. We saw this while **evaluating systems built around MCP servers**. In some cases, the underlying API returned factually correct data. The model could read the schema and reason over the raw response. On a simple task, that was enough.\n\nBut factual accuracy did not guarantee a useful final result. Some responses were technically correct but:\n\n- **difficult to aggregate** , e.g. returning separate nested records instead of a simple total,\n- **missing units or business context** , e.g. reporting “12” without specifying whether it means users, dollars, or percent,\n- **poorly shaped for a downstream agent** , e.g. returning prose where structured JSON is required,\n- **or disconnected from the user’s actual objective** , e.g. listing sales figures without explaining whether the target was met.\n\n*That creates two separate evaluation questions:* *Was the retrieved information correct?**And was it returned in a form that allowed the next system component to use it reliably?*\n\n**A tool response can pass the first test and fail the second.**\n\n### The Final Answer Is Only One Layer of the AI Evaluation\n\nWe saw the same issue in longer workflows. A model produced a plausible statement about a revised economic series, but it had not checked the available vintage dates before making the claim. Another returned a believable release date but answered from memory rather than using the release-calendar tool.\n\nIn other runs, the system skipped a required MCP call, selected the wrong series, inspected the right source but used the wrong observation, or produced a clean summary from incomplete evidence. **Checking only the final text would miss all of these failures.**\n\nA production evaluation therefore needs three layers:\n\n1. **Evidence** : Did the system retrieve the correct and sufficient information?\n2. **Execution** : Did it perform the actions the task required?\n3. **Deliverable** : Was the result complete, usable, and aligned with the user’s objective?\n\nFor a **revision-analysis task**, the evaluation may require the system to inspect vintage dates before drawing a conclusion. **A release-calendar task** may require a call to the release-calendar tool. For **a data-extraction task**, it may verify both the factual values and the presence of the labels, units, dates, and context required by the downstream agent.\n\nThis does not mean enforcing one exact reasoning path for every task. Agents may find several valid ways to solve a problem.\n\nThe rule is narrower:\n\n*Enforce actions that protect correctness, policy, or auditability.**Allow flexibility everywhere else.*\n\n### Evaluate the Trajectory, Not Just the Outcome\n\nCurrent agent-evaluation guidance already moves in this direction. [Anthropic separates](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?utm_source=chatgpt.com) the trajectory from the final environmental outcome, while [OpenAI recommends](https://developers.openai.com/api/docs/guides/agent-evals?utm_source=chatgpt.com) trace grading to inspect tool calls, routing, handoffs, and guardrails. These behaviors can be measured at multiple levels — for example, by verifying that a required action occurred somewhere in a large execution trace, comparing expected and observed action sequences, or scoring how closely the agent’s trajectory matches a reference workflow.\n\nTrajectory-aware evaluation is not a future research idea. **For critical agent workflows, it is already a requirement.**\n\nThat brings us to the second principle: if the trajectory matters, so does the environment that produces it. And that environment is the harness.\n\n## 2. The Harness Is Part of the Product\n\nIn production, model performance is shaped by more than the model itself. The surrounding harness can materially affect the result.\n\nBy **harness**, we mean the runtime around the model that includes (non exhaustively): system prompts, context construction, tool schemas and descriptions, tool execution, memory and state, permissions, retries and recovery, validation, routing and handoffs, and the logic that decides when the workflow is complete.\n\nFor example, during the [development of MCP servers](https://deepsense.ai/tech-expertise/mlops/custom-mcp-servers-as-part-of-enterprise-ai-infrastructure/?utm_source=LinkedIn&utm_medium=Newsletter_26_08_2026_2&utm_campaign=AI_Eval), **we evaluated the same underlying model in two setups**: first through a relatively direct connection to the MCP server, and then through the production-like agent harness it was actually designed to run behind.\n\nThe difference was material. We saw changes in token consumption, tool selection, the number of tool calls, final output quality, consistency across runs, and recovery from weak intermediate results.\n\n*The direct model evaluation* *did not fully predict how the system behaved in its intended runtime.*\n\nThat leads to a simple rule: **match the evaluation setup to the decision you are trying to make**. If you want to isolate model quality, keep prompts, tools, budgets, and execution rules consistent across candidates. If you want to make a deployment decision, test the configuration that would actually run in production — including the harness, prompts, tool definitions, API versions, memory, retries, validation, and recovery logic.\n\nThe same applies to tool interfaces. In our tests, technically correct raw API data was not always enough; labels, metadata, context, and response structure could determine whether the agent used that data reliably in the next step. **So compare models under common conditions** when you need a controlled baseline, but evaluate the full production setup when you need to decide what to deploy.\n\n### Memory and Artifacts Are Also Harness Decisions\n\nOpenAI recently showed [how much this can matter on ARC-AGI-3](https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/?utm_source=chatgpt.com): enabling retained reasoning and compaction in the harness increased GPT-5.6 Sol’s score from **13.3% to 38.3%**, while using roughly **6x fewer output tokens**. The model did not change; the way the system preserved reasoning and managed context did.\n\nWe saw the same principle in our own work at [deepsense.ai](http://deepsense.ai/). In our **EDA benchmark and related synthetic evaluation environments****for data-driven reasoning**, agents work through multi-step analytical tasks: inspecting files, cleaning and combining data, detecting anomalies, reconstructing events, and producing structured outputs.\n\nIn these experiments, weak runs often rebuilt parsers and simulators inside one-off scripts. They relied on printed summaries and conversational context rather than saved files or durable state.\n\n**That produced**: duplicated work, truncated evidence, inconsistent calculations, and loss of earlier evidence later in the task.\n\nPerformance improved when the workflow created **durable intermediate artifacts**: reusable parsers, candidate tables, simulation outputs, threshold sweeps, evidence traces, and final comparison files.\n\nThis is primarily good system design. But when **durable evidence is necessary for a reliable result, the evaluation should explicitly test for it**.\n\nWe saw the same issue in a multi-round task. We introduced a /next_round API and a K_DEPENDENCY rule that limited how many previous folders remained visible. The goal was not to test whether the agent could reread its entire history, but whether it could **preserve the decision-relevant state it would need later.**\n\nThat exposed a different class of failure.\n\n*The agent could solve an individual round correctly, then* *lose an earlier decision, forget a constraint**, or reconstruct an incorrect summary several rounds later.*\n\nA static benchmark might call the model capable. **A multi-round evaluation showed whether the system could remain capable over time.**\n\nThat leads directly to a broader question about memory. If a benchmark removes access to old files, it must state whether the agent may copy important information into a scratchpad before those files disappear.\n\nIn an enterprise system, this is not an artificial benchmark rule. It maps directly to real design questions:\n\n- What may the agent persist?\n- For how long?\n- With what provenance?\n- Within which tenant boundary?\n- Can the state be audited or deleted?\n- May one workflow reuse data from another?\n\n**Memory is a harness decision** — and, in production, a product, security, and governance policy.\n\n## 3. AI Reliability Is a Distribution, Not a Screenshot\n\nSo, once the full setup is under evaluation, the next question is whether it can deliver the same quality consistently, not just once.\n\nOne successful run proves very little. We saw large differences between repeated attempts on the same task. A model could produce one strong result followed by several weak or failed runs.\n\n*Reporting the best attempt would make the system look capable.**It would not make the system reliable.*\n\n### Average Model Performance Can Hide Instability\n\nWe saw the same pattern in our [EDA Benchmark](https://deepsense.ai/blog/eda-benchmark-leaderboard-july-14-2026-update/?utm_source=LinkedIn&utm_medium=Newsletter_26_08_2026_4&utm_campaign=AI_Eval), where each model completed 10 analytical tasks five times. Claude Fable 5 achieved the highest **mean score** at 0.50, while GPT-5.6 Sol scored slightly lower on average at 0.48.\n\nBut once we adjusted for variance across repeated runs, the ranking flipped: **GPT-5.6 Sol led with a reliability-adjusted score* of 0.46, versus 0.42 for Claude Fable 5.**\n\n*That is why we separate* *average quality from repeatability**. Two models may appear to have similar performance based on the mean score, yet differ significantly in how predictably they deliver that performance.*\n\n** Reliability-adjusted score combines mean task performance with run-to-run stability: mean score × exp(-2.25 × CoV^0.88), where CoV is the coefficient of variation across repeated trajectories. Higher values indicate a stronger combination of quality and repeatability; the score is bounded between 0 and 1.*\n\n### AI Reliability Has to Be Measured Across Runs\n\nFor production systems, variance must be a first-class metric. An agent with a **75% single-run** **success rate** may sound acceptable, but that number does not tell us how consistently it will succeed across repeated executions. **If the runs were independent, the probability of three consecutive successes would be about 42% (0.75³).**\n\nFor repeated executions of the same or highly similar task, runs are correlated, so the actual pass^3 sits higher, closer to the 75% single-run rate as per-task difficulty varies more.\n\nThat is the point of pass^k: **to measure how performance holds up when we require repeated success**, rather than infer reliability from one successful run.\n\nThis distinction matters because different products have different tolerance for retries. A **research assistant** may generate several candidates and keep the best one. A **customer-support agent, transaction agent, or compliance workflow** usually needs to behave correctly on both the first and subsequent attempts.\n\nFor every serious system comparison, report:\n\n- number of trials,\n- mean or median quality,\n- standard deviation or coefficient of variation,\n- first-attempt success,\n- consistency across repeated trials,\n- cost per successful outcome,\n- total execution time,\n- and the rate of critical failures.\n\n**Do not hide variance behind one average.**\n\n*A model with a* *slightly higher mean score but large variance may be a worse production choice**than one with a lower mean and much more stable behavior.*\n\nCost belongs in the same decision. A model that gains two quality points but uses five times as many tokens and takes three times as long is not automatically better.\n\nThe useful question is not: How much does one run cost? It is: **How much does one successful, reliable outcome cost?**\n\nConsistency alone is not enough. A reliable system still needs to optimize for the right outcome — and tell us why it failed when it did.\n\n## 4. Score the Business Outcome. Then Study the Failure.\n\nExact-answer scoring works well when only one answer can be correct. It works poorly for many enterprise problems.\n\n**Consider an optimization task.** Two agents may produce different schedules, order plans, or allocation dictionaries. Both outputs may be valid. One may produce slightly higher profit. Another may reduce risk or avoid a costly operational constraint.\n\nComparing both outputs against a single reference dictionary produces a brittle 0-or-1 score. It measures similarity to the reference. **It does not measure value.**\n\n### Score the Outcome, Not the Reference Answer\n\nFor these tasks, score the resulting business utility: profit, cost reduction, service level, recall of high-risk cases, time saved, or another outcome linked to the actual decision.\n\nNormalize the utility to a clear range such as 0–1. Keep hard requirements — policy violations, safety failures, invalid actions, or budget breaches — as separate constraints.\n\nA useful evaluation task should be: **bounded enough to score, but open enough to require judgment.**\n\n*If the* *environment is too open, the result becomes difficult to reproduce. If it is too constrained**,* *a frontier model can solve it mechanically without demonstrating useful planning.*\n\n**Consider a supply-chain task**. It is too open if the instruction is simply: “Improve next year’s inventory strategy.” Without a fixed dataset, planning horizon, operational constraints, or scoring criteria, different agents may make different assumptions and effectively solve different problems. The results become difficult to compare or reproduce.\n\nIt is too constrained if the task provides a complete forecast, a fixed formula, and step-by-step instructions, then asks the agent to reproduce one reference allocation. That is easy to score, but it tests whether the agent can follow a recipe—not whether it can investigate evidence, manage trade-offs, or develop a useful plan.\n\n*The useful middle defines the data, objective, and hard constraints**, while leaving the agent to decide how to analyze the evidence and construct the solution.*\n\nBut measuring the outcome is only the first step. A score tells us how much value the system created; it does not tell us how the system arrived there or why it fell short. To improve the system, evaluation must therefore do more than rank results. It must also diagnose the path that produced them.\n\n### Diagnose Why the AI System Score Was Low\n\nIn one supply-chain experiment, **a weak run did not fail because the model hallucinated**. The pipeline was broadly reasonable. The real problem was methodological. The model used an overly conservative inference strategy: it optimized for precision, missed too many true positives, and failed to inspect enough saved evidence before setting its thresholds.\n\n*A final score could show that the run was weak.* *It could not explain why.*\n\nAcross these evaluations, we can identify **several recurring failure modes**. This is not an exhaustive taxonomy, but a practical way to describe some of the patterns we repeatedly observe:\n\n1. **Evidence problems** : wrong, incomplete, or low-quality source data.\n2. **Action problems** : skipped tool, wrong tool, or incorrect parameters.\n3. **State problems** : lost context, stale state, or incorrect memory use.\n4. **Method problems** : weak planning, calibration, thresholding, or inference.\n5. **Output problems** : malformed, incomplete, or unusable results.\n6. **Unsupported generation** : claims that are not grounded in available evidence.\n7. **Operational problems** : timeouts, high cost or latency, permission errors, or failed recovery.\n\nThe point is not to force every failure into a fixed taxonomy. It is about moving beyond a single score and understanding what actually went wrong — because different failures require different fixes.\n\nThat turns evaluation into an engineering feedback system. **A leaderboard tells you which system scored higher.** **Failure analysis tells you what to fix.**\n\n### Automate the Triage, Keep Humans in the Loop\n\nParts of this analysis can be automated. In one of our project’s setups, Harbor — an open-source framework for [running and analyzing agent evaluations](https://github.com/harbor-framework/harbor?utm_source=chatgpt.com) in sandboxed environments — stores full evaluation trajectories.\n\nWe then use a lightweight model to extract observations and cluster recurring problems. Engineers review the resulting categories and inspect the most important traces. The model does not replace expert review; it reduces the amount of undifferentiated trace reading.\n\n[OpenAI used a related pattern](https://openai.com/index/separating-signal-from-noise-coding-evaluations/) when auditing SWE-Bench Pro: an automated pipeline and investigator agents flagged suspicious tasks, while experienced software engineers made the final judgments and resolved disagreements.\n\nThat is the practical model: **automate collection and first-pass analysis. Keep humans responsible for the quality** of failure categories and for expert judgment on subtle decisions.\n\n## 5. Deployment Is Not the Finish Line. It Is the First Data Point.\n\nEven a strong pre-deployment evaluation is only a snapshot — and for many teams, that is where evaluation stops. But once the system goes live, **the operating environment begins to change**, making production the next, and most important, evaluation environment.\n\nBut launch is also where the system enters a changing environment. Providers update model versions and serving behavior. APIs behind MCP servers change their schemas, latency, and failure modes. User behavior drifts away from the tasks you designed. A prompt section that earned its tokens in March may be compensating for behavior that no longer exists in September.\n\n*Within months,* *the evaluation you ran before deployment may no longer reflect how the system operates in production.*\n\nA deployment decision is therefore not the end of evaluation. **It is the point at which evaluation changes form.**\n\n### After Launch, AI Evaluation Becomes Monitoring\n\nBefore launch, evaluation acts as a gate: repeated trials, trajectory grading, failure analysis, and a go/no-go decision.\n\nAfter launch, the same six views — outcome, evidence and actions, usability, reliability, economics, and failure mode — become signals to monitor over time. Not one number on a dashboard. The distributions, trajectories, and recurring failures matter too.\n\nIn practice, **this means three loops running at different speeds.**\n\n### The Fast Loop: Monitor the Distribution, Not the Average\n\nSection 3 argued that reliability is a distribution. Production is where that distribution accumulates enough observations to characterize it more reliably. Track first-attempt success, retry rates, cost per successful outcome, latency, and the rate of critical failures — **and alert on changes in variance, not just changes in the mean**. A system whose average quality remains stable while its variance doubles is not stable. It may be failing for a subset of users even while the aggregate metric looks unchanged.\n\n### The Medium Loop: Turn Production Failures into Evaluation Tasks\n\nEvery production incident is a **candidate evaluation task you did not think to write.** When a run fails, analyze it using the same failure-mode lens used before launch — evidence, action, state, method, output, unsupported generation, or operational — while remaining open to new failure modes. Promote recurring cases into the regression suite.\n\nIn our deployments, the tasks that catch the most regressions are often not the ones we designed up front. **They are reconstructions of real failures**. The pre-launch task set reflects expectations. The post-launch task set reflects reality, and becomes more relevant as production cases accumulate.\n\n### The Slow Loop: Re-Earn the Deployment Decision\n\nAny material change to the system — a model upgrade, a new tool, a rewritten prompt section, or a harness migration — should trigger the evaluation battery again under conditions comparable to those used in the original decision. Same tasks. Same trial counts. Same reporting. Not because the new configuration is likely to be worse, but because **“likely” is exactly the kind of claim this article has argued against.**\n\nAn upgrade that gains two points of mean quality but loses first-attempt reliability is a regression for a transaction agent, whatever the leaderboard says.\n\n### Your AI Evaluation Stack Becomes Deployment Infrastructure\n\nThis is also the answer to a question every team eventually faces: **When a new model is released, how do we know whether to switch?** Not from the announcement benchmarks. From your own regression suite, your own harness, your own repeated trials, and your own cost per successful outcome.\n\nThe evaluation infrastructure built for the first deployment decision is what makes every subsequent decision faster and cheaper. There is a governance dividend, too. **Continuous evaluation produces the artifacts that risk owners and auditors need:** evidence that the system still behaves as approved, a record of what changed and when, and a documented response when performance degraded.\n\nA one-time pre-launch report cannot provide that. A living evaluation loop can.\n\n*The useful question is no longer: Did the system pass evaluation? It is:* *Would it still pass today?*\n\n## A Practical Enterprise AI Evaluation Pattern\n\nA useful evaluation flow looks like this:\n\n*In the picture: use-case definition → public benchmark shortlist → representative business tasks → production-like harness → repeated trials → outcome and trajectory grading → failure analysis → deployment decision → monitoring → new tasks from production failures → back to use case definition*\n\nFor each task, record six views:\n\n1. **Business outcome** : Did the system create useful value? Score utility on a consistent scale.\n2. **Required evidence and actions** : Did it use the necessary sources, tools, checks, and state transitions?\n3. **Output usability** : Can the user or downstream agent act on the result?\n4. **Reliability** : Does performance hold across repeated runs and changing context?\n5. **Economics and operations** : What are the cost, token use, latency, runtime, and recovery characteristics?\n6. **Failure mode** : When the system failed, what actually caused it?\n\n*Do not collapse these into one number too early.* *A single score may be useful for ranking candidates. It is not enough for approving a deployment.*\n\n**The model at the top of a public benchmark may still be the wrong choice**. A cheaper model may produce more value through a better harness. A slower model may be more reliable.\n\nThe takeaway to keep from this issue is simple:\n\n*A model wins a benchmark, but a system earns deployment.*\n\nTable of contents", "url": "https://wpnews.pro/news/show-hn-a-practical-ai-evaluation-pattern", "canonical_source": "https://deepsense.ai/blog/the-top-scoring-model-is-not-always-the-best-production-choice-how-to-evaluate-ai-systems-beyond-public-benchmarks/", "published_at": "2026-10-02 10:20:07+00:00", "updated_at": "2026-10-02 10:37:03.079693+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-agents", "ai-research"], "entities": ["deepsense.ai", "OpenAI", "SWE-Lancer", "GDPval", "SWE-Bench Pro", "Harness-Bench", "τ-bench", "EDA benchmark"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-a-practical-ai-evaluation-pattern", "markdown": "https://wpnews.pro/news/show-hn-a-practical-ai-evaluation-pattern.md", "text": "https://wpnews.pro/news/show-hn-a-practical-ai-evaluation-pattern.txt", "jsonld": "https://wpnews.pro/news/show-hn-a-practical-ai-evaluation-pattern.jsonld"}}