cd/sources/arize-auto-discovered· home› sources› Arize (auto-discovered)
cat /sources/arize-auto-discovered.feed | wc -l → 95

Arize (auto-discovered)

articles 95 domain arize.com → page 5/5 feed RSS
13:39
2026-06-02
arize.com
artificial-intelligence

AI benchmarks are breaking. Trace analysis is what comes next.

AI agents are increasingly exploiting benchmark designs, rendering pass/fail metrics unreliable for measuring true capability. In recent months, Anthropic's Claude Opus decrypted a benchmark's answer …

13:31
2026-05-29
arize.com
ai-agents

How to build a better agent harness with traces and evals

Arize AI cofounder and CPO Aparna Dhinakaran demonstrated a method for improving AI agents by building a better harness around the model, using traces and evals to debug failures. In a live demo with …

06:48
2026-05-29
arize.com
ai-agents

Agent harnesses have an expiration date

Agent harnesses built on Claude Code's implicit finish condition—which assumes a model is done when it sends a text-only response without tool calls—are failing across multiple model generations, caus…

17:57
2026-05-20
arize.com
artificial-intelligence

What we learned testing 7 models under the same agent harness

Seven large language models tested under the same agent harness showed similar correctness scores but significant differences in operational behavior, including latency, tool-call counts, and timeout …

← prev page 5 / 5