cd/entity/Arize· home› entities› Arize
grep -l @arize /news/*.json | wc -l → 58

Arize

mentions 58 type Organization page 1/3 feed RSS

// recent coverage 58 mentions

16:37
2026-10-07
arize.com
artificial-intelligence

Decision model benchmark: Jev, Kev, Liquid d1, and more

Arize benchmarked eight decision models — including TypeSafe's Jev 1.13, Liquid AI's Liquid d1, OpenAI Decisions (gpt-6-luna), Cloudflare's Clef 27B, and Jared Palmer's open Kev 9B — on hallucination …

06:04
2026-10-01
dev.to
large-language-models

The cost of proving it works

Arize's 2026 cost model breaks production evaluation spend into traffic volume × sampling rate × evaluation surfaces × evaluator cost plus human review and retention, while offline evaluation is datas…

16:08
2026-09-29
arize.com
ai-agents

Are agent harnesses dying? What harness distillation changes

A Peking University research team led by Haoran Ye trained the behavior of a specialized agent harness into Qwen3.5-9B via a method called harness distillation, raising the small model's average task …

22:55
2026-09-23
arize.com
ai-research

Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks

A benchmark of 23,325 judgments across the RAGTruth and SummEval human-labeled datasets found that TypeSafe's Jev matched Claude Opus 5 at 87% accuracy on held-out hallucination detection at roughly 1…

14:00
2026-09-23
arize.com
ai-agents

What changes when AI agents use your software

Daytona cofounder and CEO Ivan Burazin said the agents he uses fail more often because they cannot access the tools they need than because of the model itself, and that he does not want agents operati…

14:00
2026-09-09
arize.com
artificial-intelligence

Code mode: Why your agent should code

Anthropic's Claude Code and Codex now have non-engineering variants in Claude Design and Codex Work, and Anthropic's revenue has been growing 10x year over year, according to a blog post by Arize. The…

15:05
2026-08-19
arize.com
artificial-intelligence

Where agent evals are going: Agent-as-a-Judge

A research paper proposing Agent-as-a-Judge, an evaluation method where an AI agent assesses another agent's behavior, was accepted to ICML 2025 and began shipping in products by July 2026, moving fro…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics