cd/entity/Inspect AI· home entities Inspect AI
grep -l @inspect ai /news/*.json | wc -l → 9

Inspect AI

mentions 9 type Organization feed RSS

// recent coverage 9 mentions

16:35
2026-09-01
dev.to
ai-agents

How to Design AI Evaluations You Can Actually Trust

Google's Developer Relations team published a suite of Agent Skills for Google products on GitHub and outlined five rules for designing trustworthy AI evaluations. The rules emphasize understanding th…

07:00
2026-09-01
dev.to
ai-research

Step up to the Sheets: AI Eval Export and Illustrating Data

A developer from Google AI detailed a pandas-based pipeline that converts Inspect AI evaluation logs into a CSV optimized for Google Sheets, enabling non-technical stakeholders to create boardroom-rea…

00:00
2026-08-25
0xff.nu
artificial-intelligence

hexbench

The hex-bench benchmark, a personal evaluation suite for the hex project, tests personal-assistant workflows using Inspect AI's react solver against a local OpenAI-compatible llama.cpp endpoint, with …

07:00
2026-08-18
dev.to
artificial-intelligence

Designing AI Evals: Clarity Now and Visualization Next

A developer from Google Cloud's DevRel team demonstrates how to design objective evaluations for AI agent skills using open-source frameworks like Inspect AI and Harbor. The investigation uses Gemini …

20:59
2026-08-04
dev.to
large-language-models

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

EvalPort introduces a grader system with 11 types for LLM evaluation, including exact_match, semantic_similarity, llm_judge, and custom, designed to be framework-agnostic and self-describing. The syst…

03:40
2026-07-30
dev.to
large-language-models

OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval, a new open-source project, aims to standardize LLM evaluation by defining a portable JSON Schema for test cases, graders, and results. The project provides SDKs, a CLI, and converters for po…

21:51
2026-07-22
lesswrong.com
ai-safety

A Multi-Agent Extension for Petri

Meridian Labs and Anthropic have developed an extension for the open-source AI safety evaluation framework Petri that enables multi-agent evaluations, addressing a key limitation of the original singl…

// co-occurs with top 8 entities
// topics top 6 topics