cd/entity/Braintrust· home entities Braintrust
grep -l @braintrust /news/*.json | wc -l → 47

Braintrust

mentions 47 type Organization page 1/3 feed RSS

// recent coverage 47 mentions

02:44
2026-08-31
gist.github.com
ai-agents

The Verification Gap: Who Audits the Agents?

A developer argues that while AI-generated software has become cheap, the cost of trusting it remains high, creating a 'verification gap' in the AI industry. The article highlights Stellar Wave's bug-…

15:35
2026-08-30
seldon-ai.com
ai-infrastructure

Observational reconstruction, then the inverse problem

Seldon, an AI infrastructure company, argues that many production LLM calls are repeated behavioral contracts that cheaper models could serve, and its router and Import Audit tools aim to reconstruct …

00:00
2026-08-28
posthog.com
ai-tools

Best AI evaluation tools for production

PostHog is named the best overall LLM evaluation tool for production, according to a new guide, because it ties eval scores to user sessions, traces, and feature flags, while Braintrust is highlighted…

09:00
2026-08-26
pydantic.dev
artificial-intelligence

The 7 best LLM evaluation tools in 2026

Pydantic Logfire is the top pick among seven LLM evaluation tools compared in a guide updated August 26, 2026, which also covers Braintrust, Langfuse, LangSmith, Arize Phoenix, Confident AI, and Galil…

09:00
2026-08-26
pydantic.dev
ai-tools

The best Langfuse alternatives in 2026, honestly compared

A comparison of Langfuse alternatives in 2026 finds that teams leave Langfuse due to its billing model, lack of log and metric ingestion, enterprise-only governance features, and its acquisition by Cl…

21:11
2026-08-20
discuss.huggingface.co
ai-tools

Looking for simple ways to evaluate an AI agent

Promptfoo is recommended as the primary evaluation tool for AI agents focused on documentation and RAG tasks, with Ragas and LangSmith suggested for deeper analysis. The guidance from Promptfoo, Huggi…

16:01
2026-08-10
promptcube3.com
ai-agents

Claude Code agents fail because we treat them like synchronous

Claude Code agents fail because developers treat them like synchronous code, according to a developer's analysis of execution transcripts. The article identifies three failure modes—silent context ove…

09:00
2026-08-06
pydantic.dev
artificial-intelligence

Do evals the Airbnb way

Airbnb published a playbook for evaluating generative AI at scale, recommending teams read roughly 100 outputs and traces before building evaluators, then layer programmatic checks, LLM judges, and hu…

05:00
2026-08-05
vercel.com
ai-infrastructure

Export AI Gateway traces with Vercel Drains

Vercel's AI Gateway now generates an OpenTelemetry trace for every request, allowing Pro and Enterprise teams to export traces via Vercel Drains to OTLP/HTTP-compatible endpoints such as Braintrust, D…

09:00
2026-08-03
pydantic.dev
developer-tools

Score freely

Logfire, a full observability suite from Pydantic, announced it charges $0 per thousand AI evaluation scores, attaching each gen_ai.evaluation.result as an OpenTelemetry event billed at the standard t…

00:00
2026-07-24
chaliy.name
ai-tools

You Do Not Need a Server for Evals

Evals for AI projects like coding agents and sandboxed bash do not require a dedicated server or platform, according to developer Everruns. Datasets, runners, and results can be stored and versioned d…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics