cd/entity/Braintrust· home entities Braintrust
grep -l @braintrust /news/*.json | wc -l → 29

Braintrust

mentions 29 type Organization page 1/2 feed RSS

// recent coverage 29 mentions

00:00
2026-07-24
chaliy.name
ai-tools

You Do Not Need a Server for Evals

Evals for AI projects like coding agents and sandboxed bash do not require a dedicated server or platform, according to developer Everruns. Datasets, runners, and results can be stored and versioned d…

19:27
2026-07-22
dev.to
artificial-intelligence

An LLM judge is a biased instrument, not a measurement

A developer found that an LLM judge gave opposite results for the same eval run on consecutive days due to position bias, one of three systematic biases documented in the 2023 paper "Judging LLM-as-a-…

18:34
2026-07-20
news.ycombinator.com
artificial-intelligence

Ask HN: How do I reliably eval my AI models

A Hacker News user asks how to reliably evaluate AI models, mentioning Braintrust and Arize as potential tools, and inquires about sourcing experts to create gold datasets.…

09:31
2026-07-12
startupfortune.com
ai-agents

How to Evaluate AI Agents Before You Ship Them to Real Users

Most founders shipping AI agents lack systematic evaluation methods, leading to public failures like Chevrolet's chatbot agreeing to sell a car for $1 and McDonald's AI drive-thru adding bacon to ice …

21:19
2026-07-07
rightmodeler.com
large-language-models

Get recommended a cheaper model with this skill

Rightmodeler replays real agent traces through cheaper models, judges outputs against shipped versions, and shows cost-saving opportunities with evidence. The tool supports traces from Claude Code, Co…

00:00
2026-07-02
pydantic.dev
ai-agents

Observability tools agents want

Pydantic Logfire and other observability platforms are shipping MCP servers, CLIs, and SDKs that allow AI agents to directly inspect traces, logs, prompts, evals, and dashboards, shifting the focus fr…

18:05
2026-07-01
newsletter.port.io
ai-agents

How to build a context lake that saves you 80% on token costs

Port's experiment found that routing AI agents through a structured context lake instead of direct-to-MCPs cut token costs by 58%, and adding a skill file brought savings to 80%. The context lake pre-…

00:00
2026-06-30
1password.com
ai-safety

Braintrust's Ankur Goyal: Code review doesn't cover prompts

Braintrust CEO Ankur Goyal warned that prompt changes in AI agents often bypass security review, creating a gap where behavior-shaping updates can alter what agents do, what data they access, and whic…

00:00
2026-06-26
vercel.com
ai-agents

Trace and debug eve agent sessions with Vercel Observability

Vercel launched Agent Runs, a new observability feature for eve agent sessions that provides curated views of triggers, duration, token usage, and step-level details without OpenTelemetry setup. The t…

04:40
2026-06-16
discuss.huggingface.co
large-language-models

Metrics for Text Generation from T5 Model

A user training a T5 model asked for alternative metrics to Exact Match for evaluating text generation. Community members suggested ROUGE-1, ROUGE-2, and BLEU, and recommended Braintrust for running e…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics