cd/entity/DeepEval· home entities DeepEval
grep -l @deepeval /news/*.json | wc -l → 22

DeepEval

mentions 22 type Organization page 1/2 feed RSS

// recent coverage 22 mentions

00:54
2026-08-21
dev.to
large-language-models

RAG - Hallucination Detection

A developer explains how to detect hallucinations in retrieval-augmented generation (RAG) systems, where an LLM generates responses not supported by the retrieved context. Techniques include comparing…

21:11
2026-08-20
discuss.huggingface.co
ai-tools

Looking for simple ways to evaluate an AI agent

Promptfoo is recommended as the primary evaluation tool for AI agents focused on documentation and RAG tasks, with Ragas and LangSmith suggested for deeper analysis. The guidance from Promptfoo, Huggi…

18:58
2026-08-12
deepeval.com
developer-tools

DeepEval Open-Sourced for TypeScript

DeepEval has open-sourced its TypeScript SDK in beta, enabling all 47 of its 49 metrics to run as a gate in CI/CD pipelines via a single Vitest matcher. The SDK, which compiles the same language-neutr…

20:59
2026-08-04
dev.to
large-language-models

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

EvalPort introduces a grader system with 11 types for LLM evaluation, including exact_match, semantic_similarity, llm_judge, and custom, designed to be framework-agnostic and self-describing. The syst…

20:38
2026-08-03
dev.to
ai-agents

I have been Vibecoding Evals (works better than I thought)

A developer experimenting with AI coding agents found that adding evals to a support-triage app caught a subtle misclassification that manual testing missed. The app, built for a fictional shipment-tr…

03:40
2026-07-30
dev.to
large-language-models

OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval, a new open-source project, aims to standardize LLM evaluation by defining a portable JSON Schema for test cases, graders, and results. The project provides SDKs, a CLI, and converters for po…

21:56
2026-07-23
promptcube3.com
ai-safety

Open Source AI Community, AI red teaming tools, Wi

AI red teaming tools like Giskard, DeepEval, and Promptfoo automate adversarial testing to systematically find edge-case failures in model logic, moving beyond manual 'vibe checks' that risk PR disast…

19:27
2026-07-22
dev.to
artificial-intelligence

An LLM judge is a biased instrument, not a measurement

A developer found that an LLM judge gave opposite results for the same eval run on consecutive days due to position bias, one of three systematic biases documented in the 2023 paper "Judging LLM-as-a-…

15:30
2026-07-11
dev.to
large-language-models

How to Add Evals to an LLM Feature

A developer explains how to add evals to an LLM feature, using an outbound AI calling agent as an example. The process involves defining a business outcome metric, curating a representative dataset of…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics