cd/entity/DeepEval· home› entities› DeepEval
grep -l @deepeval /news/*.json | wc -l → 32

DeepEval

mentions 32 type Organization page 1/2 feed RSS

// recent coverage 32 mentions

20:47
2026-09-22
deepeval.com
ai-research

Show HN: JevEval, evals using Jev-as-a-judge

DeepEval released JevEval, a custom LLM evaluation metric that uses Jev as a judge to convert explicit questions and calibrated probabilities into deterministic scores. JevEval separates evaluation in…

18:17
2026-09-21
neo4j.com
generative-ai

The AI application your team can actually stand behind

Neo4j detailed a GraphRAG-based generative AI framework it built internally to ground LLM answers in a Neo4j Aura knowledge graph, combining vector search with an entity layer of types such as Documen…

00:00
2026-08-28
posthog.com
ai-tools

Best AI evaluation tools for production

PostHog is named the best overall LLM evaluation tool for production, according to a new guide, because it ties eval scores to user sessions, traces, and feature flags, while Braintrust is highlighted…

09:00
2026-08-26
pydantic.dev
artificial-intelligence

The 7 best LLM evaluation tools in 2026

Pydantic Logfire is the top pick among seven LLM evaluation tools compared in a guide updated August 26, 2026, which also covers Braintrust, Langfuse, LangSmith, Arize Phoenix, Confident AI, and Galil…

00:54
2026-08-21
dev.to
large-language-models

RAG - Hallucination Detection

A developer explains how to detect hallucinations in retrieval-augmented generation (RAG) systems, where an LLM generates responses not supported by the retrieved context. Techniques include comparing…

21:11
2026-08-20
discuss.huggingface.co
ai-tools

Looking for simple ways to evaluate an AI agent

Promptfoo is recommended as the primary evaluation tool for AI agents focused on documentation and RAG tasks, with Ragas and LangSmith suggested for deeper analysis. The guidance from Promptfoo, Huggi…

18:58
2026-08-12
deepeval.com
developer-tools

DeepEval Open-Sourced for TypeScript

DeepEval has open-sourced its TypeScript SDK in beta, enabling all 47 of its 49 metrics to run as a gate in CI/CD pipelines via a single Vitest matcher. The SDK, which compiles the same language-neutr…

20:59
2026-08-04
dev.to
large-language-models

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

EvalPort introduces a grader system with 11 types for LLM evaluation, including exact_match, semantic_similarity, llm_judge, and custom, designed to be framework-agnostic and self-describing. The syst…

20:38
2026-08-03
dev.to
ai-agents

I have been Vibecoding Evals (works better than I thought)

A developer experimenting with AI coding agents found that adding evals to a support-triage app caught a subtle misclassification that manual testing missed. The app, built for a fictional shipment-tr…

03:40
2026-07-30
dev.to
large-language-models

OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval, a new open-source project, aims to standardize LLM evaluation by defining a portable JSON Schema for test cases, graders, and results. The project provides SDKs, a CLI, and converters for po…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics