cd/entity/Promptfoo· home› entities› Promptfoo
grep -l @promptfoo /news/*.json | wc -l → 26

Promptfoo

mentions 26 type Organization page 1/2 feed RSS

// recent coverage 26 mentions

11:06
2026-09-18
dev.to
ai-agents

Your SKILL.md is production config. Test it like one.

A developer released skilldiff, an open-source tool that tests agent skill files like SKILL.md for behavioral regressions by running them in a real agent harness against a fixture repo and asserting o…

05:51
2026-09-04
dev.to
ai-safety

Promptfoo + Humanbound

Promptfoo, used by over 300,000 developers and 156 Fortune 500 companies for red teaming AI agents, can now integrate with Humanbound's firewall training pipeline. The integration allows teams to impo…

22:20
2026-09-03
dev.to
ai-tools

Six open-source AI workflow kits you can actually inspect

Software Sausage released six open-source AI workflow kits, each containing a README, an editable evidence ledger, and a dependency-free shell verifier. All 17 verifiers pass at release v0.17.0, and t…

04:08
2026-08-23
byteiota.com
ai-products

Claude Agent Stack Goes GA: Computer Use, Skills, Files

Anthropic made computer use, browser use, the Skills API, and the Files API generally available on August 20, marking a production-ready agent stack. The new computer toolset reduces round trips by 20…

21:11
2026-08-20
discuss.huggingface.co
ai-tools

Looking for simple ways to evaluate an AI agent

Promptfoo is recommended as the primary evaluation tool for AI agents focused on documentation and RAG tasks, with Ragas and LangSmith suggested for deeper analysis. The guidance from Promptfoo, Huggi…

12:53
2026-08-16
promptcube3.com
artificial-intelligence

AI red teaming tools

Developers often rely on manual, ad-hoc testing for LLM features, but a more systematic approach using automated red teaming tools like Giskard and Promptfoo can uncover critical vulnerabilities quick…

20:59
2026-08-04
dev.to
large-language-models

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

EvalPort introduces a grader system with 11 types for LLM evaluation, including exact_match, semantic_similarity, llm_judge, and custom, designed to be framework-agnostic and self-describing. The syst…

03:40
2026-07-30
dev.to
large-language-models

OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval, a new open-source project, aims to standardize LLM evaluation by defining a portable JSON Schema for test cases, graders, and results. The project provides SDKs, a CLI, and converters for po…

09:00
2026-07-29
pydantic.dev
ai-agents

The best AI agent optimization platforms in 2026

Pydantic Logfire ranks as the best AI agent optimization platform in 2026, according to a Pydantic analysis that evaluated tools on their ability to close the loop by proposing and shipping production…

21:56
2026-07-23
promptcube3.com
ai-safety

Open Source AI Community, AI red teaming tools, Wi

AI red teaming tools like Giskard, DeepEval, and Promptfoo automate adversarial testing to systematically find edge-case failures in model logic, moving beyond manual 'vibe checks' that risk PR disast…

06:30
2026-07-14
agentic.tracebit.com
ai-safety

Context bombs: stopping AI attackers in their tracks

A new defensive technique called a 'context bomb' — a short string hidden in decoy resources that triggers safety guardrails in offensive AI agents — reduced autonomous cyberattack success by roughly …

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics