cd /news/artificial-intelligence/peoplesearchbench-a-multi-dimensiona… · home topics artificial-intelligence article
[ARTICLE · art-76474] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms

Researchers introduced PeopleSearchBench, an open-source benchmark for evaluating AI-powered people search platforms, and found that Lessie, a specialized AI people search agent, scored 65.2 overall, 18.5% higher than the second-ranked system, and achieved 100% task completion across 119 real-world queries. The benchmark compares four platforms on relevance precision, effective coverage, and information utility using a Criteria-Grounded Verification pipeline with high human validation (Cohen's kappa = 0.84).

read1 min views1 publishedJul 28, 2026

arXiv:2603.27476v2 Announce Type: replace Abstract: AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted benchmark exists for evaluating their performance. We introduce PeopleSearchBench, an open-source benchmark that compares four people search platforms on 119 real-world queries across four use cases: corporate recruiting, B2B sales prospecting, expert search with deterministic answers, and influencer/KOL discovery. A key contribution is Criteria-Grounded Verification, a factual relevance pipeline that extracts explicit, verifiable criteria from each query and uses live web search to determine whether returned people satisfy them. This produces binary relevance judgments grounded in factual verification rather than subjective holistic LLM-as-judge scores. We evaluate systems on three dimensions: Relevance Precision (padded nDCG@10), Effective Coverage (task completion and qualified result yield), and Information Utility (profile completeness and usefulness), averaged equally into an overall score. Lessie, a specialized AI people search agent, performs best overall, scoring 65.2, 18.5% higher than the second-ranked system, and is the only system to achieve 100% task completion across all 119 queries. We also report confidence intervals, human validation of the verification pipeline (Cohen's kappa = 0.84), ablations, and full documentation of queries, prompts, and normalization procedures. Code, query definitions, and aggregated results are available on GitHub.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @peoplesearchbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/peoplesearchbench-a-…] indexed:0 read:1min 2026-07-28 ·