cd /news/artificial-intelligence/evaluating-openai-s-privacy-filter-c… · home topics artificial-intelligence article
[ARTICLE · art-87107] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

A new independent evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter PII detector, across 42 benchmarks in 22 languages and 5 domains finds that OPF outperforms Presidio and XLM-RoBERTa on PII-annotated benchmarks (F1=0.855 on AI4Privacy vs. 0.431 and 0.269) but degrades sharply on narrative prose and non-Latin scripts (Arabic F1=0.04, Cyrillic F1=0.03). The study, posted on arXiv (2608.02616v1), also shows GPT-4o leads on medical, legal, and financial PII (SPY avg 0.643, Gretel 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (0.60).

read1 min views1 publishedAug 5, 2026

arXiv:2608.02616v1 Announce Type: new Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synthetic benchmarks spanning 22 languages and 5 domains. Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated benchmarks; on multilingual NER, XLM-RoBERTa leads OPF on all 13 Indic and non-Latin languages. GPT-4o leads on medical, legal, and financial PII (SPY: 0.643 avg, Gretel: 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (0.60). OPF degrades sharply when PII is embedded in narrative prose: F1=0.04--0.57 on NER benchmarks and collapse for non-Latin scripts (Arabic: 0.04, Cyrillic: 0.03). Error analysis shows OPF is strongest on structurally regular PII types (email: 0.78, phone: 0.76) and weakest on culturally variable ones (person: 0.40, address: 0.49), and is recall-biased on customer-support and medical/legal PII (P=0.31--0.54, R=0.70--0.85); global precision spans 0.31--0.86 across all domains.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-openai-s-…] indexed:0 read:1min 2026-08-05 ·