{"slug": "evaluating-openai-s-privacy-filter-cross-lingual-cross-domain-pii-detection-42", "title": "Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks", "summary": "A new independent evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter PII detector, across 42 benchmarks in 22 languages and 5 domains finds that OPF outperforms Presidio and XLM-RoBERTa on PII-annotated benchmarks (F1=0.855 on AI4Privacy vs. 0.431 and 0.269) but degrades sharply on narrative prose and non-Latin scripts (Arabic F1=0.04, Cyrillic F1=0.03). The study, posted on arXiv (2608.02616v1), also shows GPT-4o leads on medical, legal, and financial PII (SPY avg 0.643, Gretel 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (0.60).", "body_md": "arXiv:2608.02616v1 Announce Type: new\nAbstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synthetic benchmarks spanning 22 languages and 5 domains. Zero-shot, OPF achieves F1=0.855 on AI4Privacy and 0.464 on SPY medical, outperforming Presidio (0.431, 0.273) and XLM-RoBERTa (0.269, 0.111) on PII-annotated benchmarks; on multilingual NER, XLM-RoBERTa leads OPF on all 13 Indic and non-Latin languages. GPT-4o leads on medical, legal, and financial PII (SPY: 0.643 avg, Gretel: 0.527), while OPF leads on structured synthetic PII (0.71 avg) and customer support (0.60). OPF degrades sharply when PII is embedded in narrative prose: F1=0.04--0.57 on NER benchmarks and collapse for non-Latin scripts (Arabic: 0.04, Cyrillic: 0.03). Error analysis shows OPF is strongest on structurally regular PII types (email: 0.78, phone: 0.76) and weakest on culturally variable ones (person: 0.40, address: 0.49), and is recall-biased on customer-support and medical/legal PII (P=0.31--0.54, R=0.70--0.85); global precision spans 0.31--0.86 across all domains.", "url": "https://wpnews.pro/news/evaluating-openai-s-privacy-filter-cross-lingual-cross-domain-pii-detection-42", "canonical_source": "https://arxiv.org/abs/2608.02616", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:03:10.021309+00:00", "lang": "en", "topics": ["artificial-intelligence", "natural-language-processing", "ai-research", "ai-safety"], "entities": ["OpenAI", "Privacy Filter", "Presidio", "XLM-RoBERTa", "GPT-4o", "AI4Privacy", "SPY", "Gretel"], "alternates": {"html": "https://wpnews.pro/news/evaluating-openai-s-privacy-filter-cross-lingual-cross-domain-pii-detection-42", "markdown": "https://wpnews.pro/news/evaluating-openai-s-privacy-filter-cross-lingual-cross-domain-pii-detection-42.md", "text": "https://wpnews.pro/news/evaluating-openai-s-privacy-filter-cross-lingual-cross-domain-pii-detection-42.txt", "jsonld": "https://wpnews.pro/news/evaluating-openai-s-privacy-filter-cross-lingual-cross-domain-pii-detection-42.jsonld"}}