{"slug": "institution-specific-llm-prompting-recovers-phi-that-de-identification-systems", "title": "Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss", "summary": "A study from arXiv (2608.17051v1) found that large language models (LLMs) with in-context learning outperform purpose-built de-identification systems in recovering institutionally situated protected health information (PHI) from electronic health records. On 100 annotated pediatric oncology notes from Texas Children's Hospital, the best LLM prompt achieved an F1 score of 0.918, compared to 0.779 for Stanford TiDE, and recovered 79% of missed institutional PHI categories. The final calibrated prompt reached a recall of 0.981, demonstrating that well-calibrated prompting can close the institutional PHI gap and control the precision-recall trade-off in a single LLM call per note.", "body_md": "arXiv:2608.17051v1 Announce Type: new\nAbstract: Secondary use of electronic health records requires de-identification, yet existing systems miss \\emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off.\nOn 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric.\nLLMs outperformed the purpose-built systems (best F1=0.918$\\pm$0.001 vs.\\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\\pm$0.002).\nWell-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard.\nLLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.", "url": "https://wpnews.pro/news/institution-specific-llm-prompting-recovers-phi-that-de-identification-systems", "canonical_source": "https://arxiv.org/abs/2608.17051", "published_at": "2026-08-19 04:00:00+00:00", "updated_at": "2026-08-19 04:11:45.938242+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research"], "entities": ["arXiv", "Texas Children's Hospital", "Stanford TiDE", "OpenMed PII"], "alternates": {"html": "https://wpnews.pro/news/institution-specific-llm-prompting-recovers-phi-that-de-identification-systems", "markdown": "https://wpnews.pro/news/institution-specific-llm-prompting-recovers-phi-that-de-identification-systems.md", "text": "https://wpnews.pro/news/institution-specific-llm-prompting-recovers-phi-that-de-identification-systems.txt", "jsonld": "https://wpnews.pro/news/institution-specific-llm-prompting-recovers-phi-that-de-identification-systems.jsonld"}}