{"slug": "glyph-a-multi-strategy-agentic-system-for-column-description-and-sensitivity-of", "title": "Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs", "summary": "Researchers Kostia Kudriavtsev, Parvez Rafi, and Sha Sundaram presented Glyph, a production multi-agent LLM system that generates column descriptions and assigns sensitivity-ontology tags for enterprise data catalogs, in a paper published September 2026. Glyph fine-tunes a 6-layer MiniLM metadata encoder with an in-batch contrastive objective, lifting same-tag retrieval on an in-distribution held-out split from NDCG@10 0.55 to 0.92 and MAP@100 from 0.19 to 0.90 over the stock base encoder. The Tagger assigns labels from a governed 275-leaf Data Classification Ontology using three parallel strategies fused with Reciprocal Rank Fusion, and the system is evaluated end-to-end under a recall-weighted F2 objective across three groups.", "body_md": "[content type paper](https://machinelearning.apple.com/research/)published September 2026\n\nGlyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs\n\nAuthorsKostia Kudriavtsev, Parvez Rafi, Sha Sundaram\n\nEnterprise data lakes accumulate tables faster than human stewards can document or classify them, leaving columns with missing descriptions and unassigned governance labels. This documentation debt undermines data discovery, access control, and regulatory compliance. We present Glyph, a production system that frames two coupled problems, column description generation and column type annotation for data classification, as cooperating LLM agents orchestrated as stateful graphs. The Descriptor grounds generation in the pipeline source code that produces each column, retrieved on demand from an enterprise GitHub via a reasoning–acting tool loop (active Retrieval-Augmented Generation). The Tagger assigns labels from a governed 275-leaf Data Classification Ontology by running three complementary strategies in parallel (a description tagger, a line-of-business regex tagger, and a metadata tagger backed by a fine-tuned contrastive encoder over a vector database), then fuses their ranked outputs with Reciprocal Rank Fusion (RRF). We fine-tune a 6-layer MiniLM metadata encoder with an in-batch contrastive objective, lifting same-tag retrieval on an in-distribution held-out split from NDCG@10 0.55 to 0.92 (MAP@100 0.19→ 0.90) relative to the stock base encoder. We report end-to-end multi-label tagging quality under a recall-weighted F2 objective across three evaluation groups, an ablation isolating each strategy and the RRF fusion, and the engineering decisions that distinguish Glyph from prior column-type-annotation work and from commercial value/regex sensitivity scanners: value-free and code-grounded design, per-tag provenance, and graceful degradation. Together these make multi-agent LLM cataloging auditable and operable as a production service.\n\nSemantic Regexes: Auto-Interpreting LLM Features with a Structured Language\n\nDecember 3, 2025[research area Human-Computer Interaction](https://machinelearning.apple.com/research/?domain=Human-Computer%20Interaction), [research area Methods and Algorithms](https://machinelearning.apple.com/research/?domain=Methods%20and%20Algorithms)[conference ICLR](https://machinelearning.apple.com/research/?event=ICLR)\n\nAutomated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, these natural language feature descriptions are often vague, inconsistent, and require manual relabeling. In response, we introduce semantic regexes, structured language descriptions of LLM features. By combining primitives that capture linguistic and semantic feature patterns with modifiers for contextualization,…\n\nLyric Document Embeddings for Music Tagging\n\nFebruary 2, 2022[research area Data Science and Annotation](https://machinelearning.apple.com/research/?domain=Data%20Science%20and%20Annotation), [research area Methods and Algorithms](https://machinelearning.apple.com/research/?domain=Methods%20and%20Algorithms)\n\nWe present an empirical study on embedding the lyrics of a song into a fixed-dimensional feature for the purpose of music tagging. Five methods of computing token-level and four methods of computing document-level representations are trained on an industrial-scale dataset of tens of millions of songs. We compare simple averaging of pretrained embeddings to modern recurrent and attention-based neural architectures. Evaluating on a wide range of…", "url": "https://wpnews.pro/news/glyph-a-multi-strategy-agentic-system-for-column-description-and-sensitivity-of", "canonical_source": "https://machinelearning.apple.com/research/glyph-column-description-tagging", "published_at": "2026-09-16 00:00:00+00:00", "updated_at": "2026-09-16 19:43:48.766573+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "ai-tools"], "entities": ["Glyph", "Kostia Kudriavtsev", "Parvez Rafi", "Sha Sundaram", "MiniLM", "Data Classification Ontology", "Reciprocal Rank Fusion", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/glyph-a-multi-strategy-agentic-system-for-column-description-and-sensitivity-of", "markdown": "https://wpnews.pro/news/glyph-a-multi-strategy-agentic-system-for-column-description-and-sensitivity-of.md", "text": "https://wpnews.pro/news/glyph-a-multi-strategy-agentic-system-for-column-description-and-sensitivity-of.txt", "jsonld": "https://wpnews.pro/news/glyph-a-multi-strategy-agentic-system-for-column-description-and-sensitivity-of.jsonld"}}