{"slug": "neuro-symbolic-ai-for-leed-compliance-document-centric-benchmarking-numeric-and", "title": "Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts", "summary": "A neuro-symbolic pipeline using a 4-billion-parameter language model (gemma3:4b) achieved 67.3% accuracy in screening LEED v4.1 BD+C certification documents, outperforming an 8-billion-parameter model (llama3.1:8b), according to a new arXiv preprint. The deterministic numeric checker corrected arithmetic errors on key quantitative credits, raising EA-p2 accuracy from 50% to 100%, but the full neuro-symbolic configuration trailed the best text-only baseline at 61.6% overall accuracy due to extraction failures and conservative behavior on qualitative categories. Adding low-resolution drawing images consistently reduced accuracy, and prompt effectiveness varied with project documentation density.", "body_md": "arXiv:2607.15647v1 Announce Type: new\nAbstract: LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether small, locally deployed language models can perform meaningful screening of LEED documentation and how deterministic symbolic components should share that work. A neuro-symbolic pipeline is introduced that aligns project PDFs to LEED credit sections, retrieves evidence with credit-aware keyword signatures, verifies compliance with a locally hosted 4-billion-parameter language model, and applies a LEED-specific numeric checker to quantitative thresholds. Experiments on four university buildings (484 PDFs, 153 credit-level decisions) show that a 4-billion-parameter model (gemma3:4b) is the strongest text-only core verifier, achieving 67.3% accuracy and outperforming a larger 8-billion-parameter model (llama3.1:8b) in this task. The deterministic numeric checker corrects arithmetic errors on key quantitative credits, moving EA-p2 from 50% to 100% accuracy and improving several other credits when required values are reliably extracted. At the same time, the full neuro-symbolic configuration achieves 61.6% overall accuracy, trailing the best text-only baseline due to extraction failures and conservative behavior on qualitative categories. Systematic ablations show that adding low-resolution drawing images (150-300 dpi) consistently reduces accuracy, and that prompt effectiveness depends on the building's ground-truth PASS rate: rubric prompts perform best on documentation-rich projects, while chain-of-thought prompts perform best on documentation-lean projects. Within the specific scope of LEED v4.1 BD+C compliance verification over raw project documentation, this pipeline and its baselines provide an initial reproducible reference point for both accuracy and failure modes.", "url": "https://wpnews.pro/news/neuro-symbolic-ai-for-leed-compliance-document-centric-benchmarking-numeric-and", "canonical_source": "https://arxiv.org/abs/2607.15647", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:53:17.981289+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["LEED", "gemma3:4b", "llama3.1:8b"], "alternates": {"html": "https://wpnews.pro/news/neuro-symbolic-ai-for-leed-compliance-document-centric-benchmarking-numeric-and", "markdown": "https://wpnews.pro/news/neuro-symbolic-ai-for-leed-compliance-document-centric-benchmarking-numeric-and.md", "text": "https://wpnews.pro/news/neuro-symbolic-ai-for-leed-compliance-document-centric-benchmarking-numeric-and.txt", "jsonld": "https://wpnews.pro/news/neuro-symbolic-ai-for-leed-compliance-document-centric-benchmarking-numeric-and.jsonld"}}