{"slug": "large-language-models-can-follow-instructions-but-not-many-at-once-phase-in", "title": "Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction", "summary": "A new study introduces Constraint Saturation Evaluation (CSE), a benchmark testing large language models' ability to follow multiple simultaneous constraints, and finds that per-constraint pass rates decay gradually while the chance of satisfying all constraints collapses: a model passing individual constraints at ~41% at k=8 succeeds on all eight just 5.7% of the time. The study, covering 15 models, 36 constraint types, and 369,753 checks at k=1-12, shows that structural constraints lose 2x more baseline capability per added constraint than lexical ones, and reliable instruction following breaks down beyond 5-6 simultaneous constraints, with probe-level success falling below 50% at 7 constraints for the strongest model and at 3 or fewer for 12 of 15 models.", "body_md": "arXiv:2608.12426v1 Announce Type: new\nAbstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where many must hold jointly, remains poorly characterized: how rapidly does performance degrade, what governs the degradation, and can the collapse be mitigated? We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement: 15 models, 36 constraint types, 369,753 checks at k=1-12. Three findings emerge. First, per-constraint pass rate decays gradually and predictably, while the chance of satisfying all k constraints collapses - a model passing individual constraints at ~41% at k=8 succeeds on all eight just 5.7% of the time. Second, constraints do not degrade equally: structural constraints lose 2x more baseline capability per added constraint than lexical ones, ordered by a comprehension-maintenance gap that separates constraints requiring sustained tracking from binary decisions immune to composition. Third, failures are nearly independent, which is what makes the accumulation multiplicative; the residual coupling that does exist tracks shared output features rather than pairwise interference - a wrong sentence count fails every constraint that reads it. Reliable instruction following breaks down beyond 5-6 simultaneous constraints: probe-level success falls below 50% at 7 constraints for the strongest model, and at 3 or fewer for 12 of 15.", "url": "https://wpnews.pro/news/large-language-models-can-follow-instructions-but-not-many-at-once-phase-in", "canonical_source": "https://arxiv.org/abs/2608.12426", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:09:58.205790+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety"], "entities": ["Constraint Saturation Evaluation (CSE)", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/large-language-models-can-follow-instructions-but-not-many-at-once-phase-in", "markdown": "https://wpnews.pro/news/large-language-models-can-follow-instructions-but-not-many-at-once-phase-in.md", "text": "https://wpnews.pro/news/large-language-models-can-follow-instructions-but-not-many-at-once-phase-in.txt", "jsonld": "https://wpnews.pro/news/large-language-models-can-follow-instructions-but-not-many-at-once-phase-in.jsonld"}}