{"slug": "syntax-vs-semantics-how-transformers-learn-deep-dependencies", "title": "Syntax vs. Semantics: How Transformers Learn Deep Dependencies", "summary": "A new arXiv paper (2608.26139v1) proposes a mechanistic framework showing that large language models learn deep semantic dependencies through a competition between surface statistics and deep semantics, identifying a 'Gradient Starvation' phenomenon that suppresses sparse semantic error signals early in training and causes structural reasoning to emerge as a sudden phase transition. The framework explains Chain-of-Thought effectiveness by externalizing reasoning steps to bypass suppression, validated on models from toy transformers to Llama-3.1-8B and Qwen2.5-Coder-7B, and introduces a topology-aligned contrastive objective that improves variable binding task performance by over 2x compared to standard cross-entropy fine-tuning.", "body_md": "arXiv:2608.26139v1 Announce Type: new\nAbstract: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as a competition between Surface Statistics and Deep Semantics. Our theoretical analysis identifies a ``Gradient Starvation\" phenomenon where the error signals for sparse semantic dependencies are actively suppressed during early optimization. This suppression impedes the learning of structural reasoning and causes its emergence to manifest as a sudden phase transition. Furthermore, this framework offers a mechanistic basis for the effectiveness of Chain-of-Thought (CoT) strategies. By externalizing intermediate reasoning steps into concrete tokens, CoT effectively bypasses the suppression regime inherent to implicit reasoning. We validate these findings across scales ranging from toy transformers to production models (Llama-3.1-8B, Qwen2.5-Coder-7B). Finally, guided by this theory, we propose a topology-aligned contrastive objective that explicitly rectifies the gradient geometry. Experiments on variable binding tasks demonstrate that our method achieves an improvement that is over 2x larger than that obtained via standard cross-entropy fine-tuning.", "url": "https://wpnews.pro/news/syntax-vs-semantics-how-transformers-learn-deep-dependencies", "canonical_source": "https://arxiv.org/abs/2608.26139", "published_at": "2026-08-28 04:00:00+00:00", "updated_at": "2026-08-28 04:20:38.717217+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["arXiv", "Llama-3.1-8B", "Qwen2.5-Coder-7B"], "alternates": {"html": "https://wpnews.pro/news/syntax-vs-semantics-how-transformers-learn-deep-dependencies", "markdown": "https://wpnews.pro/news/syntax-vs-semantics-how-transformers-learn-deep-dependencies.md", "text": "https://wpnews.pro/news/syntax-vs-semantics-how-transformers-learn-deep-dependencies.txt", "jsonld": "https://wpnews.pro/news/syntax-vs-semantics-how-transformers-learn-deep-dependencies.jsonld"}}