The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models A round-trip study of sixteen language models found that serializing tree-structured arithmetic expressions into natural language is a lossy, asymmetric channel, with swapping the generating and extracting model shifting accuracy by up to 60.4 points and the best mixed pair reaching 92.9%. The paper's authors — Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella and Samy Bengio — report that at least 73.6% of round-trip failures originate at generation, and that roughly 3,600 fine-tuning examples sharing the evaluation's operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro. A disjoint-domain regime with new operators and vocabulary also raised every open-weight model, though a gap to the frontier remains. content type paper https://machinelearning.apple.com/research/ published September 2026 The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models AuthorsXavier Suau, Alex Ferrando de las Morenas, Luca Zappella, Samy Bengio When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure operator count, depth, right-branching rather than model family. Third, the channel is trainable: ∼ 3600 fine-tuning examples that share the evaluation’s operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language. With Apple Intelligence, we’re integrating powerful generative AI right into the apps and experiences people use every day, all while protecting their privacy. At the 2025 Worldwide Developers Conference we introduced a new generation of language foundation models specifically developed to enhance the Apple Intelligence features in our latest software releases. We also introduced the new Foundation Models framework, which gives app developers… Syntactic Code Search with Sequence-to-Tree Matching: Supporting Syntactic Search with Incomplete Code Fragments June 20, 2024 research area Human-Computer Interaction https://machinelearning.apple.com/research/?domain=Human-Computer%20Interaction , research area Tools, Platforms, Frameworks https://machinelearning.apple.com/research/?domain=Tools%2C%20Platforms%2C%20Frameworks conference Programming Language Design and Implementation PLDI