Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models A paper on arXiv (2609.30784v1) proposes a symbiotic architecture that gives frozen large language models audio-understanding capabilities by having an injector module write audio-conditioned vectors directly into the LLM's key-value (KV) cache, without fine-tuning the model's weights. The authors report that the method activates fewer parameters during audio prefilling, outperforms the conventional frozen-LLM approach, and approaches the performance of a fine-tuned audio language model on automatic speech recognition, audio question answering, and acoustic scene classification, while preserving the backbone LLM's original text-only task performance by construction. arXiv:2609.30784v1 Announce Type: cross Abstract: This paper proposes an architecture for equipping large language models LLMs with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value KV cache, enabling the LLM to behave as an audio language model ALM . The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks automatic speech recognition, audio question answering, and acoustic scene classification and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.