X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Researchers introduced X-AuT, a progressive framework that selects layer combinations to compress audio-encoder depth for speech large language models, addressing the deletion and premature end-of-sequence errors caused by removing complete blocks. The method uses cross-scale distillation to reduce inference cost while limiting perturbation to the embeddings consumed by the decoder. Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combination