04:00
2026-09-15
arxiv.org
large-language-models
LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference
Researchers introduced LayerRoute, a parameter-efficient method for adaptive transformer layer-skipping that combines per-layer hard-gated routing with joint LoRA fine-tuning, according to an arXiv pa…