WavePrune: One period is often enough for RoPE
WavePrune, a method that restricts each RoPE channel to its first rotation period, raises the HELMET long-context score on four of five tested models without extra tuning, including 35.7 to 40.0 on Qw…
WavePrune, a method that restricts each RoPE channel to its first rotation period, raises the HELMET long-context score on four of five tested models without extra tuning, including 35.7 to 40.0 on Qw…
A developer implemented the FlashAttention-2 forward pass twice — first as a deliberately slow PyTorch reference with a Python double loop over tiles, then as a Triton kernel — to teach kernel optimiz…
Researchers introduced Elastic Threshold Attention (ETA), an end-to-end trainable sparse attention architecture that predicts dynamic, contextual thresholds from query representations to speed up long…
Engineer Celcilin C S introduced CellularFlow, a novel architecture that replaces dense feed-forward networks in transformers with addressable DNA memory banks to solve catastrophic forgetting. The de…
Proxima, an out-of-tree vLLM plugin implementing STAR-KV low-rank KV cache compression, serves 4x more concurrent requests at 8192 context and boots at 16384 context where plain vLLM refuses to start,…
A first-principles analysis of LoRA fine-tuning memory usage reveals that 87.3% of VRAM scaling with sequence length for Llama 3.1 8B is consumed by the cross-entropy loss head tensor, not the model, …
Meituan's LongCat team released version 1.5 of its open-source talking-avatar generator LongCat-Video-Avatar, which replaces the Wav2Vec2 audio encoder with Whisper-Large-v3 and uses DMD2-based step d…
A developer details practical QLoRA fine-tuning using Axolotl and Unsloth, explaining how parameter-efficient methods like LoRA and QLoRA enable training multi-billion parameter models on a single con…
KVBoost is a new open-source Python library that accelerates HuggingFace LLM inference by implementing chunk-level KV cache reuse, achieving 3–5× faster time-to-first-token (TTFT) and up to 85% cache …