WavePrune: One period is often enough for RoPE
WavePrune, a method that restricts each RoPE channel to its first rotation period, raises the HELMET long-context score on four of five tested models without extra tuning, including 35.7 to 40.0 on Qw…
WavePrune, a method that restricts each RoPE channel to its first rotation period, raises the HELMET long-context score on four of five tested models without extra tuning, including 35.7 to 40.0 on Qw…
The Allen Institute for AI released AstaBrief 8B on October 2nd, an open model built on Qwen3-8B that turns research questions and retrieved scientific papers into cited reports and now powers Fast mo…
The Allen Institute for AI (Ai2) released AstaBrief 8B on October 2, 2026, an open-weights model built on Qwen3-8B that generates fully cited scientific reports from a research question and retrieved …
The Allen Institute for AI (Ai2) open-sourced AstaBrief 8B, an 8-billion-parameter model built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited scientific repo…
DEdit, a diffusion-based speculative decoding drafter from an arXiv paper (arXiv:2609.38510v1), achieves macro-average speedups of 5.72x on Qwen3-4B and 5.97x on Qwen3-8B over autoregressive generatio…
Researchers introduced Decode-Latency Feedback Prefill (DLFP), a model-free controller implemented in vLLM that resizes prefill chunks based on observed decode intervals, cutting P99 inter-token laten…
Stanford and Nvidia released CLM-8B on September 23, an open-weights Contrastive Language Model that makes real-time agent decisions up to 9x faster than TypeSafe AI's proprietary Jev model, clocking …
Contrastive Language Models (CLMs) introduced CLM-8B, an 8-billion-parameter System One model trained with a contrastive learning objective that connects states and actions, pre-trained on 60M Nemotro…
A dual-module framework called LEGO, combining Legal Expert GraphRAG and expert Chain-of-Thought, reached 40.53% exact-match accuracy on the LawExamQA_Civil benchmark using a Qwen3-8B backbone, accord…
Researchers Ruiyang Wang and colleagues introduced GAVEL, a framework that verifies and repairs long-horizon LLM robot plans against an explicit graph world model, improving single-task success on BEH…
Software engineer Sean Goedecke outlined two techniques for programming with "System One" language models — models that output only decisions from user-provided multiple-choice questions — in a post d…
Researchers introduced Fathom, a per-query read-depth key scan for sparse decoding over offloaded KV caches, which at one million tokens on Qwen3-8B runs a decode step 1.67x faster in GPU time than th…
A study by Chiara Manna, Argentina Anna Rescigno, and Eva Vanmassenhove compared the automated WinoMT annotation pipeline with the instruction-tuned LLM Qwen3-8B for annotating grammatical gender in E…
Amazon Web Services published a walkthrough for building an AI-powered product tagging system that customizes Qwen3-8B with supervised fine-tuning (SFT) and then optimizes it with reinforcement learni…
Researchers Xianming Li, Zongxi Li and Tsz-fung Andrew Lee merged ShadowPEFT into the main branch of Hugging Face's PEFT library, adding a stateful shadow network that can be detached with unload_shad…
A second-part technical article on latency optimization for open-weight LLM inference details four techniques measured on a fixed deployment of Qwen3-8B (Apache-2.0) served on vLLM through SageMaker's…
Osprey, a target-agnostic speculative decoding method detailed in arXiv:2609.09338v1, raises mean acceptance length by 16.1% on Qwen3-8B, 21.2% on Llama-3.3-70B-Instruct, and 22.7% on MiniMax-M2.5 (22…
Researchers introduced Osprey, a target-agnostic pretraining method for speculative decoding drafters that bootstraps them from off-the-shelf pretrained small language models, according to the arXiv p…
An RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput to 17–19 GB/s from the nominal 45–60 GB/s when total weight size is an integer multiple of 1 MiB, af…
A developer has released EmbedFlow, an open-source tool that enables zero-downtime upgrades of embedding models by retrieving top-K documents with the old model and reranking them with the new model, …