Masked Language Flow Models
Researchers introduced Masked Language Flow Models (MLFMs), combining masked diffusion and flow-based methods for efficient language generation. MLFMs enable conditional generation via continuous flow…
Researchers introduced Masked Language Flow Models (MLFMs), combining masked diffusion and flow-based methods for efficient language generation. MLFMs enable conditional generation via continuous flow…
A developer analyzed the open-source ecosystem for reducing LLM costs, focusing on three cost lines: cached input, uncached/written input, and output. The key insight is that tools must preserve the m…
Researchers proposed Dynamic-dLLM, a training-free framework to accelerate Diffusion Large Language Models (dLLMs) by dynamically allocating cache budgets and calibrating decoding thresholds. The meth…
Researchers compared six offline reinforcement-learning methods for distilling reasoning from large language models into smaller ones, finding that SFT, RFT, and RIFT produce nearly identical weight u…
A developer trained a large language model from scratch using plain PyTorch, implementing the full post-training pipeline including SFT, reward modeling, DPO, PPO, and GRPO on public datasets, all run…
Headroom, an open-source context compression layer, reduces token usage in AI agent loops by 60–95% while preserving answer accuracy. The tool intercepts and compresses tool outputs, logs, and convers…
Researchers introduced SEVRA, a serving-layer controller that decides when a frozen reasoning model should verify its answer instead of thinking longer, finding that selective verification improves ac…
Researchers introduced Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads in large language models by measuring their causal impact on reasoning tasks. C…
Researchers introduced GRACE, a theoretical framework that determines the optimal verification granularity for test-time scaling in large language models based on problem difficulty, verifier accuracy…
A research project using k-sparse autoencoders found interpretable features in a small reasoning model, including features that appear to correspond to the model's reasoning process. The experiment an…
Researchers have developed Discrete Tilt Matching (DTM), a likelihood-free method for fine-tuning masked diffusion large language models using reinforcement learning. The approach recasts fine-tuning …
Researchers have developed Fast-dLLM++, a training-free extension to diffusion large language models that accelerates inference by selecting parallel token commit sets based on the full sorted confide…
A developer wrote a fused decode-attention kernel that ran 2.2× faster than the baseline in microbenchmarks, but when integrated into a HuggingFace `generate` call for an RL training loop, the decode …
A new study from arXiv reveals that advanced reasoning models can maintain a factually correct chain-of-thought while simultaneously outputting a wrong answer under sustained adversarial pressure, a f…
Researchers introduced Sequential Bayesian Belief Tracking (SBBT), a framework that estimates the likelihood of a correct final answer from partial reasoning traces by calibrating observation likeliho…
Researchers introduced SPEAR (Sandboxed Prompt Engineer with Active Roll-back), a code-augmented agentic optimizer that autonomously rewrites prompts using a Python sandbox for structural error analys…
A new study reveals that language model reasoning trajectories during test-time sampling cluster into "reasoning basins," causing majority votes to favor stable but potentially incorrect answers. The …
A new study of small language models reveals that chain-of-thought prompting for arithmetic relies on a positional shortcut: the model copies whichever number appears last before the answer delimiter,…