Backdooring Sparse Autoencoders
A paper submitted to arXiv on 5 October 2026 introduces a decoder-only sparse autoencoder (SAE) backdoor that induces attacker-chosen behavior when the modified SAE is inserted into the forward pass o…
A paper submitted to arXiv on 5 October 2026 introduces a decoder-only sparse autoencoder (SAE) backdoor that induces attacker-chosen behavior when the modified SAE is inserted into the forward pass o…
Researchers introduced Sigma, described as the first large-scale continuous diffusion language model at 3B and 8B parameters, built on steerable low-dimensional ODE/SDE latent trajectories and trained…
Researchers posted a paper on arXiv on 1 October 2026 introducing LoopCD, a training-free contrastive decoding framework for looped Transformers that contrasts the final prediction with an earlier rec…
Arize's 2026 cost model breaks production evaluation spend into traffic volume × sampling rate × evaluation surfaces × evaluator cost plus human review and retention, while offline evaluation is datas…
ELF-REG, a method that scales Embedded Language Flows (ELF) continuous diffusion language models to reasoning tasks, reached 55.96% pass@1 on GSM8K at 64 network function evaluations and 13.39% on MAT…
A developer outlined best practices for evaluating AI models, emphasizing a multi-layered framework combining standardized benchmarks like MMLU and HumanEval, adversarial red teaming for prompt inject…
Edge0-35B-A3B, a 35B-parameter sparse mixture-of-experts model built on Qwen3.5-MoE, loses an average of 3.9 points across five benchmarks when quantized to 4-bit through its edge0 inference pipeline …
Researchers from Duke University and Tsinghua University published a paper on September 3, 2026 introducing PlaidQ, a continuous Gaussian latent-diffusion language model for code generation. Distillin…
A deep-dive analysis from tamiz.pro examines the gap between AI agent hype and production deployment, highlighting that frameworks like LangChain, AutoGen, and Haystack with 10K+ GitHub stars signal i…
Researchers introduced Latent Recurrent Thoughts (LRT), a method that enables frozen large language models to reason in continuous representation space using a small recurrent reasoner, outperforming …
A new dependency-free audit script from developer Ashish Sinha found that six human-authored Hugging Face benchmarks—BIRD-CRITIC 1.0, Spider, GSM8K, and HumanEval—are clean, with only 18 genuinely dup…
MetroStar is hiring a Solutions Architect in Indianapolis, IN, offering $78,000–$115,000 per year to support secure, scalable cloud architectures that modernize legacy Marine Corps systems and enable …
A new study from arXiv (2608.23807v1) characterizes serving of masked diffusion language models (dLLMs) using LLaDA-8B-Instruct with a D2F LoRA adapter on a single NVIDIA H200 GPU, finding that reques…
Researchers propose STEP-KTODER, a framework for code preference optimization that defines steps as module-level functions in decomposed multi-function programs and assigns binary correctness labels v…
A new experiment by an unnamed researcher found that flipping a single bit in the Qwen2.5-Coder-3B large language model can reduce its coding accuracy from 85% to near zero, simulating the effect of c…
A team led by Ming Zhong at the University of Illinois Urbana-Champaign and Google DeepMind has introduced SWE-IF, a framework that aligns code evaluation with human preference, revealing that instruc…
Fortitude Omnis Group's OmnisBench benchmark initially showed small models performing surprisingly well, but community feedback revealed potential contamination from old benchmarks like HumanEval and …
Jie Tang's team at Z.ai argues that LLM scaling should optimize for fixed inference budgets, favoring smaller, deeper architectures trained longer over wider models. Their ablation shows a 7B model tr…
A new measurement study across four models finds that the default LoRA target_modules of q_proj and v_proj underperforms a data-driven placement for capability retention: on Llama-3-8B fine-tuning cod…
A new empirical analysis from arXiv (2608.19072v1) finds that large language model (LLM) agents post-training an LLM lock in their training strategy at the very beginning and spend the remaining budge…