Speculation Is All You Need
Modal Labs released state-of-the-art DFlash speculators for Qwen 3.5 and Qwen 3.6 models on Hugging Face, achieving 5-20% additional speedups and enabling Qwen 3.5 122B-A10B to run at over 1000 tok/s …
Hugging Face is an AI community platform and company providing a hub for open-source machine learning models, datasets, and demo spaces. It hosts over 500,000 models and is widely used by the AI research community.
Modal Labs released state-of-the-art DFlash speculators for Qwen 3.5 and Qwen 3.6 models on Hugging Face, achieving 5-20% additional speedups and enabling Qwen 3.5 122B-A10B to run at over 1000 tok/s …
A developer released PaneTrans, a browser extension for drag-select region translation and OCR on video/canvas, built on Transformers.js with local processing by default. The tool uses an offscreen do…
Empero AI released Qwythos-9B-Claude-Mythos-5-1M, a 9-billion-parameter open-weights reasoning model distilled from Claude Mythos 5, featuring a 1-million-token context window, native tool use, and a …
Modal and Z Lab released DFlash, a speculative decoding model for Qwen 3.5 397B-A17B, achieving over 4.3x throughput versus baseline and 1.5x versus MTP on HumanEval at concurrency 1. The model uses a…
Microsoft, Hugging Face, Meta's PyTorch team, NVIDIA, and others launched OpenEnv, an open protocol for agent learning environments that standardizes how agents practice and improve. The protocol aims…
Headroom, an open-source local context compression layer, reduces LLM prompt sizes by 60-95% while preserving accuracy, solving cost and latency issues from massive context windows. Its multi-modal pi…
A developer building a fashion-discovery feed found that off-the-shelf AI image detectors perform poorly on custom data, achieving only 0.68 AUC versus a simple linear model's 0.82. The article argues…
Microsoft rolled out AI Citation Share and other features in Bing Webmaster Tools, while new data from Google and Ahrefs cast doubt on the effectiveness of llms.txt for AI search visibility. Google, M…
Liquid AI released two new retrieval models, LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M, for fast multilingual search across 11 languages. The models, available on Hugging Face, are designed as dro…
Researchers at AIS2C2 2025 presented a prompt-driven tool-calling framework enabling lightweight open-source LLMs to perform complex multi-step tasks, reducing reliance on large proprietary models and…
Researchers introduced SEVRA, a serving-layer controller that decides when a frozen reasoning model should verify its answer instead of thinking longer, finding that selective verification improves ac…
DeepSeek AI released preview versions of its DeepSeek-V4 series, including two Mixture-of-Experts language models with up to 1.6 trillion parameters and support for one-million-token contexts. The mod…
Hugging Face released ML-Intern, an open-source agent that automates the machine learning research loop from literature review to training runs, on June 18. The tool has been used over 12,000 times si…
Salesforce released a tutorial demonstrating an end-to-end workflow for its CodeGen model, covering loading from Hugging Face, generating Python functions from natural-language prompts, and adding val…
AMD Radeon GPUs, both integrated and discrete, now support running large language models locally through open-source tools like Lemonade, LM Studio, Ollama, and llama.cpp. A new guide provides step-by…
LectuLibre developed a chunking strategy to translate entire books using large language models while preserving narrative coherence. The pipeline parses documents into logical units like chapters, spl…
Mass General Brigham researchers developed BRIDGE, a multilingual benchmark that evaluates large language models on real-world clinical tasks, revealing significant gaps between AI performance on medi…
Researchers at Sina Weibo released VibeThinker-3B, a 3-billion-parameter language model that matches or exceeds the reasoning performance of much larger systems from Google DeepMind, OpenAI, Anthropic…
Co/Core launches an AI cooperative where members share their own compute hardware to run AI inference jobs, using an open standard and public receipts to ensure transparency. The platform supports exi…
Eleven companies including Google, Microsoft, GitHub, and Hugging Face published the Agentic Resource Discovery (ARD) open specification on June 17, enabling AI agents to dynamically discover and veri…