Release 4.0.0 · HuggingFace/Transformers.js
HuggingFace released Transformers.js v4, a major update featuring a new WebGPU backend rewritten in C++ for faster AI model inference in browsers, Node, Bun, and Deno. The release adds support for lar…
HuggingFace released Transformers.js v4, a major update featuring a new WebGPU backend rewritten in C++ for faster AI model inference in browsers, Node, Bun, and Deno. The release adds support for lar…
Researchers from Sina Weibo Inc (China) released VibeThinker-3B, a 3-billion-parameter dense reasoning model built on Qwen2.5-Coder-3B using the Spectrum-to-Signal post-training pipeline. The open-sou…
A developer fixed a broken fallback chain in their OpenClaw agent that was causing request timeouts during peak hours. The new chain includes seven entries: two local Ollama models, three OpenRouter f…
NVIDIA introduced advanced fused MLP kernels for mixture-of-experts (MoE) models, built with the CuTe DSL, delivering 1.3x–2x kernel-level speedups and enabling sync-free MoE execution. The optimizati…
Microsoft AI released a detailed technical report on the development of its first model, MAI-Thinking-1, emphasizing a controlled, reproducible training process built on human-generated data and propr…
The article describes the author's successful setup for running local AI models on an M4 Mac with 24GB of memory, specifically highlighting Qwen 3.5-9B (Q4 quantized) as the best performing model at ~…