Mana: 2-3 Seconds to Feeling Human
A developer has built Mana, a voice AI assistant that runs entirely on a local machine with an 8GB VRAM GPU, achieving 2-3 second response times by using a single 4B model for reasoning, code generati…
A developer has built Mana, a voice AI assistant that runs entirely on a local machine with an 8GB VRAM GPU, achieving 2-3 second response times by using a single 4B model for reasoning, code generati…
DeepSeek's V4 Flash and Qwen-3.8-Max have rewritten the cost equation for large language models, with Nous Research offering DSV4 Flash at a 90% discount and OpenCode reporting 8 trillion tokens proce…
A developer on tamiz.pro argues that traditional supply chain security tools are inadequate for AI agent workflows, proposing local LLMs and agent eval harnesses as the new defense pillars. The post h…
OpenAI's latest model, Sol, is the only OpenAI model in the top 10 of the coding arena leaderboard, while four open-weight Chinese models have achieved frontier-level performance, with one being affor…
OpenAI's GPT-5.6 Sol rewrote production GPU kernels, cutting serving costs by 20%, and an internal model Astra produced ten formally proved mathematical results, while OpenAI cut GPT-5.6 Luna API pric…
Alibaba's Qwen team released Qwen-MM-Plugins, an open-source repository that packages multimodal capabilities as installable skills for six agent harnesses: Claude Code, Codex, Qoder, OpenClaw, Qwen C…
New checkpoints for the Qwen 3.8B family have been detected on a model hub, suggesting additional instruction-tuned or specialized versions, though the Qwen team has not officially announced them. The…
Moonshot AI released Kimi K3 in July 2026 as a 2.8T-parameter open-weight MoE model with a 1M token context window and native vision support, scoring 88.3 on Terminal-Bench 2.1, close to top closed mo…
A post on r/LocalLLaMA by an Ant Group employee explains why Chinese AI labs release high-quality local models, highlighting four distinct strategies: Alibaba's Qwen focuses on distribution across mod…
In 2026, edge inference has advanced to the point where an 80B Qwen model runs in just 4.3GB of RAM on a MacBook Pro, and a 35B model runs on an iPhone 18 Pro. This is achieved through a combination o…
A new arXiv study (2608.00077v1) audits visual-token pruning in OCR-critical multimodal large language models (MLLMs), finding that accuracy alone misses failures where correct answers lack local toke…
The US lead over China in AI has essentially disappeared, according to an analysis of model releases and deployment trends. Chinese models like DeepSeek's R1 and V3 and Qwen now match or beat US open-…
Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters and a million-token context window, and demonstrated an autonomous 16-day …
Alibaba Group Holding Ltd.'s AI model, Qwen, autonomously coded for 16 consecutive days, with every commit publicly available on GitHub, according to a report by The New Stack. The AI, developed by Al…
DingTalk and Feishu have evolved into central hubs for enterprise AI, with DingTalk's AI Assistant Platform and Feishu's Intelligent Partner enabling conversational BI deployment directly within their…
Qwen 3.8 Max, a new API model, bills $2 per million input tokens and $6 per million output, but its cheapest reliable configuration is not the default. Testing through the Synthorai gateway revealed t…
Steve Berry warns that AI companies, like Uber before them, will raise prices once users are dependent, and credits China's Alibaba with releasing free open-source models like Qwen as a market check. …
Chinese-developed open-weight AI models accounted for 41% of all downloads on Hugging Face between February 2025 and February 2026, surpassing US-developed models at 36.5%, according to Hugging Face C…
Headroom, a new browser extension, provides a real-time visual gauge of context window usage in AI chat conversations, warning users before the AI forgets earlier information. It supports DeepSeek, Ch…
Alibaba's Qwen team unveiled Qwen3.8-Max, a 2.4-trillion-parameter language model with 95 billion active parameters, designed to autonomously complete complex tasks over extended periods. In testing, …