Qwen 3.8 Omni Flash
Alibaba's Qwen released Qwen 3.8 Omni Flash, a 3.8-billion-parameter model that natively integrates text, vision, and audio processing in a single low-latency architecture. The model is positioned to run fully local, rea…
AI Infrastructure news and analysis on Web Pulse: 34314 curated articles tracking the latest AI Infrastructure developments, tools, and research, updated continuously from vetted sources.
Alibaba's Qwen released Qwen 3.8 Omni Flash, a 3.8-billion-parameter model that natively integrates text, vision, and audio processing in a single low-latency architecture. The model is positioned to run fully local, rea…
Bend, a high-level programming language that compiles mathematically verified code to run natively and in parallel on both CPUs and GPUs without manual multi-threading, is pitched as a constraint layer for LLM agents gen…
ASUS has begun shipping the Ascent QN10, the first mini PC built on Qualcomm's Snapdragon X2 Elite platform and featuring an 80 TOPS NPU, according to gHacks. The launch pairs Qualcomm's new Snapdragon X2 Elite silicon w…
A developer building JAZZ, a Linux desktop environment for AI engineering, documented a series of failures encountered while developing the system, including filesystem corruption that broke the bootloader on a dual-boot…
Amazon called for "rigorous testing and strong safeguards" before new AI models are released, with an Amazon spokesperson saying "We don't see it as a choice between progress and safety," but the company stopped short of…
Qualcomm has patented a machine learning pipeline that trains a model on data from two separate network simulators to predict wireless network performance metrics — throughput, latency, and signal reliability — before ha…
Colibri is a new inference architecture that runs GLM-5.2, a 744-billion-parameter Mixture-of-Experts model, locally on 25GB of RAM by streaming expert weights from disk, since the model activates only about 40 billion p…
Huawei unveiled the Ascend 960 SuperPoD at its 2026 Connect conference, a cluster linking 15,488 Ascend AI processors across 220 cabinets in 2,200 square meters that operates as a single logical machine. Huawei rotating …
A tester ran Bonsai 2 27B, a ternary model that stores each weight as one of three values (1, 0, -1), locally on a 12GB Nvidia GPU using a special build of llama-cpp from PrismML's GitHub. The 7.2 GB model, based on Qwen…
A user request asks that the AMD Radeon AI PRO R9700 be added to the "My Hardware" database, citing its RDNA4 architecture, 32 GB GDDR6 memory, and 383 FP8 TFLOPS / 383 INT8 TOPS throughput as making it a fit for 30B–70B…
Salesforce and Google Cloud announced an expanded partnership at Dreamforce 2026 that runs Salesforce infrastructure on Google Cloud via Hyperforce and connects Salesforce's Model Context Protocol-based architecture to G…
Astra reportedly produced proofs for ten decades-old mathematics and theoretical computer science problems for only a few thousand dollars of compute, a result the piece frames less as a capability milestone than as an e…
PrismML released a ternary-weight version of Qwen3.8-27B under its Bonsai family on Tuesday, cutting the model from more than 50 GB of weights at full precision to just under 6 GB, which lets it run on an 8 GB PC graphic…
A practitioner building agent pipelines argues that agent-driven task management succeeds or fails on the orchestration layer rather than the interface, citing state drift rather than the LLM as the source of most failur…
Z.ai released GLM-5.3-FlashX, a native multimodal model that delivers inference speeds of up to 200 tokens/s, faster than GLM-5.3-Flash. The model uses a hybrid sparse and linear attention architecture to maintain accura…
Uber cut its cost per AI coding agent session by 52% from its June peak while weekly agent requests grew 9.4x and weekly active users grew 7x between February and August 2026, the company's engineering team reported. Mor…
A developer argues that running AI agents in production requires its own infrastructure discipline distinct from standard microservice practices, because agent requests break four core assumptions: bounded latency, const…
Microsoft Foundry Agent Service has adopted the Model Context Protocol (MCP) as a first-class remote tool type, introducing a Toolbox construct that acts as a governance layer for agent tool calling. The design separates…
A cost-analysis writeup argues that cache economics, not sticker price, determines real AI spending for small businesses, since repeated context tokens can be cached and billed at a fraction of full input rates. Using a …
Huawei Technologies rotating chairman Eric Xu Zhijun said at the Huawei Connect 2026 conference in Shanghai on Thursday that a major domestic shift toward Huawei's Ascend-based AI computing infrastructure for model train…