First Principles of Model Routing
Model routing principles emphasize keeping models distinct, maintaining a small model pool, and using relative real-world benchmarks to optimize routing decisions between AI models based on speed, quality, and cost.
AI Infrastructure news and analysis on Web Pulse: 23076 curated articles tracking the latest AI Infrastructure developments, tools, and research, updated continuously from vetted sources.
Model routing principles emphasize keeping models distinct, maintaining a small model pool, and using relative real-world benchmarks to optimize routing decisions between AI models based on speed, quality, and cost.
Kioxia Holdings has begun shipping samples of its next-generation 332-layer 3D flash memory chips to AI data center operators, aiming to capture a larger share of the growing market. The new chips offer 59% more storage …
SAP is limiting hiring and travel spending to fund its AI transformation, shaking up executive ranks, and pushing customers toward cloud migrations with AI incentives. The company unveiled its "Autonomous Enterprise" vis…
Frontier large language models should be regulated as common carriers, allowing customers to use outputs for any lawful purpose including model distillation, argues a recent analysis. The proposal compares LLMs to teleph…
Yann LeCun's AMI Labs raised $1.03 billion at a $3.5 billion valuation to build AI systems that understand the physical world, arguing that current large language models cannot achieve human-level intelligence. The Paris…
OpenAI released the GPT-OSS-20B model on Hugging Face under the Apache-2.0 license, a 38.5 GB text-generation model with over 7 million downloads. The model requires large hardware such as multi-GPU or high-RAM machines …
Digi International launched DANI, an AI network operations agent embedded in its Digi Remote Manager platform, to automate network diagnostics and device management. The product was among the week's notable infosec relea…
Meta AI chief Alexandr Wang told employees in an internal town hall that the company's unreleased model, codenamed Watermelon, has matched OpenAI's GPT-5.5 on key benchmarks, giving CEO Mark Zuckerberg a talking point fo…
Microsoft launched the Microsoft Frontier Company, a new business focused on AI engineering, with a $2.5 billion investment and 6,000 experts to co-design and deploy AI systems for customers. The company aims to help cus…
Microsoft and Amazon Web Services announced multibillion-dollar investments in Forward Deployed Engineer (FDE) services, embedding thousands of AI experts directly into customer teams to help build and deploy AI systems.…
A federal lawsuit has been filed against Elon Musk's xAI over its Colossus supercomputer in Memphis, Tennessee, where residents of the predominantly Black Boxtown neighborhood report noise and air pollution from up to 35…
Anthropic is in early-stage talks with Samsung Electronics to manufacture a custom AI accelerator chip, The Information reported on July 2, 2026. The move follows OpenAI's custom chip and reflects AI labs' push to reduce…
Tesla will cap employee AI spending at $200 per week starting July 6, according to an internal memo. The move comes months after the company urged workers to use AI more aggressively, highlighting challenges in managing …
Vercel Chief of Software Andrew Qu argues that AI agents represent a new form of software distinct from traditional web applications, citing their dynamic interaction and unpredictable outputs. Qu led the development of …
Vercel launched Agent Runs observability in its MCP and CLI for eve, the open-source agent framework, enabling developers to inspect traces including reasoning, tool calls, and token usage. The new tools allow finding pr…
A comprehensive analysis of 15+ large language model quantization methods categorizes them into four paradigms: CPU-optimized GGUF-based approaches, GPU-native weight-only kernels, NVIDIA's floating-point hierarchy, and …
AMD published a guide demonstrating a GPU-resident YOLO26 object detection pipeline on the Radeon AI PRO R9700 GPU using ROCm 7.2. The pipeline keeps video frames in VRAM from decode to bounding boxes by chaining the VCN…
AMD announced the acceleration of large-scale LLM inference on its Instinct MI350X/MI355X GPUs using Eagle3 speculative decoding and AMD Quark quantization. The work enables lossless inference speedups for models like Ki…
Anthropic's Claude Sonnet 5 offers a 1-million-token context window at standard pricing with no long-context tax, but the cost of filling it remains significant. Whole-repo Q&A with caching costs about $1.17 for eight qu…
Gartner predicts 40% of enterprise applications will feature task-specific AI agents by 2026, up from under 5% in 2025, while over 40% of agentic AI projects are expected to be canceled by end of 2027 due to inadequate d…