2026-08
In August 2026, TTVL released Refmode CLI, a macOS command-line tool for switching Apple display reference modes, and shipped a large redesign of ttvl.co while rebuilding Hypercontext, a shared context layer for people, …
AI Infrastructure news and analysis on Web Pulse: 35533 curated articles tracking the latest AI Infrastructure developments, tools, and research, updated continuously from vetted sources.
In August 2026, TTVL released Refmode CLI, a macOS command-line tool for switching Apple display reference modes, and shipped a large redesign of ttvl.co while rebuilding Hypercontext, a shared context layer for people, …
GeoTessera 0.10 has been released, moving hosting to Source Cooperative with UTM-native Zarr routing and Matryoshka reads, while the full global run of Tessera v1.1 embeddings for 2017-2025 has been delayed by a month du…
Google Ads + Meta Ads + GA4 MCP, an MCP server integrating Google Ads, Meta Ads, and Google Analytics 4, released version 1.0.0 to the official MCP registry on 2026-08-29, offering over 250 tools for campaigns, creatives…
The MCP registry index mcpindex.ai lists io.github.epistemedeus/x402-data-gateway v1.23.28, published to the official MCP registry on 2026-08-29, as a hosted remote over streamable-http offering 22 paid x402 and MPP tool…
A developer reports that Azure OpenAI's rate-limit counter for gpt-5-mini deployments decrements by exactly the input token count, not by the documented formulas that include max_tokens. In 29 paired samples, 20 charges …
LightNow MCP Proxy (io.github.lightnow-ai/lightnow-proxy) version 1.9.1 was published to the official MCP registry on 2026-08-29, distributed as lightnow-proxy on PyPI. The mcpindex gate pins each MCP server's tool contr…
OpenContext has released a Model Context Protocol (MCP) server that gives AI coding agents persistent, project-local memory by reading and writing markdown files in a .opencontext/ directory, with tools save_context and …
A growing number of livestreams are using AI pipelines that combine LLM agents, prompt engineering, diffusion models, and text-to-speech to generate endless, viewer-driven 'slop' content in real time, according to a tech…
X Global Government Affairs announced Thursday that it identified around 200,000 inauthentic accounts, including 200 focused on swaying AI infrastructure debates, suspected of originating from China and amplifying argume…
OpenRouter retired monitoring of the deepseek/deepseek-v4-flash-0731 endpoint at akashml/fp8 on 2026-09-11 after it stopped yielding usable samples, leaving its B3IT tracking status retired and stalled. The endpoint's B3…
Monitoring of the z-ai/glm-5.2 endpoint at gmicloud/fp8 was retired as of 2026-09-04 because a single query costs more than the per-query guard allows, according to the tracking dashboard. The endpoint had been tracked w…
A developer built EidosStack AI Gateway, an OpenAI-compatible API gateway that provides unified access to 48+ models from providers including DeepSeek, Anthropic, OpenAI, Google, Alibaba, Moonshot, and Zhipu. The gateway…
VLLM's Decode Context Parallelism (DCP) enables higher concurrency and throughput for long-context agentic workloads by splitting the KV cache across GPUs, according to a new blog post from the vLLM team. In tests on a s…
Deploying LLM-powered vision agents in chaotic school zones remains unreliable due to latency and reasoning gaps, according to a technical analysis. The article proposes a tiered architecture using on-device TinyML, edge…
As of late August 2026, the llama.cpp and vLLM projects are advancing rapidly, but community members report that the most recent information on tensor-parallel inference for Intel Arc GPUs remains a Reddit post from earl…
A round-robin load balancer across three Azure OpenAI resources caused a 400 error on the Responses API because the API is stateful and requires requests to hit the same resource that created the items, with one failed c…
An engineer built TrueSRE, an autonomous multi-agent platform for Site Reliability Engineering that diagnoses and remediates production outages in under 25 seconds. Built on the TrueForge Agent Harness, FastMCP, and Dayt…
A developer has published a public, implementation-independent operating model for building and governing AI-assisted systems, emphasizing explicit authority, evidence-based action, bounded mutation, and recoverability. …
Fal engineer Rehan Sheikh launched a livestream on August 29, 2026, using H3 Max, fal's optimized version of the MiniMax H3 video generation model, which produces 15 seconds of AI video every 9 seconds, faster than real-…
Epoch AI published a case study on Anthropic to examine whether financing could bottleneck AI compute. The analysis suggests that while capital availability is a growing constraint, it is not yet the primary limiting fac…