Meta Muse Spark Dashboard
SigNoz released a Meta Muse Spark dashboard that tracks cost, token spend, prompt cache effectiveness, latency, and finish reasons for Meta Model API calls, requiring SigNoz v0.135.0 or newer and the V2 dashboard schema.…
AI Infrastructure news and analysis on Web Pulse: 34996 curated articles tracking the latest AI Infrastructure developments, tools, and research, updated continuously from vetted sources.
SigNoz released a Meta Muse Spark dashboard that tracks cost, token spend, prompt cache effectiveness, latency, and finish reasons for Meta Model API calls, requiring SigNoz v0.135.0 or newer and the V2 dashboard schema.…
Session Replay provider Replay Vision built a rasterizer to convert rrweb session data into MP4 videos so multimodal AI models like Gemini can analyze user recordings, addressing the problem that models cannot understand…
Meta released Muse, a personal agent that runs inside a dedicated Secure VM with a host-side control component called Sentinel that independently evaluates connector policies and network requests, according to a blog pos…
The MCP server io.github.theoddden/stamp, version 0.5.1, was published to the official MCP registry on 2026-09-08 and is distributed as stamp-mcp on PyPI and as a hosted remote over streamable-http. The mcpindex.ai scree…
Inception Labs launched Mercury 2.5 on September 8th, a diffusion-based language model API that the company claims generates 1,107 tokens per second on widely available Nvidia GPUs, with a 260,000-token context window an…
OpenAI is hiring a Software Engineer for AI for Chip Design in San Francisco, offering a salary range of $266k–468k per year, which sits 61% above the $228k median for AI Agents roles in the United States. The role invol…
Tether CEO Paolo Ardoino, in an interview on Eye on AI, criticized Silicon Valley's gigawatt data center approach to AI and unveiled Tether's QVAC, a 4-billion-parameter edge model that he claims outperforms Meta's Gemma…
ServiceNow is hiring a Senior Staff AI Security Engineer in Santa Clara, California, offering $191k–334k/yr, a salary 11% above the $238k median for Core ML roles on the board. The role involves building ML-driven securi…
Perplexity has open-sourced Lily, a local inference runtime for Apple silicon that uses custom Metal kernels to run Qwen-based models, enabling faster processing of local files with fewer cloud tokens and better privacy.…
OpenRouter retired behavioral monitoring of the z-ai/glm-5.3-flash endpoint hosted at baseten/fp8 on 2026-09-17 after re-initialization repeatedly timed out following a detected change, leaving no logprob tracking becaus…
Agent identity for AI agents depends on four distinct layers — KYA (Know Your Agent), authentication, authorization, and payment authorization — with payment authorization remaining the least-solved layer, according to a…
Alif Semiconductor introduced two $49 development boards, the Ensemble E1C StartKit (SK-E1C) and the wireless Balletto B1 StartKit (SK-B1), on September 8th, pairing a 160 MHz Arm Cortex-M55 CPU with an Arm Ethos-U55 NPU…
A new arXiv paper submitted on 27 Aug 2026 systematizes the parallel and distributed systems foundations of training Reasoning Language Models (RLMs) such as DeepSeek-R1, o3, and Kimi k1.5, noting that state-of-the-art R…
Latitude, an AI observability platform, introduced a workflow for managing fleets of AI agents, demonstrated with a Hermes fleet of five client deployments and 845 conversations over six weeks. The system uses flaggers, …
The Model Context Protocol (MCP) specification offers two client registration options for remote MCP servers: pre-registration, akin to a dinner party where clients are known in advance, and automated registration, akin …
GoDaddy is hiring a Staff Software Engineer - Backend for its Machine Learning Engineering team, offering $182k–273k per year for a remote position in the United States. The role involves leading the design and evolution…
Prompt caching can cut LLM API bills by 50 to 90 percent, but only if prompts are structured correctly, according to Anthropic's pricing and case studies. Anthropic charges 1.25x standard input price for a 5-minute cache…
Crusoe, a vertically integrated AI infrastructure company, is hiring a Staff Applied AI Inference Engineer in San Francisco with a salary range of $215,000–260,000 per year, which sits 8% above the $219,000 median for ML…
OpenAI will shut down its Sora Videos API on September 24, 2026, affecting models sora-2, sora-2-pro, and associated snapshots, with all account data deleted and no replacement offered by OpenAI. The shutdown follows Sor…
OpenAI's GPT-6 Astra, its latest and most capable model, is now generally available on Amazon Bedrock, enabling organizations to run AI agents for coding, data analysis, and complex workflows at production scale. The mod…