Running Claude Code via NVIDIA NIM Proxy
A developer created a local proxy server that routes Claude Code traffic to NVIDIA NIM, enabling use of NVIDIA's AI models through the Claude Code interface. The proxy tool, available on GitHub, requi…
A developer created a local proxy server that routes Claude Code traffic to NVIDIA NIM, enabling use of NVIDIA's AI models through the Claude Code interface. The proxy tool, available on GitHub, requi…
Hippocratic AI partnered with Modular to integrate the MAX framework into its inference pipelines, achieving sub-500ms mean time to first token and approximately 30% faster P99 end-to-end latency for …
Co-packaged optics (CPO), a technology that integrates optical components directly with switch ASICs or AI accelerators to reduce power consumption and increase bandwidth density for AI data centers. …
Cerebras completed its IPO this week, with shares closing at $280 for a $60 billion market capitalization, following a pulled S-1 filing and a major partnership with OpenAI. The company’s CFO Bob Komi…
NVIDIA's Vera Rubin platform, powered by the Groq 3 LPX accelerator, addresses the scale-up challenges of agentic AI by delivering deterministic, low-latency execution across thousands of chips for tr…
NVIDIA released a new version of its Metropolis Blueprint for video search and summarization (VSS) that transforms live and recorded video into searchable, actionable intelligence using AI agents and …
NVIDIA and Ineffable Intelligence, the London-based AI lab founded by AlphaGo architect David Silver, announced a collaboration to build infrastructure for large-scale reinforcement learning. The part…
Nous Research released Hermes Agent, an open-source AI agent framework that crossed 140,000 GitHub stars and became the most-used agent on OpenRouter, designed for self-improvement and reliability on …
NVIDIA and SAP announced an expanded collaboration at SAP Sapphire to embed NVIDIA's open-source OpenShell runtime into the SAP Business AI Platform, providing security and governance controls for spe…
Modular CEO and Google AI infrastructure veteran Chris Lattner launched Inkwell, a real-time illustrated storybook app built on the company's Modular Cloud inference platform, demonstrating sub-second…
Convergent infrastructure requirements for the foundation model lifecycle—including tightly coupled accelerator compute, high-bandwidth networking, and distributed storage—and highlights the growing r…
Ollama, an AI model server, ran on a MacBook with no NVIDIA GPU by using GTAP software to intercept CUDA calls and forward them to a remote DGX Spark workstation with a 128 GB Blackwell GPU. The setup…
NVIDIA CEO Jensen Huang told Carnegie Mellon University graduates Sunday that they are entering the workforce at the start of the AI revolution, calling it a "once-in-a-generation opportunity to reind…
Researchers from NVIDIA have developed a new sparse data format and custom GPU kernels, called TwELL, that reshape unstructured sparsity in transformer language models to align with GPU architecture, …
Infobip Shift 2026 will return to Zadar this September, featuring speakers from Apple and NVIDIA alongside a program focused on developer careers and responsible AI. The conference, which drew 5,500 a…
U.S. Energy Secretary Chris Wright and NVIDIA Vice President Ian Buck argued Thursday at the SCSP AI+ Expo that American leadership in artificial intelligence depends on expanding domestic energy prod…
NVIDIA's Spectrum-X Ethernet fabric, now featuring Multipath Reliable Connection (MRC) technology, has been deployed by OpenAI, Microsoft, and Oracle to set a new standard for gigascale AI networking.…
NVIDIA and ServiceNow announced an expanded partnership at ServiceNow Knowledge 2026 to deliver specialized autonomous AI agents for enterprise environments. The collaboration introduces Project Arc, …
Researchers analyzing weight files from over a dozen open-weight language models — spanning labs including Qwen, DeepSeek, Google, and OpenAI, from 0.6 billion to 1.4 trillion parameters — found that …
NVIDIA CEO Jensen Huang declared at GTC 2026 that agentic AI systems are shifting infrastructure priorities from training to inference, as inference now accounts for 80-90% of the total lifetime cost …