Nemotron 3 Ultra
NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open-weight MoE model with hybrid Mamba-2 and Transformer architecture, under the Linux Foundation's OpenMDW-1.1 license. The model requires s…
NVIDIA released Nemotron 3 Ultra, a 550-billion-parameter open-weight MoE model with hybrid Mamba-2 and Transformer architecture, under the Linux Foundation's OpenMDW-1.1 license. The model requires s…
Independent benchmarks of NVIDIA's DGX Spark desktop through mid-2026 show Qwen 3.5 27B as the most consistent all-rounder on the easy Ollama path, while GPT-OSS 120B pushes nearly 4x the throughput b…
AMD's Ryzen AI Max+ 395 processor powers several new desktop and workstation systems, including the Framework Desktop and GMKtec EVO-X2, offering 128GB unified memory and 256 GB/s bandwidth for local …
ZML launched LLMD, a free inference server that runs LLaMA, Gemma, Qwen, and Mistral models on NVIDIA, AMD, Google TPU, Intel, and Apple hardware from a single Docker image. Built in Zig and compiled …
NVIDIA GPUs deliver eight standout graphics and AI features including Ray Tracing, Path Tracing, DLSS, ShadowPlay, and RTX HDR, according to a July 9, 2026 BGR roundup. The features differentiate NVID…
Micron Technology committed $200 billion to US memory chip manufacturing and R&D through 2035, a $30 billion increase from prior pledges, aiming to produce 40% of its DRAM domestically. The expansion,…
Microsoft opened its Fairwater AI datacenter in Wisconsin, equipped with NVIDIA GB200 GPUs, as part of a $190 billion annual capex spree that has pressured its stock. The facility supports Azure and A…
SpaceXAI launched Grok 4.5, its first model trained with AI company Cursor, claiming it excels at coding, agentic tasks, and knowledge work. The model is now the default for the Grok Build coding agen…
Analyst estimates from BofA Global Research and Morgan Stanley project an NVIDIA Rubin Ultra rack costing nearly $21 million, with HBM4e memory alone accounting for about $1.5 million per rack. The fi…
NVIDIA, AMD, and Intel compete in the 2026 AI GPU market, with NVIDIA's Blackwell RTX 50-series, AMD's Radeon AI Pro R9700, and Intel's Arc Pro B70 targeting local LLM inference. VRAM capacity and mem…
SK Hynix's $28 billion US share sale was oversubscribed seven times, drawing $196 billion in demand, as Wall Street investors rush to gain exposure to the AI infrastructure boom. The South Korean memo…
NVIDIA and Microsoft are jointly developing the DGX Station for Windows, a system designed to run hundreds of AI agents securely within enterprise Windows environments, targeting Japan's industrial se…
NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE LLM that achieves 2.03x server throughput over Nemotron-3-Super at matched user throughput. The model uses iterative puzzle comp…
SK Hynix, the South Korean memory chipmaker, is preparing to list on the Nasdaq through American depositary receipts, targeting approximately $28 billion in proceeds. The offering is already more than…
XAI released Grok 4.5 on July 8, topping the SWE Marathon benchmark with a 29.0% resolution rate and surpassing Anthropic's Claude Opus 4.8. The model, trained on tens of thousands of NVIDIA GB300 GPU…
Hewlett Packard Enterprise reported fiscal Q2 2026 revenue of $10.68 billion, a 40% year-over-year increase, driven by AI infrastructure demand. The company's AI server backlog reached nearly $6 billi…
NVIDIA's next-generation Rosa CPU may use TSMC's A16 process with back-side power delivery for a 2028 data-center platform, according to TrendForce and Wccftech reports. The technology could improve p…
A Fortran developer proposes adding a RESIDENT locality specifier to DO CONCURRENT to prevent implicit offload copying in systems with separate memory address spaces, allowing selective host-only para…
Ollama, an AI platform that simplifies running open-source models locally, has reached 8.9 million developers and is used by 85% of the Fortune 500. The company raised $65 million in Series B funding …
NVIDIA's closed-source cuda-checkpoint tool, which freezes and restores CUDA processes, suffers from slow PCIe transfers. Engineers reverse-engineered the tool to understand why checkpoint transfers f…