Micron stock heats up again, crossing $1,000
Micron Technology Inc. (MU) shares crossed the $1,000 level in premarket trading on Monday, marking the first time since July 6, and are up 240% year to date. The rally is driven by strong demand for …
Micron Technology Inc. (MU) shares crossed the $1,000 level in premarket trading on Monday, marking the first time since July 6, and are up 240% year to date. The rally is driven by strong demand for …
Modular shipped Mojo 1.0 on August 12 as part of its 26.5 release, marking the first production-ready version of its Python-syntax AI hardware language, with changes including uniform `var` syntax, a …
Nvidia has pre-booked TSMC's 1.6-nanometer A16 process capacity for its second-half 2028 Feynman GPUs, locking up advanced packaging and silicon photonics to deny rivals the manufacturing capacity nee…
Linux 7.2 has been released, introducing Cache Aware Scheduling after more than a year of work, along with AMD Zen 6 prep, HDMI 2.1 Fixed Rate Link support, and initial support for 2023 MacBooks. The …
Linux 7.2, released by Linus Torvalds on August 16, introduces cache-aware load balancing for chiplet-based servers like AMD EPYC and Intel Xeon 6, but the feature is opt-in via CONFIG_SCHED_CACHE. Th…
AWS is pairing its Trainium chips with Cerebras CS-3 systems to split transformer inference into prefill and decode phases, with Trainium handling prefill and Cerebras handling decode, shipping as a p…
Anthropic's Claude Code can now run on-premises with AMD Instinct GPUs, serving GLM 5.2 at full quality via SGLang and LiteLLM, eliminating cloud API dependency and per-token costs. The setup uses an …
AMD has launched a multi-part study on memory instruction scheduling for lock-stepped kernels on its Instinct MI300X GPU, beginning with a tiled GEMM kernel example. The series aims to address bandwid…
Meta Platforms Inc. unveiled four new generations of its in-house Meta Training and Inference Accelerator (MTIA) chips on March 11, 2026, and expanded a partnership with Broadcom Inc. to co-develop th…
Unsloth, the open-source fine-tuning library, released Unsloth Desktop, a free beta app for macOS, Windows, and Linux that runs and trains LLMs, diffusion models, and audio models locally, directly co…
Amazon Web Services (AWS) has instructed its engineers to conserve CPU cycles at all costs as AI workloads, particularly agentic AI systems, strain its cloud infrastructure and cause wait times for CP…
Riot Platforms signed a $9 billion, 20-year compute agreement with Anthropic for 191 megawatts of capacity at its Rockdale, Texas facility, marking a strategic shift from Bitcoin mining to AI data-cen…
A daily AI news digest published on Aug 16 lists 30 stories, including reports that OpenAI is losing key personnel ahead of a potential IPO, Qwen 3.8 27B outperforms the larger Qwen 3.7 Plus in coding…
A developer contributed optimizations to llama.cpp for AMD's gfx1151 GPU, specifically targeting Qwen3.8 27B's SSM convolution input pattern. The patch introduces a fast LDS-transpose path for dim-0 c…
HuggingFace's State of Open Models report for summer 2026 reveals that Chinese labs now dominate frontier-scale open models, with the largest reaching 2.78 trillion parameters, while US labs peaked at…
AMD's Ryzen AI Halo may outperform NVIDIA's DGX Spark for local AI development, according to a hardware comparison. The Ryzen AI Halo offers better power efficiency and thermal performance, while the …
Alibaba released Qwen3.8-27B, a 27.78-billion-parameter open-weights multimodal model under Apache 2.0, with a 262,144-token native context window and configurable reasoning. Vendor-reported benchmark…
Semiconductor stocks are on pace for their best August since 2003, with the PHLX Semiconductor Sector Index up more than 10% in August 2026 after a 20%-plus decline in July, according to semiconsociet…
A technical article by Piyush Srivastava, Karnik Modi, Stephen Varela, and Rithish Ramesh examines LLM inference benchmarking, focusing on the interplay between latency, throughput, concurrency, and c…
French startup Kog claims its software can deliver 30x faster LLM inference on standard datacenter GPUs, demoing 3,000 tokens per second on a 2-billion-parameter model using AMD MI300X and NVIDIA H200…