cd/entity/GPU· home› entities› GPU
grep -l @gpu /news/*.json | wc -l → 116

GPU

mentions 116 type Organization page 5/6 feed RSS

// recent coverage 116 mentions

12:20
2026-06-15
dev.to
artificial-intelligence

The CPU Is Back in the Stack — and Nobody Budgeted for It

Agentic AI workloads are shifting the compute ratio, making CPUs the critical coordination substrate rather than a support component for GPUs. This inversion is exposed by current Xeon supply tightnes…

09:00
2026-06-15
infoworld.com
large-language-models

33 LLM metrics to watch closely

A comprehensive list of 33 metrics for evaluating large language models (LLMs) has been compiled, covering performance indicators such as time to first token, average tokens per second, throughput, er…

23:51
2026-06-14
neurophos.com
ai-chips

Neurophos OPU

Neurophos announced its Optical Processing Unit (OPU), a photonic AI chip achieving 0.47 ExaOPS at 300 TOPS/W, claiming 10,000x size reduction over traditional photonic tensor cores. The company says …

14:00
2026-06-08
blog.r-lopes.com
artificial-intelligence

Quorum Math And Cache TTLs Are The Same Conversation

Engineering teams face a convergence of latency budgets, quorum math, and cache TTLs as the same underlying constraint, with ML systems intensifying the squeeze by placing LLM-as-judge loops and predi…

08:53
2026-06-05
letsdatascience.com
artificial-intelligence

AI-Powers Worm Exploits Stolen Compute to Infect Mixed Devices

Researchers published a proof-of-concept AI-driven worm that embeds an open-weight LLM on compromised GPUs to autonomously scan, exploit, and propagate across Linux, Windows, and IoT devices. The worm…

00:31
2026-06-04
dev.to
ai-agents

Your Agent Has a Memory That Runs While You Sleep

A developer built a continuous AI agent memory system called `akm improve` that runs autonomously on local hardware, processing 14,189 memories across 48 scheduled runs in 24 hours with zero failures.…

00:00
2026-05-31
cefboud.com
large-language-models

Exploring Speculative Decoding: From Concept to Implementation

Speculative decoding optimizes LLM inference by using a cheap draft model to predict multiple tokens, which are then verified in a single forward pass of the target model, reducing memory-bandwidth bo…

23:18
2026-05-30
categoryvc.com
ai-chips

AI Hardware

Modern GPUs spend most of their time during AI inference waiting for data, as memory bandwidth cannot keep pace with compute throughput. This fundamental bottleneck has driven the AI hardware market, …

← prev page 5 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics