cd/entity/GPU· home› entities› GPU
grep -l @gpu /news/*.json | wc -l → 116

GPU

mentions 116 type Organization page 4/6 feed RSS

// recent coverage 116 mentions

19:04
2026-06-21
devclubhouse.com
ai-chips

TPU vs GPU: The Architecture and Software Trade-offs

Google's TPU uses a systolic array architecture optimized for tensor algebra, offering higher throughput and energy efficiency than GPUs for dense matrix operations, but requires XLA compilation and i…

14:39
2026-06-19
letsdatascience.com
large-language-models

DigitalOcean Demonstrates LLM Compression with SparseGPT

DigitalOcean published a tutorial on June 19 demonstrating how to compress large language models using SparseGPT and Wanda pruning methods for GPU cloud deployment, targeting reduced inference costs a…

00:18
2026-06-19
dev.to
ai-agents

How I Run a 50-Agent AI Workforce on a Single 6GB GPU

A developer describes running ~50 local AI agents on a single 6GB GPU by using a lock-based queue, an eviction monitor, a resource governor, and a model router. The system serializes GPU access so onl…

00:00
2026-06-19
fergusfinn.com
ai-infrastructure

InfiniBand, RoCE, and all that

InfiniBand, a high-performance interconnect technology designed for Remote Direct Memory Access (RDMA), has become critical for AI training and inference workloads that require direct data movement be…

13:24
2026-06-17
letsdatascience.com
robotics

Nvidia showcases robots installing GPUs autonomously

Nvidia demonstrated an agentic robotics system called ENPIRE that taught fleets of robots to perform dexterous tasks, including handing a GPU to another arm and positioning it over a motherboard. The …

22:12
2026-06-16
0mean1sigma.com
machine-learning

2678x Faster Matrix Multiplication with a GPU

A developer achieved 2678x faster matrix multiplication using a GPU with CUDA, demonstrating how parallel processing on thousands of GPU cores reduces the O(N³) complexity of sequential matrix multipl…

18:57
2026-06-16
injuly.in
large-language-models

Inference cost at scale with napkin math

A technical analysis calculates the dollar cost per user for serving large language models at scale using napkin math, breaking down GPU resources, matrix multiplication costs, and attention mechanism…

20:40
2026-06-15
gilesthomas.com
machine-learning

Jax: Commitment Issues

JAX's default_device context manager places arrays on the specified device but does not commit them, allowing JAX to move them to other devices. This caused array lookups to take over a second by trig…

← prev page 4 / 6 next →
// co-occurs with top 8 entities
// topics top 6 topics