cd/entity/CUDA· home entities CUDA
grep -l @cuda /news/*.json | wc -l → 318

CUDA

mentions 318 type Organization page 7/16 feed RSS

// recent coverage 318 mentions

21:29
2026-07-30
wheels.astral.sh
developer-tools

Astral GPU Indexes

Astral has launched pre-built GPU-enabled Python wheels for the PyTorch ecosystem, supporting packages like FlashAttention across multiple CUDA versions (11.8 through 13.2) and PyTorch versions. The A…

21:20
2026-07-30
modal.com
artificial-intelligence

Host overhead is killing your inference efficiency

Host overhead, caused by the CPU blocking the GPU, is a major source of inefficiency in AI inference, leading to low GPU kernel utilization and doubling GPU costs when at 50%. Modal recommends using t…

09:00
2026-07-29
infoworld.com
developer-tools

Why open source matters in an AI world

Open source remains a critical strategy for tech companies in the AI era, argues a former Borland employee drawing on lessons from Borland's failed open-source move with InterBase. Nvidia leverages it…

13:01
2026-07-28
github.com
artificial-intelligence

Adaptive speculative decoding on a $300 GPU

A developer achieved up to 9.27× speedup on code editing tasks using adaptive speculative decoding on a €300 RTX 5060 GPU, with a 0.6B draft model nearly doubling math and JSON throughput. The project…

08:17
2026-07-28
probablydance.com
artificial-intelligence

If AI Writes All the Code, What Do the Programmers Do?

Eight months ago, software engineer Malte Skarupke wrote roughly 90% human code and 10% AI code; now his code is about 90% AI-generated. In a recent optimization of a matrix-multiply kernel using Nvid…

18:28
2026-07-27
promptcube3.com
artificial-intelligence

Kimi K3 Weights: Initial Deployment Notes

A developer deploying the Kimi K3 model encountered a CUDA out-of-memory error caused by KV cache allocation during initial inference passes, not the model weights themselves. The developer resolved t…

16:52
2026-07-27
i-programmer.info
artificial-intelligence

Programming Massively Parallel Processors, 5th Ed(Morgan Kaufmann)

The 5th edition of 'Programming Massively Parallel Processors' by Wen-mei W. Hwu, David B. Kirk, and Izzat El Hajj, published by Morgan Kaufmann, introduces new chapters on filtering, wavefront parall…

02:02
2026-07-27
promptcube3.com
artificial-intelligence

Gemma Model Deployment: Handling VRAM Spikes

A developer reports that deploying Google's Gemma model causes a CUDA out-of-memory error during the weights loading sequence, with 18.1GB already allocated on a 24GB GPU. The spike occurs during the …

11:48
2026-07-26
pub.towardsai.net
artificial-intelligence

How I Fit a Model That “Shouldn’t” Fit on a 6GB Laptop GPU

A self-study AI engineer demonstrates how to fit a large language model on a 6GB laptop GPU using quantization, a technique that reduces model precision to enable fine-tuning on consumer hardware. The…

23:03
2026-07-25
promptcube3.com
artificial-intelligence

AMD ISA: Why Machine-Readable Specs Change GPU Programming

AMD's move to provide machine-readable Instruction Set Architecture (ISA) specifications could transform GPU programming by enabling LLM agents to directly generate optimized kernels, bypassing the ne…

22:43
2026-07-25
github.com
developer-tools

Nvprobe – Open-source, zero-setup CLI for CUDA benchmarks

Nvprobe, an open-source, zero-setup CLI tool for CUDA benchmarks, has been released on GitHub. The tool automates CUDA workloads including HPL, HPCG, MLPerf inference, and custom kernels, and generate…

← prev page 7 / 16 next →
// co-occurs with top 8 entities
// topics top 6 topics