cd/entity/CUDA· home entities CUDA
grep -l @cuda /news/*.json | wc -l → 318

CUDA

mentions 318 type Organization page 2/16 feed RSS

// recent coverage 318 mentions

14:00
2026-08-31
kdnuggets.com
large-language-models

Speed Up LLM Inference with DSpark Speculative Decoding

DeepSeek's DSpark speculative decoding technique, which combines parallel drafting with a lightweight sequential component, can improve local LLM generation speed on the same GPU, with DeepSeek report…

17:20
2026-08-30
cryptobriefing.com
artificial-intelligence

Nvidia’s RTX Spark seen as direct challenge to Apple in local AI

Nvidia launched the RTX Spark Arm-based superchip at Computex Taipei, targeting high-end laptops priced between $3,000 and $4,000 and capable of running 120-billion-parameter large language models loc…

11:00
2026-08-30
promptcube3.com
ai-infrastructure

Nvidia is winning the AI war by selling entire ecosystems

Nvidia is winning the AI war by selling entire ecosystems rather than just chips, leveraging CUDA software gravity, TensorRT, NeMo, and Mellanox-owned InfiniBand networking to lock developers into its…

16:48
2026-08-29
dev.to
developer-tools

I Tried Getting Closer to the GPU With Triton

A developer explores Triton, a Python-based language and compiler for writing GPU kernels, to understand low-level GPU programming and performance. The developer explains how Triton abstracts GPU thre…

03:13
2026-08-29
forum.level1techs.com
artificial-intelligence

MoE with little models

A forum user comparing CUDA and ROCm for local AI inference reports that ROCm issues have diminished and is considering an all-AMD build with dual R9700 GPUs by 2027, citing AMD's lower cost. Another …

21:26
2026-08-28
twitter.com
artificial-intelligence

I have made an OS in CUDA

A developer claims to have built a complete operating system, AOTX-1 (Ahead of Time eXecutive One), entirely in CUDA, with the kernel written in raw PTX. The OS includes networking (TCP/IP), graphics,…

19:07
2026-08-28
dev.to
machine-learning

GPU Architecture for ML (CUDA basics)

An engineer's blog post explains GPU architecture for machine learning, focusing on NVIDIA's CUDA platform and the parallel processing capabilities of GPUs. It breaks down key components like Streamin…

23:53
2026-08-27
promptcube3.com
artificial-intelligence

Jensen Huang thinks we already hit AGI and it's basically

Nvidia CEO Jensen Huang said that for a vast array of specific, high-level tasks, Nvidia's ecosystem has already reached AGI-level capability, arguing that the term AGI lacks scientific rigor and has …

04:00
2026-08-27
arxiv.org
large-language-models

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

A new benchmark, DataKernelBench, evaluates whether large language models can optimize database queries on GPUs, achieving up to 2.11x speedup over torch.compile on TPC-H SF10 with an H100 GPU. The be…

17:08
2026-08-26
promptcube3.com
ai-infrastructure

Nvidia is basically the only thing keeping the entire AI market

Nvidia Corp. is the primary driver of the AI market, with demand for its H100 and upcoming Blackwell architecture remaining insatiable, but the market is shifting from hardware demand to ROI reality, …

17:06
2026-08-26
forum.level1techs.com
large-language-models

FreeToken : An LLM Engine to max the bandwidth of All-The-Things

FreeToken, an LLM engine designed to maximize bandwidth across all hardware, faces a fundamental challenge: bit-exact reproducibility is impossible across different CPU and GPU architectures due to di…

← prev page 2 / 16 next →
// co-occurs with top 8 entities
// topics top 6 topics