cd/entity/GPU· home› entities› GPU
grep -l @gpu /news/*.json | wc -l → 116

GPU

mentions 116 type Organization page 6/6 feed RSS

// recent coverage 116 mentions

03:52
2026-05-29
dev.to
neural-networks

Tensors Explained Part 2: Why Tensors Are Useful

Tensors enable hardware acceleration by leveraging GPUs and TPUs to perform parallel mathematical operations efficiently, making them essential for training neural networks. They also support automati…

19:50
2026-05-28
arxiv.org
artificial-intelligence

SIA: Self Improving AI with Harness and Weight Updates

Researchers have developed SIA, a self-improving AI system that updates both its own software scaffolding and internal model weights without human intervention, combining two previously separate appro…

18:39
2026-05-28
letsdatascience.com
machine-learning

MoE Transforms Open Model Ecosystem Costs

Mixture of Experts (MoE) models are reshaping the economics of open-model deployments by reducing GPU inference costs and altering serving stack requirements. The shift toward MoE architectures in 202…

11:53
2026-05-28
github.com
large-language-models

Why LLM decode is memory-bound, not compute-bound

LLM inference costs 100x more than traditional machine learning inference because autoregressive generation requires a separate forward pass through the entire model for each output token. A Llama 3.1…

12:51
2026-05-26
klongpy.org
machine-learning

KlongPy: PyTorch Back End and Autograd

KlongPy now supports a PyTorch backend that enables GPU acceleration and automatic differentiation for gradient-based computations. The torch backend outperforms NumPy by up to 8x on large arrays and …

16:35
2026-05-24
thedeepview.com
artificial-intelligence

How the compute crisis is defining the next stage of AI

Lambda Chief Commercial Officer Robert Brooks IV argued that computing power is becoming one of the most strategically important resources in the AI economy, with his company building supercomputers f…

15:38
2026-05-22
dwarkesh.com
ai-chips

Reiner Pope – Chip design from the bottom up

Reiner Pope, CEO of AI chip startup MatX and former Google engineer, delivered a blackboard lecture explaining chip design from basic logic gates to the architectures of GPUs, TPUs, FPGAs, and the hum…

04:54
2026-05-22
arxiv.org
machine-learning

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

CODA, a GPU kernel abstraction that reparameterizes memory-bound Transformer operations like normalization and activations to execute as GEMM-plus-epilogue programs, keeping data on-chip to reduce glo…

11:37
2026-05-21
dev.to
large-language-models

End-to-End Observability for vLLM and TGI: from DCGM to Tokens

Running large language model inference servers like vLLM and TGI in production requires specialized observability because they behave differently from standard web services, with key metrics like late…

18:18
2026-05-13
newsletter.semianalysis.com
ai-chips

Cerebras — Faster Tokens Please

Cerebras Systems has secured a 750MW compute deal with OpenAI, positioning the company for its upcoming IPO as demand for fast token generation surges. The wafer-scale chip maker's speed advantages, p…

18:56
2026-04-30
pytorch.org
large-language-models

SMG: The Case for Disaggregating CPU from GPU in LLM Serving

Shepherd Model Gateway (SMG) has disaggregated all CPU-bound workloads from GPU inference in large language model serving, moving tokenization, detokenization, and parsing into a dedicated Rust gatewa…

15:19
2023-02-25
alexselimov.com
open-source

Hosting your own git frontend service using Gitea

The article provides a step-by-step guide on self-hosting a Git frontend service using Gitea on a Debian server with Nginx. It covers setting up a PostgreSQL database for Gitea, downloading and instal…

← prev page 6 / 6
// co-occurs with top 8 entities
// topics top 6 topics