cd/entity/Triton· home entities Triton
grep -l @triton /news/*.json | wc -l → 54

Triton

mentions 54 type Organization page 1/3 feed RSS

// recent coverage 54 mentions

06:22
2026-08-24
frontierroles.com
machine-learning

ML Systems Performance Engineer (MFU) — Higgsfield

Higgsfield AI, a generative AI company with $500M in annual revenue run rate and 25M+ users, is hiring an ML Systems Performance Engineer (MFU) for its Almaty, Kazakhstan office. The role focuses on o…

10:44
2026-08-20
promptcube3.com
machine-learning

Pine AI tops τ³-Voice leaderboard at 75.

Pine AI's 1.2B-parameter speech recognition model tops the τ³-Voice leaderboard with a score of 75, achieving significant gains on accented speech, medical dictation, and code-switching, but only marg…

21:01
2026-08-18
pub.towardsai.net
artificial-intelligence

What a Kernel Is, and Why Everyone Is Writing New Ones

A kernel is a single function that runs on a GPU, and the steep memory hierarchy—with a 1,600x gap between L2 cache and HBM—makes kernel design crucial for AI performance. FlashAttention, developed by…

14:30
2026-08-18
hiraditya.github.io
artificial-intelligence

The KV Cache Has No ABI

The KV cache has no standard ABI, with vLLM's FlashAttention backend alone reporting its cache shape as a four-dimensional tensor that varies by backend, attention variant, and model family, complicat…

23:17
2026-08-15
byteiota.com
developer-tools

Triton 3.7 Plugin Extensions: Drop Your Fork Now

Triton 3.7's new plugin extension system lets GPU kernel developers load custom MLIR compiler passes at runtime as shared libraries, eliminating the need to maintain a Triton fork. Meta's TLX extensio…

13:08
2026-08-15
sourcefeed.dev
artificial-intelligence

What a 232x AI Kernel Speedup Actually Proves

A solo developer with no professional GPU background placed 12th of 183 in GPU MODE's batched QR decomposition contest, beating the cuSolver-backed baseline by 232x (419,000 µs down to 1,805 µs on NVI…

19:23
2026-08-11
arxiv.org
machine-learning

What Irregularity Costs: CUDA C++, Rust, and Triton

A new study comparing CUDA C++, Rust via NVIDIA's cuda-oxide, and Triton on a hash-blocked TSDF fusion kernel found that Triton is more than an order of magnitude slower than hand-written CUDA C++ on …

10:58
2026-08-04
discuss.huggingface.co
artificial-intelligence

Decay-Gated O(N) Causal Linear Attention with Fused Triton Kernel

A developer has open-sourced a Decay-Gated O(N) Causal Linear Attention architecture with fused Triton/CUDA kernels, aiming to bypass quadratic multi-head attention bottlenecks. The project includes a…

00:26
2026-08-04
promptcube3.com
artificial-intelligence

US vs China AI: the lead is basically gone

The US lead over China in AI has essentially disappeared, according to an analysis of model releases and deployment trends. Chinese models like DeepSeek's R1 and V3 and Qwen now match or beat US open-…

13:00
2026-07-29
hiraditya.github.io
machine-learning

When XLA Isn't Enough: Pallas, Mosaic, and Triton

JAX's Pallas kernel system, which lowers through Triton on GPU and Mosaic on TPU, lets developers write custom kernels when XLA's automatic fusion is insufficient for operations like flash attention, …

17:48
2026-07-24
promptcube3.com
artificial-intelligence

Jensen Huang's New X Account: Why It Matters for AI

NVIDIA CEO Jensen Huang's new X account signals a shift toward real-time, decentralized communication with the developer community, potentially accelerating feedback loops on AI hardware and software …

00:00
2026-07-24
blog.getutm.app
developer-tools

Bringup Notes: Building Triton

A developer building the Triton DirectX 11 driver for QEMU argues that AI tools amplify human developers rather than replace them, citing challenges such as sparse documentation, multi-boundary debugg…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics