cd/sources/hiraditya-auto-discovered· home› sources› Hiraditya (auto-discovered)
cat /sources/hiraditya-auto-discovered.feed | wc -l → 24

Hiraditya (auto-discovered)

articles 24 domain hiraditya.github.io → page 1/2 feed RSS
13:00
2026-08-26
hiraditya.github.io
ai-infrastructure

Heterogeneity Moved Inside the Chip

At Hot Chips 2026, OpenAI presented Jalapeño, an inference ASIC built with Broadcom, which integrates heterogeneous compute, memory, and network resources on a single chip to handle the three phases o…

07:00
2026-08-25
hiraditya.github.io
artificial-intelligence

The Illusion of Determinism in Disaggregated Inference

SGLang and vLLM, the two major LLM serving engines, have attempted to enforce deterministic inference but face fundamental challenges due to floating-point non-associativity in GEMM kernels, which cau…

07:00
2026-08-24
hiraditya.github.io
artificial-intelligence

A Bug Is a Violation of a Specification

A bug is a violation of a specification, and no specification exists that prefix caching's variable logits violate, according to an analysis of vLLM and SGLang issues. The vLLM PR #34046 adds an opt-i…

14:30
2026-08-19
hiraditya.github.io
artificial-intelligence

Two Schedulers, One SLO

A vLLM RFC from the llm-d team warns that disaggregated inference deployments, where prefill and decode run on separate schedulers, can trigger recomputation-based preemption inside the decode instanc…

14:30
2026-08-18
hiraditya.github.io
artificial-intelligence

The KV Cache Has No ABI

The KV cache has no standard ABI, with vLLM's FlashAttention backend alone reporting its cache shape as a four-dimensional tensor that varies by backend, attention variant, and model family, complicat…

03:00
2026-08-17
hiraditya.github.io
ai-infrastructure

Prefill and Decode Want Different Computers

AWS is pairing its Trainium chips with Cerebras CS-3 systems to split transformer inference into prefill and decode phases, with Trainium handling prefill and Cerebras handling decode, shipping as a p…

14:00
2026-08-14
hiraditya.github.io
artificial-intelligence

Provenance Is Not Correctness

Anthropic has begun watermarking text output from Claude, embedding imperceptible watermarks in text and attaching C2PA metadata to files, with models launched on or after 2 August 2026 supporting mar…

13:00
2026-07-29
hiraditya.github.io
machine-learning

When XLA Isn't Enough: Pallas, Mosaic, and Triton

JAX's Pallas kernel system, which lowers through Triton on GPU and Mosaic on TPU, lets developers write custom kernels when XLA's automatic fusion is insufficient for operations like flash attention, …

19:00
2026-07-28
hiraditya.github.io
machine-learning

What JAX Traces, and What It Refuses

JAX, a numerical computing library, is fundamentally a tracing machine that converts pure Python functions into a typed intermediate representation called a jaxpr, refusing to execute code with contro…

19:00
2026-07-25
hiraditya.github.io
machine-learning

XLA Up Close: What It Optimizes, and What It Won't

XLA, the compiler behind JAX, TensorFlow, and PyTorch/XLA, optimizes array programs by freezing shapes, statically allocating buffers, and fusing operations against a global cost model, which makes it…

03:00
2026-07-24
hiraditya.github.io
machine-learning

A Tour of XLA: Where MLIR Lives (and Where It Doesn't)

XLA, the compiler under JAX, TensorFlow, and PyTorch/XLA, uses two intermediate representations: classic HLO (a hand-built C++ IR) and MLIR dialects such as StableHLO and CHLO, with a translation laye…

19:00
2026-07-21
hiraditya.github.io
machine-learning

How torch.compile Actually Works

Torch.compile is not a traditional compiler but a system that intercepts Python bytecode at runtime, extracts compilable regions through a multi-stage pipeline, and stitches them back with eager Pytho…

19:00
2026-07-19
hiraditya.github.io
artificial-intelligence

Triton: The Compiler That Pretends to Be a Library

Triton is a compiler with a Python frontend that parses a function's AST, runs it through an MLIR pipeline, and emits a GPU binary, never executing the Python function as Python. The compiler handles …

19:00
2026-07-14
hiraditya.github.io
developer-tools

Anatomy of a CUDA Binary

A cubin (CUDA binary) is a standard ELF64 file with NVIDIA-specific sections that encode everything needed to load and launch a kernel, including machine code, parameter layout, and register allocatio…

19:00
2026-07-13
hiraditya.github.io
developer-tools

Where Should Your Code Live?

A survey of code repository hosting in 2026 finds that GitHub remains the default for most developers, but concerns over DMCA compliance, AI training on code, account suspension, and Microsoft ownersh…

13:00
2026-06-24
hiraditya.github.io
machine-learning

Three Ways to Take a Gradient: Tape, Trace, and Source Transform

PyTorch, JAX, and compiler-level tools like Enzyme use three different representations for automatic differentiation—runtime tape, functional trace-and-transform, and source (IR) transformation—each w…

12:00
2026-06-23
hiraditya.github.io
machine-learning

A tour of MLIR: The Dialect Stack Everyone Depends On

MLIR, a compiler infrastructure framework, has become the foundation for numerous machine learning compilers including XLA, Triton, Mojo, Torch-MLIR, IREE, and ONNX-MLIR. It provides a reusable IR con…

15:00
2026-06-21
hiraditya.github.io
machine-learning

Why VLIW Architecture is Popular Again

VLIW architecture is experiencing a resurgence driven by machine learning workloads, which rely on regular, statically analyzable computation that VLIW exploits efficiently. Unlike out-of-order design…

page 1 / 2 next →