# Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton

> Source: <https://developer.nvidia.com/blog/deploying-an-hstu-generative-recommender-with-nvidia-dynamo-triton/>
> Published: 2026-09-30 20:54:59+00:00

Generative recommender (GR) systems are emerging as a powerful new direction for large-scale personalization. Instead of treating recommendation as a set of isolated retrieval, ranking, and prediction stages, GRs reformulate recommendation as sequence modeling over user behavior. A user’s interactions, context, candidate items, and actions become tokens in a high-cardinality event stream, and the model learns to generate or score the next relevant items from that sequence.

This approach is especially attractive for modern recommendation workloads, where user histories can be long, item catalogs are constantly changing, and personalization quality depends on modeling rich sequential behavior. But it also introduces a serving challenge: GR models need low-latency inference despite long histories, large embedding tables, and sequence-heavy model architectures.

[NVIDIA Dynamo-Triton](https://developer.nvidia.com/dynamo-triton) (formerly NVIDIA Triton Inference Server) now supports an end-to-end Hierarchical Sequential Transduction Unit (HSTU) GR inference workflow through the NVIDIA [`recsys-examples`](https://github.com/NVIDIA/recsys-examples/tree/v26.06.01/examples/hstu/inference_aoti) repository. The workflow combines [HTSUs](https://github.com/NVIDIA/recsys-examples/blob/main/examples/hstu/README.md), [PyTorch Ahead-of-Time Inductor compilation](https://github.com/triton-inference-server/pytorch_backend#pytorch-20-models), [FlexKV-backed KV caching](https://github.com/taco-project/FlexKV), [native C++ validation](https://github.com/NVIDIA/recsys-examples/blob/v26.06.02/examples/hstu/inference_aoti/README.md#exported-model-package), [NV embedding cache](https://github.com/NVIDIA/nv-embedding-cache), and [Dynamo-Triton deployment](https://github.com/triton-inference-server).

The result is a practical path for serving HSTU ranking models with strong latency performance. At dynamic batch size 8 on an [NVIDIA RTX PRO 6000 Blackwell Workstation Edition](https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/) GPU, Dynamo-Triton with PyTorch AOTI achieved best-case speedups of up to 4.47x for the three-layer HSTU model and 5.93x for the eight-layer model under a 100% GPU KV-cache hit rate, relative to the same AOTI configuration without KV caching.

This post shows you how to move an HSTU generative recommender from PyTorch development to production inference with NVIDIA Dynamo-Triton, PyTorch AOTI, and FlexKV. You will learn how to export and ahead-of-time compile the model, validate the resulting deployment artifact in Python and native C++, and serve it through Dynamo-Triton without rewriting the model for a separate runtime.

It also examines how GPU-backed KV caching reduces repeated computation and presents benchmarks demonstrating up to 5.93x lower latency, highlighting the practical performance benefits of this deployment workflow.

## Why use HSTUs for generative recommendation?

HSTUs were introduced for GR workloads that operate over high-cardinality, nonstationary event streams. In a traditional recommender system, retrieval and ranking are often built from a collection of specialized models and feature pipelines. GRs instead model recommendation as a sequential prediction problem, allowing the model to reason over user context, item history, action history, and candidate items in one sequence-aware architecture.

In the NVIDIA HSTU ranking example, the model input is built from categorical tokens. Contextual tokens represent user-side information, item tokens represent items, and optional action tokens represent user interactions with those items.

The HSTU preprocessing path retrieves embeddings, interleaves item and action embeddings when action tokens are present, appends contextual information, and applies positional encoding. HSTU blocks then process the sequence, and a prediction head produces multitask ranking outputs.

This structure is well-suited for recommendation systems where recency, order, and repeated interaction patterns matter. However, it also means inference can become expensive when each request repeatedly processes long historical sequences. Production systems need to preserve HSTU modeling benefits while reducing redundant computation during serving.

## Why is serving large sequential recommenders challenging?

Serving large sequential recommenders is different from serving a small dense ranking model. The serving stack must handle jagged sequence inputs, large categorical embedding state, long histories, and request patterns where the same user may return repeatedly with only a small amount of new information. Recomputing the full key-value state for a user history on every request wastes work and increases latency.

This is where KV caching becomes important. A KV cache stores reusable key-value data from prior sequence computation, allowing the model to avoid recomputing cached portions of the user history. For recommender inference, this is particularly useful when a user’s long-term history remains mostly stable while new candidate items or recent actions arrive.

The NVIDIA HSTU inference workflow includes a [`KVCacheManager`](https://github.com/NVIDIA/recsys-examples/tree/v26.06.01/corelib/recsys_kvcache_manager) that uses GPU memory and host storage for KV data caches. The GPU cache is organized as a paged KV data table and supports lookup, allocation, append, and eviction. When GPU cache space is constrained, older users can be evicted according to an LRU-style policy. Host-side storage provides another tier for cached KV data, and the workflow includes a FlexKV-backed backend for the KV-cache runtime.

The HSTU attention kernel can consume KV data from the paged cache, and the exported inference path includes cache-aware custom operations for lookup, allocation, onboarding, appending, and offloading. This allows the serving path to preserve the model’s sequence semantics while reducing redundant computation.

## PyTorch AOTI for native inference

The Pytorch AOTI (Ahead-of-Time Inductor) workflow starts from a PyTorch model and exports it using `torch.export` and PyTorch AOTI. AOTI compiles the model ahead of time into a package that can be loaded by a native C++ runtime. This reduces Python runtime overhead and provides a deployment-friendly artifact for the Dynamo-Triton PyTorch AOTI backend.

The exported model package contains the AOTI model archive plus metadata and embedding table files. In the NVIDIA example, the embedding implementation combines DynamicEmb inference embedding tables and [NV Embedding Cache](https://github.com/NVIDIA/nv-embedding-cache) that reduces GPU memory usage by storing only the popular embeddings in GPU memory while keeping the entire table in CPU memory. The export path writes layer metadata and embedding table data alongside the compiled `.pt2` archive so the model can be loaded without unnecessary duplicate embedding table copies.

The workflow validates the same exported artifacts in multiple ways. Python export scripts generate the package and replay tensors. Native C++ executables load and replay the exported model for correctness and performance validation. The Dynamo-Triton deployment then uses the same AOTI package and replay path, which helps keep development validation and production serving aligned.

## Dynamo-Triton deployment path

Dynamo-Triton provides the production serving layer for the exported HSTU model. The AOTI deployment uses the [Dynamo-Triton PyTorch backend](https://github.com/triton-inference-server/pytorch_backend#aot-inductor-support-beta) with `platform: "torch_aoti"`. This allows Dynamo-Triton to load and serve the ahead-of-time compiled PyTorch model package.

The full workflow includes the following five stages:

- Build the required custom operators and runtime libraries
- Export the HSTU ranking model with PyTorch AOTI
- Start the FlexKV-backed KV-cache service
- Validate the exported artifacts with native C++ replay
- Serve the exported KV-cache AOTI model with Dynamo-Triton

This approach is important because recommender serving requires more than just a fast model kernel. Dynamo-Triton brings model repository management, request handling, backend integration, metrics, and deployment structure. AOTI brings a lower-overhead compiled model artifact. [NV Embedding Cache](https://github.com/NVIDIA/nv-embedding-cache) lowers GPU memory requirements by keeping only the hot parts of embedding tables in GPU memory. [FlexKV](https://github.com/taco-project/FlexKV/tree/main) reduces recomputation for long user histories by caching attention blocks. Together these form a serving stack aimed at realistic generative recommender inference.

## Benchmarking HSTU serving latency

This benchmark compares HSTU serving latency across Dynamo-Triton backends, model sizes, batch sizes, and KV-cache states.

The goal is to quantify the performance benefits of the production HSTU serving stack. It compares the Dynamo-Triton PyTorch AOTI backend with the Python backend, measures the additional latency reduction from GPU KV caching, and evaluates how those benefits scale across model depths and batch sizes. Ultimately, it shows developers the performance they can expect when moving from uncached Python-based inference to compiled, cache-aware HSTU deployment with Dynamo-Triton.

The [benchmark results in](https://github.com/NVIDIA/recsys-examples/blob/v26.06.02/examples/hstu/inference_aoti/benchmark/README.md#benchmark-results) [`recsys-examples`](https://github.com/NVIDIA/recsys-examples/blob/v26.06.02/examples/hstu/inference_aoti/benchmark/README.md#benchmark-results) use the KuaiRand-1K ranking configuration on a single GPU. The model structure includes three-layer and eight-layer HSTU variants, hidden size 512, four attention heads, BF16 model weights, BF16 KV cache, maximum history sequence length of 8,192 in total (4,096 item plus action pairs) history stream, maximum candidate sequence length of 100, and six contextual features. The effective sequence length before alignment is 8,298 tokens, exported with a maximum aligned sequence length of 8,320.

The benchmark protocol reports latency per logical request. For the Dynamo-Triton AOTI benchmark, each Dynamo-Triton call contains one logical batch, and latency per logical request is calculated by dividing end-to-end pass time by the number of Dynamo-Triton calls times the logical batch size. Dataset loading, validation, rebatching, user-ID generation, server startup, warmup, and post-warmup sleep are excluded from the measured time.

The hardware used for the Dynamo-Triton backend comparison and batch-size results is an NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU.

### Dynamo-Triton backend comparison

At Dynamo-Triton batch size 2, PyTorch AOTI improves latency over the Dynamo-Triton Python backend even without KV-cache hits. With hits from the GPU KV cache of 20 GB, the latency improvement is significantly larger.

These results show two different gains. First, AOTI reduces serving overhead compared with the Python backend. Second, KV cache hits reduce model work by reusing cached sequence state. The cache benefit is especially visible on the deeper eight-layer model, where avoiding recomputation has more impact.

### Batch-size scaling with AOTI and KV cache

The PyTorch AOTI backend results by batch size show that KV caching becomes increasingly effective as logical batch size grows.

At batch size 8, the three-layer HSTU model reaches 0.423 ms latency per logical request with GPU KV-cache hits. The eight-layer HSTU model reaches 0.678 ms. These are strong results for long-sequence ranking inference and demonstrate the value of combining compiled model execution with cache-aware serving.

## What are the benefits of accelerating HSTU GR inference?

Recommendation systems operate under tight latency budgets. Additional ranking latency can affect page-load times, feed responsiveness, and ad-serving deadlines. Meanwhile, increasingly sequence-aware and personalized models can require more inference compute, particularly as user histories grow.

The HSTU serving workflow combines complementary technologies to address this challenge. HSTU provides the generative recommendation architecture, PyTorch AOTInductor produces an ahead-of-time compiled deployment artifact, FlexKV-backed KV caching enables reuse of previously computed attention state, and NVIDIA Dynamo-Triton provides the production serving environment.

This approach is especially valuable when successive requests share an unchanged prefix of a user’s interaction history. Rather than recomputing attention over that portion of the sequence, the model can reuse its cached key-value state and compute only what is required for newly appended tokens. The potential savings increase with longer sequences and deeper models, where repeated computation across multiple HSTU layers would otherwise add significant latency.

## Get started accelerating HSTU GR inference

You can reproduce and extend this workflow from the [NVIDIA/](https://github.com/NVIDIA/recsys-examples/tree/v26.06.01/examples/hstu/inference_aoti)[`recsys-examples`](https://github.com/NVIDIA/recsys-examples/tree/v26.06.01/examples/hstu/inference_aoti) GitHub repo. The HSTU overview introduces the GR model structure, including contextual tokens, item tokens, action tokens, embedding tables, HSTU blocks, and prediction heads.

The AOTI inference guide walks through building the required images and libraries, preparing KuaiRand-1K data, training a checkpoint, exporting the KV-cache AOTI model, validating it with C++ replay, packaging the Dynamo-Triton runtime image, and replaying requests through the Dynamo-Triton server.

To learn more, check out these related resources:

### Acknowledgments

*This post is a cross-functional effort across several NVIDIA teams. We would like to thank J, Runchu Zhao, Yulu Liu, Lin Hu, Zhuofan Li, Jacob Subag, and Tomer Bar-On for their contributions.*
