Adding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems.
Do Inference Now (DIN) Deploy is an open-source collection of practical C++ samples that bridges that gap. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware-accelerated application on Windows and Linux. The same ONNX Runtime API can also be accessed through WinML 2.0.
From ONNX export to C++ implementation #
Each DIN Deploy sample starts with a Python exporter that downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact. The application side is a native C++ CLI built on ONNX Runtime (ORT). That split keeps model conversion separate from deployment logic, and developers can take an exported model into a local application without requiring a model-specific runtime.
Most sample code uses ONNX Runtime session and tensor APIs in C++. Vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths. Execution providers that support the required ONNX Runtime tensor APIs can run the shared code.
ORT’s copy tensor API keeps data locality manageable without dedicated vendor API usage in shared code.
For preprocessing and postprocessing around exported model inference, the FLUX.2 sample uses ONNX Runtime’s graphics interop capability, introduced in version 1.25, with Vulkan and DirectX for sampling. The repository provides CMake presets for Windows and Linux, including Arm64 variants. DirectX is available only on Windows.
AI tasks supported by DIN Deploy #
For automatic speech recognition (ASR), DIN Deploy supports offline and streaming pipelines. OpenAI Whisper covers offline transcription across model sizes, while NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming provide streaming pipelines. The samples show how to move audio through a native application and return transcription results while using GPU acceleration where it is available.
Meta SAM 2.1 samples support interactive masking for images and video. They turn model outputs into segmentation masks that native applications can use for selection, tracking, and other computer-vision workflows.
Table 1 compares GPU and CPU performance for selected DIN Deploy workloads measured on DGX Spark.
| **Model** | **GPU (DGX Spark)** | **CPU (DGX Spark)** |
|---|---|---|
| `openai/whisper-large-v3-turbo` | 58.5x | 3.8x |
| `nvidia/nemotron-3.5-asr-streaming-0.6b` | 39.01x | 3.24x |
| `nvidia/parakeet-tdt-0.6b-v3` | 206.41x | 14.44x |
| `facebook/sam2.1-hiera-base-plus` | 38.3 FPS | 0.5 FPS |
Table 1. DGX Spark GPU and CPU performance for selected DIN Deploy workloads. Results for audio are reported as multiples of real time (higher is faster)
The FLUX.2-klein-4B sample provides prompt-driven image generation. The sample includes graphics-API interops with Vulkan and DirectX, allowing applications to integrate GPU-resident resources with a cross-vendor shader interface. It also shows how post-training quantization (PTQ) with NVIDIA Model Optimizer produces a quantized ONNX model. Quantization is hardware-dependent, but due to ONNX interfaces remaining unchanged, the quantized model is a drop-in replacement requiring no application-code changes.
Get started with DIN Deploy #
Start with the repository’s CMake presets for Windows, Linux, x86-64, and Arm64. After configuring and building the project, export a model to ONNX and run the CLI with TensorRT RTX, or copy the code into your own application and use the pipeline implementations. CMake downloads ONNX Runtime and TensorRT RTX by default.
Learn more about TensorRT for RTX, NVIDIA Local AI, and the DIN Deploy repository, and the NVIDIA blog series on model quantization.