Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples
NVIDIA's open-source Do Inference Now (DIN) Deploy offers new C++ samples combining ONNX Runtime and TensorRT RTX for easier and faster local AI inference.
What changed?
NVIDIA released Do Inference Now (DIN) Deploy, an open-source set of C++ code samples. These examples demonstrate how to go from a model checkpoint (ONNX format) to a running native application using ONNX Runtime with the NVIDIA TensorRT RTX execution provider. The samples are aimed at practical, local deployment of AI models.

Why does it matter to an everyday developer?
Everyday developers often face challenges in deploying AI models locally: bridging the gap between exported models and efficient runtime on user hardware, especially for performance-critical applications. DIN Deploy simplifies this path for C++ projects, providing ready-made, practical code and enabling fast inference with GPU acceleration using TensorRT. This helps developers bring AI features to desktop, edge, or real-time applications without deep ML system integration work.
What can the developer do now?
- Download and explore the DIN Deploy C++ samples to understand end-to-end model inference workflows.
- Use provided code to integrate ONNX-format AI models with local or edge applications using C++.
- Accelerate inference in their apps with TensorRT on supported NVIDIA RTX hardware.
- Adapt and extend the open-source samples for production projects requiring performant, local AI.
