NVIDIA Announces Do Inference Now (DIN) Deploy for Local AI Inference
NVIDIA introduced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that accelerate local AI inference on Windows and Linux.
NVIDIA announced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that combine ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. DIN Deploy supports various AI tasks including automatic speech recognition and interactive segmentation.
Key facts
| Fact | Detail | The source says |
|---|---|---|
| Supported Tasks | Automatic speech recognition, interactive segmentation | “For automatic speech recognition (ASR), DIN Deploy supports offline and streaming pipelines. OpenAI Whisper covers offline transcription…” |
| Performance Gain | 39x to 206x real-time | “Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet…” |
| Supported Platforms | Windows, Linux | “The repository supports automatic speech recognition with OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming, as well…” |
What happened
NVIDIA announced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that combine ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. DIN Deploy bridges the gap between model checkpoints and native, hardware-accelerated applications. Each sample starts with a Python exporter that converts Hugging Face checkpoints into ONNX artifacts, followed by a native C++ CLI built on ONNX Runtime session and tensor APIs. The repository supports automatic speech recognition with OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming, as well as interactive segmentation with Meta SAM 2.1 and prompt-driven image generation with FLUX.2-klein-4B. Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet TDT, while SAM 2.1 achieves 38.3 FPS on GPU versus 0.5 FPS on CPU.
Source: NVIDIA
