Skip to content
Intelligence RadarNews & Launches2 MIN READ

NVIDIA Announces Do Inference Now (DIN) Deploy for Local AI Inference

NVIDIA introduced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that accelerate local AI inference on Windows and Linux.

Fathom Intelligence
Fathom IntelligenceFathom Layer Expert

NVIDIA announced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that combine ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. DIN Deploy supports various AI tasks including automatic speech recognition and interactive segmentation.

Key facts

Fact Detail The source says
Supported Tasks Automatic speech recognition, interactive segmentation “For automatic speech recognition (ASR), DIN Deploy supports offline and streaming pipelines. OpenAI Whisper covers offline transcription…”
Performance Gain 39x to 206x real-time “Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet…”
Supported Platforms Windows, Linux “The repository supports automatic speech recognition with OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming, as well…”

What happened

NVIDIA announced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples that combine ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. DIN Deploy bridges the gap between model checkpoints and native, hardware-accelerated applications. Each sample starts with a Python exporter that converts Hugging Face checkpoints into ONNX artifacts, followed by a native C++ CLI built on ONNX Runtime session and tensor APIs. The repository supports automatic speech recognition with OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming, as well as interactive segmentation with Meta SAM 2.1 and prompt-driven image generation with FLUX.2-klein-4B. Performance measurements on DGX Spark show GPU acceleration ranging from 39x real-time for Nemotron ASR streaming to 206x for Parakeet TDT, while SAM 2.1 achieves 38.3 FPS on GPU versus 0.5 FPS on CPU.


Source: NVIDIA

END OF SIGNAL

Don't Miss the Next Signal

Get high-impact hardware and AI launches distilled into your inbox. No noise, just the changes that matter.

Occasional verified technology updates. Unsubscribe anytime. See our privacy policy.