Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents
NVIDIA introduces Nemotron 3 Nano Omni, a long-context multimodal intelligence model for documents, audio, and video agents, offering best-in-class accuracy and efficiency in various benchmarks. The model extends the Nemotron multimodal line to a broader text, image, video, and audio model. It achieves top accuracy on several leaderboards, including MMlongbench-Doc, OCRBenchV2, and VoiceBench.