{"slug": "nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one", "title": "Nemotron-3-Omni GGUFs with working audio + video in llama.cpp (9 tested quants, one-pass A/V)", "summary": "NVIDIA's Nemotron-3-Nano-Omni-30B multimodal model now runs fully in llama.cpp with working audio and video, thanks to a new GGUF release by engram-ae and a llama.cpp fork by VincentKaufmann. The release includes nine quantizations from Q3_K_M to BF16 and a unified projector file supporting both C-RADIO vision and Parakeet audio encoders, enabling one-pass video with soundtrack on a single DGX Spark. The fork adds the Parakeet/FastConformer audio graph and C-RADIO video graph, with prebuilt Linux arm64 CUDA 13 binaries, and an upstream PR for the video path is planned.", "body_md": "NVIDIA’s Nemotron-3-Nano-Omni-30B is a genuinely multimodal model: image, audio, and video in. But until now the GGUF ecosystem only carried its vision half. The existing GGUF repos ship vision-only projectors, and the video path did not exist in llama.cpp at all.\n\nI built the missing pieces and released the whole package:\n\n**Models:** [engram-ae/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF · Hugging Face](https://huggingface.co/engram-ae/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF)\n\nNine quantizations from Q3_K_M to BF16, plus a unified projector file that carries both encoder towers (C-RADIO vision and Parakeet audio). One llama-server instance takes every modality on /v1/chat/completions, including the built-in web UI: drag in a photo, a wav, or an mp4.\n\n**Code:** [GitHub - VincentKaufmann/llama.cpp-omni: llama.cpp fork: Parakeet/FastConformer audio graph for Nemotron-3-Omni (mtmd) · GitHub](https://github.com/VincentKaufmann/llama.cpp-omni)\n\nA llama.cpp fork with the Parakeet/FastConformer audio graph and the C-RADIO video graph (temporal patches plus EVS token pruning), including one-pass video WITH its soundtrack: the model describes what it sees and quotes what it hears from a single request. Prebuilt Linux arm64 CUDA 13 binaries in the releases (DGX Spark, Jetson Thor).\n\n**Verification, since claims are cheap:**\n\nEverything was built and tested on a single DGX Spark (GB10). Quick start is in the model card: one server flag pair, then open the browser.\n\nRoadmap: the next release rebases onto current llama.cpp master and adopts the upstream audio projector layout, so image, text and audio will run on stock llama.cpp with these same files; the fork remains for video and one-pass A/V. An upstream PR for the video path is planned.\n\nBase model and encoders are NVIDIA’s, under the NVIDIA Open Model Agreement. I wrote the GGUF conversion, the llama.cpp graphs, and the projector unification. Happy to answer questions or fix what breaks.", "url": "https://wpnews.pro/news/nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one", "canonical_source": "https://discuss.huggingface.co/t/nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one-pass-a-v/179220#post_1", "published_at": "2026-08-24 17:47:55+00:00", "updated_at": "2026-08-24 18:15:00.324149+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "generative-ai", "ai-infrastructure", "developer-tools"], "entities": ["NVIDIA", "Nemotron-3-Nano-Omni-30B", "engram-ae", "VincentKaufmann", "llama.cpp", "Hugging Face", "GitHub", "DGX Spark"], "alternates": {"html": "https://wpnews.pro/news/nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one", "markdown": "https://wpnews.pro/news/nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one.md", "text": "https://wpnews.pro/news/nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one.txt", "jsonld": "https://wpnews.pro/news/nemotron-3-omni-ggufs-with-working-audio-video-in-llama-cpp-9-tested-quants-one.jsonld"}}