17:47
2026-08-24
discuss.huggingface.co
artificial-intelligence
Nemotron-3-Omni GGUFs with working audio + video in llama.cpp (9 tested quants, one-pass A/V)
NVIDIA's Nemotron-3-Nano-Omni-30B multimodal model now runs fully in llama.cpp with working audio and video, thanks to a new GGUF release by engram-ae and a llama.cpp fork by VincentKaufmann. The releβ¦