cd/entity/NVFP4· home entities NVFP4
grep -l @nvfp4 /news/*.json | wc -l → 11

NVFP4

mentions 11 type Organization feed RSS

// recent coverage 11 mentions

16:00
2026-07-16
blog.getzep.com
artificial-intelligence

Evaluating Nemotron 3 Embed for agent memory

NVIDIA's Nemotron 3 Embed 1B embedding model took first place on every recall task in Zep's benchmark of 5,954 production queries, outperforming a 4B competitor with a third of the parameters and surp…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators

AMD has integrated an NVFP4 emulation pipeline into vLLM that enables AMD Instinct MI355 accelerators to serve standard NVFP4 quantized checkpoints directly, dequantizing weights to BF16 on-the-fly at…

20:26
2026-07-10
machinebrief.com
large-language-models

ARCQuant: Redefining Efficiency in LLM Inference with NVFP4

ARCQuant, a new framework for Large Language Model inference, uses the NVFP4 numerical format to achieve up to 3x speedup on GPUs while maintaining accuracy comparable to full-precision baselines. The…

13:10
2026-06-26
byteiota.com
ai-infrastructure

DGX Spark June 2026: Four Nodes, 700B Models Locally

NVIDIA's June 2026 DGX Spark update introduces automated four-node clustering via Cluster Assistant, enabling local inference of models up to 700B parameters. The update also delivers a 2.6x throughpu…

10:44
2026-06-19
discuss.huggingface.co
large-language-models

Gemma 4 bug fixes and Research Request

A critical bug in Google's Gemma 4 causes it to malform tool calls under real load, affecting vLLM, llama.cpp, Ollama, and oobabooga. A developer open-sourced a diagnosis, repair, and experimental LoR…

// co-occurs with top 8 entities
// topics top 6 topics