cd/entity/NVFP4· home entities NVFP4
grep -l @nvfp4 /news/*.json | wc -l → 20

NVFP4

mentions 20 type Organization feed RSS

// recent coverage 20 mentions

04:00
2026-09-02
arxiv.org
artificial-intelligence

OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

Researchers introduced OCGQuant, a post-training quantization method that pairs outlier channels with low-magnitude companions to reduce quantization error in NVFP4 low-bit inference, achieving the lo…

02:46
2026-09-02
forgeeks.net
machine-learning

LLM inference now has two ways to get cheaper

Baseten's technical breakdown of LLM inference efficiency identifies two categories of engineering choices: those that trade latency for throughput, such as batch sizing, tensor parallelism, expert pa…

00:00
2026-08-28
inco.ai
artificial-intelligence

Inco AI launches Day-0 support for GLM 5.3

Inco AI launched day-0 support for Z.ai's GLM 5.3 model, releasing DFlash 2 and NVFP4 checkpoints alongside an Inco Engine endpoint that delivers up to 4.4× throughput versus the native FP8 checkpoint…

14:00
2026-08-25
dev.to
machine-learning

A Better FP4 Gradient Quantizer That Training Couldn't Notice

A developer found a scale-selection rule for NVFP4 gradient quantization that reduces mean squared error by 14% on real gradient tensors compared to the state-of-the-art MS-EDEN estimator, but trainin…

00:00
2026-08-11
mindstudio.ai
large-language-models

How to Run Nemotron 3.5 Lightning Locally on Your Own GPU

NVIDIA's open-weights Nemotron 3.5 Lightning mixture-of-experts model, with 30 billion total parameters but only 3 billion active, is designed for agentic grunt work and can be run locally on consumer…

00:00
2026-07-23
huggingface.co
artificial-intelligence

Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Hugging Face has integrated Nunchaku 4-bit diffusion inference natively into Diffusers, enabling users to load quantized checkpoints with a simple from_pretrained() call and no local CUDA compilation.…

16:00
2026-07-16
blog.getzep.com
artificial-intelligence

Evaluating Nemotron 3 Embed for agent memory

NVIDIA's Nemotron 3 Embed 1B embedding model took first place on every recall task in Zep's benchmark of 5,954 production queries, outperforming a 4B competitor with a third of the parameters and surp…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators

AMD has integrated an NVFP4 emulation pipeline into vLLM that enables AMD Instinct MI355 accelerators to serve standard NVFP4 quantized checkpoints directly, dequantizing weights to BF16 on-the-fly at…

20:26
2026-07-10
machinebrief.com
large-language-models

ARCQuant: Redefining Efficiency in LLM Inference with NVFP4

ARCQuant, a new framework for Large Language Model inference, uses the NVFP4 numerical format to achieve up to 3x speedup on GPUs while maintaining accuracy comparable to full-precision baselines. The…

13:10
2026-06-26
byteiota.com
ai-infrastructure

DGX Spark June 2026: Four Nodes, 700B Models Locally

NVIDIA's June 2026 DGX Spark update introduces automated four-node clustering via Cluster Assistant, enabling local inference of models up to 700B parameters. The update also delivers a 2.6x throughpu…

10:44
2026-06-19
discuss.huggingface.co
large-language-models

Gemma 4 bug fixes and Research Request

A critical bug in Google's Gemma 4 causes it to malform tool calls under real load, affecting vLLM, llama.cpp, Ollama, and oobabooga. A developer open-sourced a diagnosis, repair, and experimental LoR…

// co-occurs with top 8 entities
// topics top 6 topics