cd/entity/FP16· home› entities› FP16
grep -l @fp16 /news/*.json | wc -l → 11

FP16

mentions 11 type Organization feed RSS

// recent coverage 11 mentions

20:54
2026-08-31
promptcube3.com
large-language-models

Why Qwen3.

A hands-on benchmark of Qwen3.8 27B on an RTX 4090 found that 4-bit quantization (NF4 and AWQ INT4) preserves near-baseline quality, with MMLU scores of 59.8% and 59.5% versus 61.2% for FP16, while 1-…

08:28
2026-07-01
promptcube3.com
artificial-intelligence

AI\'s Memory Crunch Hits Indian Smartphones

Indian smartphone manufacturers are pushing 12GB or 16GB of RAM as the new standard for mid-range devices to support on-device AI features, which consume significant memory through model weights and K…

14:39
2026-06-19
letsdatascience.com
large-language-models

DigitalOcean Demonstrates LLM Compression with SparseGPT

DigitalOcean published a tutorial on June 19 demonstrating how to compress large language models using SparseGPT and Wanda pruning methods for GPU cloud deployment, targeting reduced inference costs a…

01:44
2026-06-16
dev.to
artificial-intelligence

Balanced Ternary for optimizing AI

A developer argues that balanced ternary (-1, 0, +1) could replace binary for AI hardware, citing 20× model compression, 3× inference speedup, and 8× power reduction. Microsoft's BitNet b1.58 demonstr…

14:58
2026-06-06
vettedconsumer.com
large-language-models

GGUF vs. GPTQ vs. AWQ: The Plain-English Guide to LLM Quantization

GGUF, GPTQ, and AWQ are the three dominant formats for running quantized large language models locally, each optimized for different hardware and use cases. GGUF, the format used by llama.cpp and its …

// co-occurs with top 8 entities
// topics top 6 topics