cd/entity/Activation-aware Weight Quantization (AWQ)Β· homeβ€Ί entitiesβ€Ί Activation-aware Weight Quantization (AWQ)
grep -l @activation-aware weight quantization (awq) /news/*.json | wc -l β†’ 1

Activation-aware Weight Quantization (AWQ)

mentions 1 type Person feed RSS

// recent coverage 1 mentions

12:00
2026-08-04
kdnuggets.com
large-language-models

7 Approaches to Reduce Inference Latency in Your LLM Workflows

Seven engineering strategies to reduce inference latency in large language model (LLM) workflows are outlined, including model quantization, key-value caching, and speculative decoding. The approaches…

// co-occurs with top 1 entities
// topics top 3 topics