cd/entity/GPTQ· home entities GPTQ
grep -l @gptq /news/*.json | wc -l → 18

GPTQ

mentions 18 type Organization feed RSS

// recent coverage 18 mentions

16:45
2026-08-29
promptcube3.com
large-language-models

Running massive LLMs on consumer hardware is a financial

Quantization and pruning techniques such as GPTQ, AWQ, GGUF, SparseGPT, and LoRA-based pruning enable running large language models on consumer hardware by reducing memory footprint, with 4-bit quanti…

04:29
2026-08-23
pub.towardsai.net
artificial-intelligence

How BitNet Run a Transformer With (Almost) No Multiplication?

Microsoft's BitNet research demonstrates that large language models can run with weights restricted to just −1, 0, or +1, eliminating multiplications and reducing memory by an order of magnitude. The …

10:44
2026-08-20
promptcube3.com
machine-learning

Pine AI tops τ³-Voice leaderboard at 75.

Pine AI's 1.2B-parameter speech recognition model tops the τ³-Voice leaderboard with a score of 75, achieving significant gains on accented speech, medical dictation, and code-switching, but only marg…

12:00
2026-08-04
kdnuggets.com
large-language-models

7 Approaches to Reduce Inference Latency in Your LLM Workflows

Seven engineering strategies to reduce inference latency in large language model (LLM) workflows are outlined, including model quantization, key-value caching, and speculative decoding. The approaches…

00:00
2026-07-03
deepresearch.ninja
large-language-models

LLM Quantization Methods: A Comprehensive Comparative Analysis

A comprehensive analysis of 15+ large language model quantization methods categorizes them into four paradigms: CPU-optimized GGUF-based approaches, GPU-native weight-only kernels, NVIDIA's floating-p…

10:53
2026-06-27
dev.to
machine-learning

How I Implemented GPTQ from Scratch (and What I Learned)

A developer implemented GPTQ quantization from scratch on a nanoGPT model, achieving only 1.1% perplexity degradation across 61 quantized layers. The implementation uses second-order optimization to r…

13:01
2026-06-24
gist.github.com
large-language-models

NVIDIA GenAI LLM Certification Lab

NVIDIA has released a GenAI LLM Certification Lab that guides developers through building a production-ready fine-tuning and optimization pipeline. The lab covers data preparation, LoRA fine-tuning wi…

14:58
2026-06-06
vettedconsumer.com
large-language-models

GGUF vs. GPTQ vs. AWQ: The Plain-English Guide to LLM Quantization

GGUF, GPTQ, and AWQ are the three dominant formats for running quantized large language models locally, each optimized for different hardware and use cases. GGUF, the format used by llama.cpp and its …

// co-occurs with top 8 entities
// topics top 6 topics