cd/entity/AWQ· home› entities› AWQ
grep -l @awq /news/*.json | wc -l → 25

AWQ

mentions 25 type Organization page 1/2 feed RSS

// recent coverage 25 mentions

23:01
2026-09-26
pub.towardsai.net
large-language-models

Fine-Tuning to Quantization: What Free-Tier Hardware Can Prove

QLoRA fine-tuning on 150 synthetic examples raised a base Qwen3.5-4B model's adversarial-prompt refusal rate from 79.5% (31 of 39) to 89.7% (35 of 39), while GPTQ and AWQ 4-bit quantization added roug…

00:00
2026-09-08
unite.ai
artificial-intelligence

How Long Before a Real Crackdown on AI Model Decensoring?

Open-source AI models are increasingly being stripped of their safety filters and redistributed at scale, echoing the warez scene of 1995–2010, according to an analysis on Unite.AI. The practice invol…

16:45
2026-08-29
promptcube3.com
large-language-models

Running massive LLMs on consumer hardware is a financial

Quantization and pruning techniques such as GPTQ, AWQ, GGUF, SparseGPT, and LoRA-based pruning enable running large language models on consumer hardware by reducing memory footprint, with 4-bit quanti…

04:29
2026-08-23
pub.towardsai.net
artificial-intelligence

How BitNet Run a Transformer With (Almost) No Multiplication?

Microsoft's BitNet research demonstrates that large language models can run with weights restricted to just −1, 0, or +1, eliminating multiplications and reducing memory by an order of magnitude. The …

22:34
2026-08-18
promptcube3.com
large-language-models

Llama 3.

Meta's Llama 3.1 70B model can now run on a single 24GB consumer GPU using GGUF or EXL2 quantization, achieving 5-10 tokens per second on an RTX 3090, according to a deployment guide. The guide recomm…

20:57
2026-07-31
dev.to
machine-learning

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

A developer's analysis shows that INT4 weight-only quantization speeds up decode but not prefill, because prefill is compute-bound while decode is memory-bound. The crossover point where a GEMM become…

00:00
2026-07-03
deepresearch.ninja
large-language-models

LLM Quantization Methods: A Comprehensive Comparative Analysis

A comprehensive analysis of 15+ large language model quantization methods categorizes them into four paradigms: CPU-optimized GGUF-based approaches, GPU-native weight-only kernels, NVIDIA's floating-p…

00:00
2026-06-30
aclanthology.org
large-language-models

Diet-KIT: Post-Training Quantization for Speech LLMs

Researchers from the Karlsruhe Institute of Technology developed Diet-KIT, a post-training quantization system that compresses the Qwen2-Audio-7B speech LLM from 16 GB to 3.98 GB while maintaining tra…

10:30
2026-06-26
aazar.me
large-language-models

Stop generating what you already have

A developer reduced LLM extraction latency from 42 seconds to 6 seconds by replacing verbatim text copying with pointer-based extraction and splitting a single large call into multiple parallel calls.…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics