cd/entity/AirLLM· home entities AirLLM
grep -l @airllm /news/*.json | wc -l → 8

AirLLM

mentions 8 type Organization feed RSS

// recent coverage 8 mentions

19:41
2026-08-05
dev.to
large-language-models

AirLLM: Running 70B Parameter LLMs on a Single 4GB GPU

Lyogavin has released AirLLM, an open-source Python library that enables running 70B parameter large language models on GPUs with as little as 4GB of VRAM. By streaming model layers sequentially from …

15:09
2026-08-04
sourcefeed.dev
machine-learning

An 8B Fine-Tune Now Fits in 4 GB of VRAM

Independent researcher Alpamys Makazhan released Soup, a Show HN project that fine-tunes a full Llama-3.1-8B model in NF4 quantization with a 3.32 GB VRAM peak at 119.6 tokens/sec on a 4 GB RTX 3050 L…

08:23
2026-08-04
snipvote.com
artificial-intelligence

AirLLM 70B inference with single 4GB GPU

AirLLM, a new tool from GitHub user lyogavin, enables inference of 70B-parameter models on a single 4GB GPU by aggressively quantizing and offloading layers, potentially reducing hardware costs by 10–…

16:09
2026-08-03
sourcefeed.dev
artificial-intelligence

AirLLM's 4GB 70B Trick Is Real, and Beside the Point

AirLLM, an open-source tool by Gavin Li, can run a 70B-parameter model on a 4GB GPU by streaming layers from disk, but this method is limited by disk bandwidth, resulting in seconds to minutes per tok…

09:02
2026-08-02
github.com
large-language-models

AirLLM: Inference 2.8T Kimi K3 on a single 4GB GPU

AirLLM, an open-source inference library, now supports running the 2.8T-parameter Kimi K3 model on a single 4GB GPU, using only 3.72GB of VRAM on an RTX 6000 Ada, by streaming one expert at a time for…

18:29
2026-06-16
dev.to
large-language-models

70B AI Model Runs on 8GB Laptop

A developer released AirLLM, an open-source tool that enables running 70-billion-parameter AI models on laptops with as little as 8GB RAM, without requiring a GPU. By using memory mapping and layer sw…

// co-occurs with top 8 entities
// topics top 6 topics