{"slug": "sub-1-bit-llm-compression-via-latent-factorization", "title": "Sub-1-Bit LLM Compression via Latent Factorization", "summary": "Banseok Lee and Youngmin Kim released the official implementation of LittleBit (NeurIPS 2025) and LittleBit-2 (ICML 2026), which compress large language models into the sub-1-bit regime down to 0.1 bits per weight by factorizing dense weight matrices into low-rank latent factors, binarizing them, and restoring magnitude with learned scales. LittleBit-2 adds Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ) to align SVD-derived latent factors with the binary hypercube before quantization-aware training, enabled via the opt-in --use_itq flag with no additional inference overhead. The codebase supports OPT, Llama and Llama 2/3, Phi-4, Qwen2.5, QwQ, Gemma 2, Gemma 3, and Qwen3, and recommends Python 3.12 with transformers 4.51.x for reproducing paper results.", "body_md": "Official implementation of **LittleBit** (NeurIPS 2025) and **LittleBit-2** (ICML 2026).\n\n**LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment** *(ICML 2026)*\n\nBanseok Lee, Youngmin Kim\n\n**LittleBit: Ultra Low-Bit Quantization via Latent Factorization** *(NeurIPS 2025)*\n\nBanseok Lee*, Dongkyu Kim*, Youngcheon You, Youngmin Kim\n\n**LittleBit** compresses large language models into the sub-1-bit regime by factorizing each dense weight matrix into low-rank latent factors, binarizing those factors, and restoring magnitude information through lightweight learned scales. This enables extreme compression, including the 0.1 bits-per-weight setting, while preserving the original model architecture at inference time.\n\n**LittleBit-2** improves this recipe by addressing latent geometry misalignment in the initialization stage. It applies Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ), aligning the SVD-derived latent factors with the binary hypercube before QAT. LittleBit-2 initialization is available as an opt-in (`--use_itq`) and produces no additional inference overhead.\n\n- **Sub-1-bit compression:** Designed for 1.0 to 0.1 bits per weight.\n- **LittleBit-2 opt-in:** Enable Joint-ITQ initialization with`--use_itq` for improved latent geometry alignment.\n- **No inference-time change:** LittleBit-2 modifies initialization only; the deployed factorized layer remains the same.\n- **QAT-friendly:** Supports Quantization-Aware Training with SmoothSign and optional residual factorization.\n\nThe codebase currently supports:\n\n- OPT\n- Llama and Llama 2/3\n- Phi-4\n- Qwen2.5 and QwQ\n- Gemma 2 and Gemma 3\n- Qwen3\n\nWe recommend Python 3.12.\n\n```\nconda create -n littlebit python=3.12\nconda activate littlebit\n\n# Install CUDA toolkit. Adjust the CUDA version if needed.\nconda install nvidia/label/cuda-12.4.1::cuda-toolkit -c nvidia/label/cuda-12.4.1\n\n# Install PyTorch.\npip install torch==2.8.0+cu124 torchvision==0.23.0+cu124 torchaudio==2.8.0+cu124 --index-url https://download.pytorch.org/whl/cu124\n\n# Install dependencies.\npip install -r requirements.txt\n```\n\nImportant\n\nFor reproducing the paper results, use `transformers` 4.51.x. Newer `transformers` releases may change model internals or evaluation behavior.\n\n```\npip install \"transformers==4.51.*\"\n```\n\nTrain a model with Quantization-Aware Training. By default, `LittleBitLinear` uses the original SVD-only initialization. To enable LittleBit-2 (Joint-ITQ), pass `--use_itq True`.\n\n**Single GPU**\n\n```\nCUDA_VISIBLE_DEVICES=0 python -m main \\\n    --model_id meta-llama/Llama-2-7b-hf \\\n    --dataset c4_wiki \\\n    --save_dir ./outputs/Llama-2-7b-LittleBit-2 \\\n    --num_train_epochs 5.0 \\\n    --per_device_train_batch_size 4 \\\n    --lr 4e-05 \\\n    --warmup_ratio 0.02 \\\n    --report wandb \\\n    --quant_func SmoothSign \\\n    --quant_mod LittleBitLinear \\\n    --residual True \\\n    --eff_bit 1.0 \\\n    --kv_factor 1.0 \\\n    --min_split_dim 8 \\\n    --l2l_loss_scale 10.0\n\n# Opt-in to LittleBit-2 initialization\n# --use_itq True\n```\n\n**Multi-GPU with DeepSpeed**\n\n```\ndeepspeed --num_gpus=4 main.py \\\n    --model_id meta-llama/Llama-2-7b-hf \\\n    --dataset c4_wiki \\\n    --save_dir ./outputs/Llama-2-7b-LittleBit-2 \\\n    --ds_config_path configs/zero3.json \\\n    --num_train_epochs 5.0 \\\n    --per_device_train_batch_size 4 \\\n    --lr 4e-05 \\\n    --report wandb \\\n    --quant_func SmoothSign \\\n    --quant_mod LittleBitLinear \\\n    --residual True \\\n    --eff_bit 1.0 \\\n    --kv_factor 1.0 \\\n    --min_split_dim 8\n```\n\nEvaluate a local checkpoint or a model hosted on the Hugging Face Hub.\n\n```\n# From a local directory\nCUDA_VISIBLE_DEVICES=0 python eval.py \\\n    --model_id ./outputs/Llama-2-7b-LittleBit-2 \\\n    --seqlen 2048 \\\n    --ppl_task wikitext2,c4 \\\n    --zeroshot_task boolq,piqa,hellaswag,winogrande,arc_easy,arc_challenge,openbookqa\n\n# From the Hugging Face Hub\nCUDA_VISIBLE_DEVICES=0 python eval.py \\\n    --model_id username/littlebit-llama-7b-0.1bpw \\\n    --seqlen 2048 \\\n    --ppl_task wikitext2\n```\n\nOlder checkpoints may not include `littlebit_config.json`. In that case, pass the quantization arguments explicitly:\n\n```\nCUDA_VISIBLE_DEVICES=0 python eval.py \\\n    --model_id ./outputs/Legacy-Llama-2-7b \\\n    --quant_func SmoothSign \\\n    --quant_mod LittleBitLinear \\\n    --split_dim 1024\n```\n\nParameter loading priority:\n\n1. Explicit CLI arguments\n2. `littlebit_config.json` in the model directory\n3. `config.json` fallback for older checkpoints\n\nIf you find this work useful, please cite:\n\n```\n@inproceedings{lee2026littlebit2,\n  title={LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment},\n  author={Lee, Banseok and Kim, Youngmin},\n  booktitle={Proceedings of the 43rd International Conference on Machine Learning},\n  year={2026}\n}\n@inproceedings{lee2025littlebit,\n  title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization},\n  author={Lee, Banseok and Kim, Dongkyu and You, Youngcheon and Kim, Youngmin},\n  booktitle={Advances in Neural Information Processing Systems},\n  year={2025}\n}\n```\n\nThis project is licensed under the [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) license.", "url": "https://wpnews.pro/news/sub-1-bit-llm-compression-via-latent-factorization", "canonical_source": "https://github.com/SamsungLabs/LittleBit", "published_at": "2026-10-08 13:29:46+00:00", "updated_at": "2026-10-08 13:47:28.476325+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "ai-infrastructure", "developer-tools"], "entities": ["LittleBit", "LittleBit-2", "Banseok Lee", "Youngmin Kim", "NeurIPS 2025", "ICML 2026", "Joint-ITQ", "Llama 2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/sub-1-bit-llm-compression-via-latent-factorization", "markdown": "https://wpnews.pro/news/sub-1-bit-llm-compression-via-latent-factorization.md", "text": "https://wpnews.pro/news/sub-1-bit-llm-compression-via-latent-factorization.txt", "jsonld": "https://wpnews.pro/news/sub-1-bit-llm-compression-via-latent-factorization.jsonld"}}