Sub-1-Bit LLM Compression via Latent Factorization Banseok Lee and Youngmin Kim released the official implementation of LittleBit (NeurIPS 2025) and LittleBit-2 (ICML 2026), which compress large language models into the sub-1-bit regime down to 0.1 bits per weight by factorizing dense weight matrices into low-rank latent factors, binarizing them, and restoring magnitude with learned scales. LittleBit-2 adds Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ) to align SVD-derived latent factors with the binary hypercube before quantization-aware training, enabled via the opt-in --use_itq flag with no additional inference overhead. The codebase supports OPT, Llama and Llama 2/3, Phi-4, Qwen2.5, QwQ, Gemma 2, Gemma 3, and Qwen3, and recommends Python 3.12 with transformers 4.51.x for reproducing paper results. Official implementation of LittleBit NeurIPS 2025 and LittleBit-2 ICML 2026 . LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment ICML 2026 Banseok Lee, Youngmin Kim LittleBit: Ultra Low-Bit Quantization via Latent Factorization NeurIPS 2025 Banseok Lee , Dongkyu Kim , Youngcheon You, Youngmin Kim LittleBit compresses large language models into the sub-1-bit regime by factorizing each dense weight matrix into low-rank latent factors, binarizing those factors, and restoring magnitude information through lightweight learned scales. This enables extreme compression, including the 0.1 bits-per-weight setting, while preserving the original model architecture at inference time. LittleBit-2 improves this recipe by addressing latent geometry misalignment in the initialization stage. It applies Internal Latent Rotation with Joint Iterative Quantization Joint-ITQ , aligning the SVD-derived latent factors with the binary hypercube before QAT. LittleBit-2 initialization is available as an opt-in --use itq and produces no additional inference overhead. - Sub-1-bit compression: Designed for 1.0 to 0.1 bits per weight. - LittleBit-2 opt-in: Enable Joint-ITQ initialization with --use itq for improved latent geometry alignment. - No inference-time change: LittleBit-2 modifies initialization only; the deployed factorized layer remains the same. - QAT-friendly: Supports Quantization-Aware Training with SmoothSign and optional residual factorization. The codebase currently supports: - OPT - Llama and Llama 2/3 - Phi-4 - Qwen2.5 and QwQ - Gemma 2 and Gemma 3 - Qwen3 We recommend Python 3.12. conda create -n littlebit python=3.12 conda activate littlebit Install CUDA toolkit. Adjust the CUDA version if needed. conda install nvidia/label/cuda-12.4.1::cuda-toolkit -c nvidia/label/cuda-12.4.1 Install PyTorch. pip install torch==2.8.0+cu124 torchvision==0.23.0+cu124 torchaudio==2.8.0+cu124 --index-url https://download.pytorch.org/whl/cu124 Install dependencies. pip install -r requirements.txt Important For reproducing the paper results, use transformers 4.51.x. Newer transformers releases may change model internals or evaluation behavior. pip install "transformers==4.51. " Train a model with Quantization-Aware Training. By default, LittleBitLinear uses the original SVD-only initialization. To enable LittleBit-2 Joint-ITQ , pass --use itq True . Single GPU CUDA VISIBLE DEVICES=0 python -m main \ --model id meta-llama/Llama-2-7b-hf \ --dataset c4 wiki \ --save dir ./outputs/Llama-2-7b-LittleBit-2 \ --num train epochs 5.0 \ --per device train batch size 4 \ --lr 4e-05 \ --warmup ratio 0.02 \ --report wandb \ --quant func SmoothSign \ --quant mod LittleBitLinear \ --residual True \ --eff bit 1.0 \ --kv factor 1.0 \ --min split dim 8 \ --l2l loss scale 10.0 Opt-in to LittleBit-2 initialization --use itq True Multi-GPU with DeepSpeed deepspeed --num gpus=4 main.py \ --model id meta-llama/Llama-2-7b-hf \ --dataset c4 wiki \ --save dir ./outputs/Llama-2-7b-LittleBit-2 \ --ds config path configs/zero3.json \ --num train epochs 5.0 \ --per device train batch size 4 \ --lr 4e-05 \ --report wandb \ --quant func SmoothSign \ --quant mod LittleBitLinear \ --residual True \ --eff bit 1.0 \ --kv factor 1.0 \ --min split dim 8 Evaluate a local checkpoint or a model hosted on the Hugging Face Hub. From a local directory CUDA VISIBLE DEVICES=0 python eval.py \ --model id ./outputs/Llama-2-7b-LittleBit-2 \ --seqlen 2048 \ --ppl task wikitext2,c4 \ --zeroshot task boolq,piqa,hellaswag,winogrande,arc easy,arc challenge,openbookqa From the Hugging Face Hub CUDA VISIBLE DEVICES=0 python eval.py \ --model id username/littlebit-llama-7b-0.1bpw \ --seqlen 2048 \ --ppl task wikitext2 Older checkpoints may not include littlebit config.json . In that case, pass the quantization arguments explicitly: CUDA VISIBLE DEVICES=0 python eval.py \ --model id ./outputs/Legacy-Llama-2-7b \ --quant func SmoothSign \ --quant mod LittleBitLinear \ --split dim 1024 Parameter loading priority: 1. Explicit CLI arguments 2. littlebit config.json in the model directory 3. config.json fallback for older checkpoints If you find this work useful, please cite: @inproceedings{lee2026littlebit2, title={LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment}, author={Lee, Banseok and Kim, Youngmin}, booktitle={Proceedings of the 43rd International Conference on Machine Learning}, year={2026} } @inproceedings{lee2025littlebit, title={LittleBit: Ultra Low-Bit Quantization via Latent Factorization}, author={Lee, Banseok and Kim, Dongkyu and You, Youngcheon and Kim, Youngmin}, booktitle={Advances in Neural Information Processing Systems}, year={2025} } This project is licensed under the CC BY-NC 4.0 https://creativecommons.org/licenses/by-nc/4.0/ license.