Show HN: Spectra-FFT – Cutting AdamW VRAM Usage by 50% for LLMs Spectra-FFT, a new optimizer presented on Hacker News, claims to cut AdamW optimizer-state VRAM usage by roughly 55% for full-parameter LLM fine-tuning, replacing the two 32-bit states AdamW keeps per parameter (8 bytes each) with a compact frequency-domain representation. The project reports a logged run completing full-parameter fine-tuning inside a strict 5GB VRAM partition of an NVIDIA A100, where the same job with AdamW hits CUDA out of memory, and offers a one-line swap of AdamW(...) for Spectra(...) on CUDA with no CPU offloading or low-rank adapters. The implementation, parameters and tuning are available only to evaluation partners under an NDA, with access reviewed manually, usually within 48 hours. The Next Layer of Super Intelligence Train powerful AI models on half the GPU memory. Run bigger AI on the hardware you already own, and stop paying for chips you don’t need. Get Started colab View Architecture architecture In plain English Training an AI model needs a lot of GPU memory, and that memory is the most expensive part of the hardware. Much of it is spent on bookkeeping rather than on the model itself. Spectra-FFT is a drop-in replacement that does the same bookkeeping in far less space. AI training keeps running out of GPU memory. Teams either rent bigger, pricier machines or give up on the model they wanted to build. We found a smarter way to store the training "bookkeeping" so it takes roughly half the memory, while the model learns the same way. Train larger models on the GPUs you already have, or run the same job on smaller, cheaper cloud machines. Switching takes one line of code. The problem · technical detail AdamW keeps two 32-bit states for every parameter: 8 bytes each, on top of the weights. That state is what pushes full-parameter fine-tuning into out-of-memory crashes on constrained hardware. Spectra-FFT is a new optimizer that stores far less of it. Gradients carry signal and noise. Spectra works in the frequency domain, where the two separate cleanly. Optimizer state is stored for the part of the signal that drives learning, not for every parameter. Full-parameter training with no low-rank adapters and a much smaller memory footprint. Architecture Spectra-FFT sits exactly where your optimizer does today. Nothing upstream or downstream changes: same model, same data, same loss, same scheduler. One line replaces AdamW ... with Spectra ... . Works with your existing training scripts, mixed precision and checkpointing. Instead of two full-size states per parameter, Spectra maintains a compact representation of the training signal, which is where the 50%+ optimizer-memory reduction comes from. Runs entirely on-device on CUDA using standard NVIDIA libraries. No CPU offloading, no custom hardware, no extra data movement. Every weight is trained. This is not LoRA or an adapter, so there is no low-rank ceiling on what the model can learn. The concept, the measured results, the paper and the training logs. Everything needed to evaluate whether Spectra works. The implementation, parameters and tuning that make it work. Available to evaluation partners under NDA. VRAM engine Estimate assumes bf16/fp16 weights 2 bytes/param and fp32 AdamW momentum + variance 8 bytes/param . Spectra figure applies the ~55% optimizer-state reduction reported in our paper. Activations, gradients and KV cache are excluded and depend on batch size and sequence length. Evidence Full-parameter fine-tuning inside a strict 5GB VRAM partition of an NVIDIA A100, logged end to end. Evaluation access We share a private Google Colab notebook that trains the same model twice on a standard GPU: once with AdamW, which runs out of memory, and once with Spectra-FFT, which completes. Access is granted to engineers and teams evaluating Spectra. php What the notebook demonstrates opt = AdamW model.parameters - CUDA out of memory opt = Spectra model.parameters - trains within budget Requests are reviewed manually, usually within 48 hours. Evaluation access is provided under a short NDA. FAQ Yes. It is a drop-in optimizer for full-parameter training, aimed at memory-constrained GPUs. Not currently. The paper describes the method and the demo lets you verify the behavior. Evaluation and partnership access is available on request. Our logged runs track baseline convergence on the evaluated setup. See the paper and W&B logs for the exact configuration and limits. Teams fine-tuning models on-premise or on consumer and mid-range GPUs, where optimizer state is the bottleneck.