{"slug": "show-hn-spectra-fft-cutting-adamw-vram-usage-by-50-for-llms", "title": "Show HN: Spectra-FFT – Cutting AdamW VRAM Usage by 50% for LLMs", "summary": "Spectra-FFT, a new optimizer presented on Hacker News, claims to cut AdamW optimizer-state VRAM usage by roughly 55% for full-parameter LLM fine-tuning, replacing the two 32-bit states AdamW keeps per parameter (8 bytes each) with a compact frequency-domain representation. The project reports a logged run completing full-parameter fine-tuning inside a strict 5GB VRAM partition of an NVIDIA A100, where the same job with AdamW hits CUDA out of memory, and offers a one-line swap of AdamW(...) for Spectra(...) on CUDA with no CPU offloading or low-rank adapters. The implementation, parameters and tuning are available only to evaluation partners under an NDA, with access reviewed manually, usually within 48 hours.", "body_md": "# The Next Layer of Super Intelligence\n\nTrain powerful AI models on half the GPU memory. Run bigger AI on the hardware you already own, and stop paying for chips you don’t need.\n\n[Get Started](#colab)\n\n[View Architecture](#architecture)\n\nIn plain English\n\nTraining an AI model needs a lot of GPU memory, and that memory is the most expensive part of the hardware. Much of it is spent on bookkeeping rather than on the model itself. Spectra-FFT is a drop-in replacement that does the same bookkeeping in far less space.\n\nAI training keeps running out of GPU memory. Teams either rent bigger, pricier machines or give up on the model they wanted to build.\n\nWe found a smarter way to store the training \"bookkeeping\" so it takes roughly half the memory, while the model learns the same way.\n\nTrain larger models on the GPUs you already have, or run the same job on smaller, cheaper cloud machines. Switching takes one line of code.\n\nThe problem · technical detail\n\nAdamW keeps two 32-bit states for every parameter: 8 bytes each, on top of the weights. That state is what pushes full-parameter fine-tuning into out-of-memory crashes on constrained hardware. Spectra-FFT is a new optimizer that stores far less of it.\n\nGradients carry signal and noise. Spectra works in the frequency domain, where the two separate cleanly.\n\nOptimizer state is stored for the part of the signal that drives learning, not for every parameter.\n\nFull-parameter training with no low-rank adapters and a much smaller memory footprint.\n\nArchitecture\n\nSpectra-FFT sits exactly where your optimizer does today. Nothing upstream or downstream changes: same model, same data, same loss, same scheduler.\n\nOne line replaces `AdamW(...)` with `Spectra(...)`. Works with your existing training scripts, mixed precision and checkpointing.\n\nInstead of two full-size states per parameter, Spectra maintains a compact representation of the training signal, which is where the 50%+ optimizer-memory reduction comes from.\n\nRuns entirely on-device on CUDA using standard NVIDIA libraries. No CPU offloading, no custom hardware, no extra data movement.\n\nEvery weight is trained. This is not LoRA or an adapter, so there is no low-rank ceiling on what the model can learn.\n\nThe concept, the measured results, the paper and the training logs. Everything needed to evaluate whether Spectra works.\n\nThe implementation, parameters and tuning that make it work. Available to evaluation partners under NDA.\n\nVRAM engine\n\nEstimate assumes bf16/fp16 weights (2 bytes/param) and fp32 AdamW momentum + variance (8 bytes/param). Spectra figure applies the ~55% optimizer-state reduction reported in our paper. Activations, gradients and KV cache are excluded and depend on batch size and sequence length.\n\nEvidence\n\nFull-parameter fine-tuning inside a strict 5GB VRAM partition of an NVIDIA A100, logged end to end.\n\nEvaluation access\n\nWe share a private Google Colab notebook that trains the same model twice on a standard GPU: once with AdamW, which runs out of memory, and once with Spectra-FFT, which completes. Access is granted to engineers and teams evaluating Spectra.\n\n``` php\n# What the notebook demonstrates\nopt = AdamW(model.parameters())    # -> CUDA out of memory\nopt = Spectra(model.parameters())  # -> trains within budget\n```\n\nRequests are reviewed manually, usually within 48 hours. Evaluation access is provided under a short NDA.\n\nFAQ\n\nYes. It is a drop-in optimizer for full-parameter training, aimed at memory-constrained GPUs.\n\nNot currently. The paper describes the method and the demo lets you verify the behavior. Evaluation and partnership access is available on request.\n\nOur logged runs track baseline convergence on the evaluated setup. See the paper and W&B logs for the exact configuration and limits.\n\nTeams fine-tuning models on-premise or on consumer and mid-range GPUs, where optimizer state is the bottleneck.", "url": "https://wpnews.pro/news/show-hn-spectra-fft-cutting-adamw-vram-usage-by-50-for-llms", "canonical_source": "https://www.spectralabs.si/", "published_at": "2026-10-05 19:30:42+00:00", "updated_at": "2026-10-05 19:50:26.796840+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-infrastructure", "mlops"], "entities": ["Spectra-FFT", "AdamW", "NVIDIA A100", "CUDA", "Google Colab", "NVIDIA", "LoRA"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-spectra-fft-cutting-adamw-vram-usage-by-50-for-llms", "markdown": "https://wpnews.pro/news/show-hn-spectra-fft-cutting-adamw-vram-usage-by-50-for-llms.md", "text": "https://wpnews.pro/news/show-hn-spectra-fft-cutting-adamw-vram-usage-by-50-for-llms.txt", "jsonld": "https://wpnews.pro/news/show-hn-spectra-fft-cutting-adamw-vram-usage-by-50-for-llms.jsonld"}}