CascadeLUT: Info.-Ordered Streaming Inference for Bandwidth-Constrained FPGAs Researchers submitted CascadeLUT, an information-ordered streaming inference framework for bandwidth-constrained FPGAs, to arXiv on 1 Aug 2026. CascadeLUT partitions input features into ordered subsets and progressively refines predictions as subsets arrive, achieving 4.0 to 12.5 times lower latency, 3.0 to 5.0 times higher throughput, and up to 13.8 times lower energy per sample than prior LUT baselines, while using 1.2 to 4.4 times the LUTs of the smallest DWN baseline per task. The framework also integrates on-device input quantization with LUT-based inference, delivering 5 times reductions in quantization overhead on real-world FPGA workloads. Computer Science Hardware Architecture Submitted on 1 Aug 2026 Title:CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs View PDF /pdf/2608.00720 HTML experimental https://arxiv.org/html/2608.00720v1 Abstract:Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table LUT -based models that eliminate multipliers and map directly to reconfigurable fabric. While prior work achieves high compute efficiency, it typically assumes full-sample availability, causing pipeline stalls in bandwidth-limited streaming scenarios. Here, the bottleneck shifts from computation to data movement, as large input transfers limit throughput and energy efficiency. We present CascadeLUT, an information-structured inference framework organized around bandwidth constraints. Instead of buffering the full input, features are partitioned into ordered subsets and predictions are progressively refined as subsets arrive. The cascade statically controls which layers consume incoming features, enabling deterministic streaming inference without runtime branching. By co-designing feature scheduling with hardware dataflow, CascadeLUT reduces data movement while maintaining accuracy. Across datasets, it achieves 4.0 to 12.5 times lower latency, 3.0 to 5.0 times higher throughput and up to 13.8 times lower energy/sample than prior LUT baselines, using 1.2 to 4.4 times the LUTs of the smallest DWN baseline per task. We also demonstrate on-device input quantization integrated with LUT-based inference and present end-to-end FPGA results on real-world workloads, with 5 times reductions in quantization overhead. References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .