Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips Researchers introduced the Groundhog Bit-Flip Attack (GBFA), the first bit-flip-based Denial-of-Wallet availability attack against Mixture-of-Experts (MoE) large language models, which inflates decoding token usage by an average of 5912% across four real-world MoE-based LLMs by deactivating fewer than 4 experts on average. The attack targets routing-layer bits to extend output length while preserving semantic fidelity, exposing a robustness vulnerability in MoE architectures. arXiv:2608.25276v1 Announce Type: new Abstract: Mixture-of-Experts MoE architectures enable scalable and efficient large language models LLMs by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproportionately correlated with certain tokens e.g., end-of-sequence , allowing adversaries to manipulate model behavior via lightweight perturbations. In this work, we present \textbf{Groundhog Bit-Flip Attack GBFA }, the first bit-flip-based \textit{ Denial-of-Wallet availability attack} against MoE-based LLMs. By identifying and flipping routing-layer bits associated with related expert activations, we demonstrate that GBFA substantially extends the decoding token usage across three different LLM modes: conversational, reasoning, and agentic tasks, while largely preserving semantic fidelity. Across four main real-world MoE-based LLMs, manually deactivating on average fewer than \textbf{4 experts} drives average output inflation to $\mathbf{5912\%}$, with the majority of test samples reaching max tokens. These results reveal a robustness vulnerability of MoE architectures to bit flip, and highlight the potential of GBFA as an availability attack against LLMs.