{"slug": "sage-surrogate-gradient-adaptation-via-attention-guided-entropy-for-spiking", "title": "SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers", "summary": "Researchers introduced SAGE, an uncertainty-modulated surrogate-gradient mechanism for Transformer-based spiking neural networks (SNNs), which adapts the surrogate-gradient slope during training using normalized self-attention entropy without changing the inference model. Experiments on CIFAR-10/100 showed consistent accuracy gains of 1-2% over fixed-surrogate baselines across multiple simulation time steps, highlighting attention-derived uncertainty as a lightweight training signal for adaptive surrogate-gradient learning in transformer-based SNNs.", "body_md": "arXiv:2608.13702v1 Announce Type: new\nAbstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires surrogate gradients whose fixed shape may be suboptimal across layers and training stages. In this work, we introduce SAGE, an uncertainty-modulated surrogate-gradient mechanism for Transformer-based SNNs. SAGE estimates block-level uncertainty from normalized self-attention entropy and uses this signal to adapt the surrogate-gradient slope during training while leaving the inference model unchanged. By modulating only the training-time surrogate parameter, the proposed method preserves the original architecture and deployment cost while improving optimization flexibility. Experiments on CIFAR-10/100 demonstrate that SAGE achieves improved accuracy over fixed-surrogate baselines, with results up to 1-2\\% consistent gains across multiple simulation time steps. These results highlight the potential of attention-derived uncertainty as a lightweight training signal for adaptive surrogate-gradient learning in transformer-based SNNs.", "url": "https://wpnews.pro/news/sage-surrogate-gradient-adaptation-via-attention-guided-entropy-for-spiking", "canonical_source": "https://arxiv.org/abs/2608.13702", "published_at": "2026-08-17 04:00:00+00:00", "updated_at": "2026-08-17 04:12:59.937088+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "neural-networks"], "entities": ["SAGE", "CIFAR-10", "CIFAR-100"], "alternates": {"html": "https://wpnews.pro/news/sage-surrogate-gradient-adaptation-via-attention-guided-entropy-for-spiking", "markdown": "https://wpnews.pro/news/sage-surrogate-gradient-adaptation-via-attention-guided-entropy-for-spiking.md", "text": "https://wpnews.pro/news/sage-surrogate-gradient-adaptation-via-attention-guided-entropy-for-spiking.txt", "jsonld": "https://wpnews.pro/news/sage-surrogate-gradient-adaptation-via-attention-guided-entropy-for-spiking.jsonld"}}