SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers

#SAGE #Surrogate-gradient #Adaptation #Attention-Guided #Entropy

arXiv:2608.13702v1 Announce Type: new Abstract: Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires…