cd /news/artificial-intelligence/fourierqk-filter-shape-admissibility… · home › topics › artificial-intelligence › article
[ARTICLE · art-144281] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law

A controlled ablation on character-level language modelling with a 6-layer GPT on TinyShakespeare found that the optimal single-scale bandwidth for FourierQK bandpass-filtered attention is sigma ~= 2 bins centred at paragraph scale (~70 tokens), yielding a gain of Delta = +1.15 nats over BASE-DOT, according to the arXiv paper arXiv:2610.00009v1 by Zeris. The study reports that DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), that admissible zero-mean Mexican Hat DOG m = 2 filters outperform non-admissible Gaussians at the same scale, and that bilateral FFT leakage scales monotonically with spectral coverage, with narrowband filters (gap > +4) clean and wideband filters (gap < +2) leaky. The paper concludes FourierQK suits bidirectional encoder-style attention such as BERT, while autoregressive generation requires a causal spectral variant such as MorletQK, since causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context).

by read1 min views1 publishedOct 3, 2026

arXiv:2610.00009v1 Announce Type: new Abstract: Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about filter properties -- DC suppression, Nyquist suppression, bandwidth, centre frequency, and multi-scale coverage -- using a controlled ablation on character-level language modelling (TinyShakespeare, 6-layer GPT). Our main findings are: (1) DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), confirming that oscillatory bandpass structure is essential, not just any low-dimensional spectral summary; (2) the optimal single-scale bandwidth is sigma ~= 2 bins centred at paragraph scale (~70 tokens), giving a clean gain of Delta = +1.15 nats over BASE-DOT; (3) admissible filters (zero-mean, Mexican Hat DOG m = 2) outperform non-admissible Gaussians at the same scale and provide partial protection against bilateral FFT leakage; (4) bilateral FFT leakage scales monotonically with spectral coverage -- narrowband filters (gap > +4) are clean, wideband filters (gap < +2) are leaky; and (5) causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context), motivating word-level experiments in the companion MorletQK paper [Zeris, 2026f]. Together, findings (1)-(5) characterise FourierQK as effective in bidirectional attention settings (encoder-style, e.g. BERT), where full-sequence context is available at both training and inference time; autoregressive generation requires a causal spectral variant such as MorletQK [Zeris, 2026f] (decoder-style, e.g. GPT). Code available at: https://github.com/AthanasiosZeris/energy-gated-attention

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @fourierqk 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fourierqk-filter-sha…] indexed:0 read:1min 2026-10-03 · —