cd /news/machine-learning/disentangling-attention-in-deep-oper… · home topics machine-learning article
[ARTICLE · art-121898] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures

A controlled study of five DeepONet variants with distinct attention mechanisms finds that per-sensor tokenization with cross-attention reduces mean relative L2 error by factors of 2.4–28.0 across all benchmark-training combinations, with the best configurations reaching 3.5–32.3 improvements. The study, posted on arXiv (2609.04407v1), evaluated models on nonlinear diffusion-reaction, viscous Burgers, and Poisson heat-conduction problems, concluding that query-dependent cross-attention is the most reliable mechanism while branch self-attention helps only with large, spatially complex inputs.

read1 min views1 publishedSep 7, 2026

arXiv:2609.04407v1 Announce Type: new Abstract: Deep neural operators learn mappings between input functions and complete PDE solution fields, enabling forward evaluations of new problem instances orders of magnitude faster than conventional numerical solvers. Attention mechanisms have recently been introduced into neural operators, but most studies change several architectural components at once, making it difficult to identify what actually improves accuracy. This work presents a controlled and systematic study of five deep operator network (DeepONet) variants with distinct attention mechanisms, trained under both data-driven and physics-informed regimes, to isolate the effects of cross-attention, self-attention, tokenization, and attention depth. We evaluate them on a source-driven transient one-dimensional nonlinear diffusion-reaction equation, a transient one-dimensional viscous Burgers equation with variable initial conditions, and a two-dimensional Poisson heat-conduction problem with heterogeneous source fields. Per-sensor tokenization with cross-attention reduces the mean relative L_2 error of the classical DeepONet in all benchmark-training combinations by factors of 2.4-28.0, while the best attention configurations reach 3.5-32.3. Branch self-attention paired only with dot-product fusion is inconsistent, degrading the one-dimensional problems while helping the more complex two-dimensional source field; added on top of cross-attention it improves all six cases, though by less than cross-attention fusion alone. Global pre-mixing provides no consistent benefit. Increasing cross-attention depth further improves accuracy, but with diminishing returns and a substantially higher cost under physics-informed training. Overall, query-dependent cross-attention is the most reliable mechanism, whereas branch self-attention is most useful for large, spatially complex functional inputs.

── more in #machine-learning 4 stories · sorted by recency
── more on @deeponet 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/disentangling-attent…] indexed:0 read:1min 2026-09-07 ·