cd /news/machine-learning/temper-tensorized-efficient-manifold… · home topics machine-learning article
[ARTICLE · art-91409] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

Researchers propose TEMPER, a tensorized parameterization for residual routing in deep neural networks that reduces additional parameters by about 84% compared to manifold-constrained hyper-connections (mHC) at eight residual streams while achieving the best CORE score on language modeling and commonsense reasoning tasks. The method represents generators as multi-way tensors using tensor networks, preserving manifold-constrained routing with lower parameter growth.

read1 min views1 publishedAug 11, 2026

arXiv:2608.07851v1 Announce Type: new Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expressivity of residual routing by incorporating multiple residual streams and learning dynamic information flow, while manifold-constrained (mHC) variants stabilize training through doubly stochastic residual mixing. However, a generator-level bottleneck remains in existing methods: they use dense, unstructured generators for pre-branch aggregation, residual mixing, and post-branch redistribution, which results in parameter count growing rapidly with the number of streams. To address this issue, we propose \underline{\textbf{T}}ensorized \underline{\textbf{E}}fficient \underline{\textbf{M}}anifold-constrained \underline{\textbf{P}}arameterization for \underline{\textbf{E}}xpressive Residual \underline{\textbf{R}}outing (\textbf{TEMPER}), which represents these generators as multi-way tensors over the input-stream, feature, and output-stream modes, and parameterizes them using tensor networks. Such a structured low-rank formulation is shown to preserve token-dependent manifold-constrained routing interface while substantially reducing parameter growth. It also promotes interpretability and intuition, as: i) tensor ranks control the dimensionality of the learned routing subspace, with full ranks recovering dense routing; while ii) the generator approximation errors bound differences in routing logits and, consequently, in the routed-block outputs. Comprehensive experiments show that TEMPER matches or outperforms existing methods across language modeling and commonsense reasoning tasks, while requiring substantially fewer additional parameters. At eight residual streams, TEMPER achieves the best CORE score while using about $84%$ fewer additional parameters than mHC, thus showing a stronger performance-parameter efficiency trade-off.

── more in #machine-learning 4 stories · sorted by recency
── more on @temper 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/temper-tensorized-ef…] indexed:0 read:1min 2026-08-11 ·