cd /news/machine-learning/geopair-geometry-preserving-cross-la… · home topics machine-learning article
[ARTICLE · art-137862] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

A new arXiv paper (2609.25963v1) introduces GeoPair, a training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations to compress Transformer architectures while preserving each layer's distinct calibration geometry. The authors report that GeoPair, coupled with structured sparsity, achieves state-of-the-art results across diverse architectures, scales, and modalities, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations that rely on heuristic grouping strategies. The work claims to replace heuristic engineering with a convergent, optimization-driven pipeline for scalable Transformer compression.

by read1 min views1 publishedSep 23, 2026

arXiv:2609.25963v1 Announce Type: new Abstract: Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.

── more in #machine-learning 4 stories · sorted by recency
── more on @geopair 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/geopair-geometry-pre…] indexed:0 read:1min 2026-09-23 ·