cd /news/machine-learning/structure-guided-masked-autoencoders… · home › topics › machine-learning › article
[ARTICLE · art-140766] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

Researchers introduced SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images, detailed in arXiv paper 2609.30682v1. SGMA combines a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence with a structure-conditioned masking process and Damped Accumulation, which aggregates signal-dependent responses across the tree into a structure canvas to guide masking. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA reached 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset (+13.00 points over the same-architecture MAE baseline) and 83.21% Dice on the 32K^2 WSI PAIP dataset (+16.84 points), with up to a 24.8x inference speedup.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30682v1 Announce Type: new Abstract: Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.

── more in #machine-learning 4 stories · sorted by recency
── more on @sgma 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/structure-guided-mas…] indexed:0 read:1min 2026-09-28 · —