Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding Researchers introduced SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images, detailed in arXiv paper 2609.30682v1. SGMA combines a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence with a structure-conditioned masking process and Damped Accumulation, which aggregates signal-dependent responses across the tree into a structure canvas to guide masking. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA reached 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset (+13.00 points over the same-architecture MAE baseline) and 83.21% Dice on the 32K^2 WSI PAIP dataset (+16.84 points), with up to a 24.8x inference speedup. arXiv:2609.30682v1 Announce Type: new Abstract: Self-supervised pre-training with Vision Transformers, including Masked Autoencoders MAE , is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O N^2 $ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation DA , which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, SGMA consistently outperforms MAE baselines. It achieves 95.68% Dice on the 8K x 8K x 28K SpringXCT dataset, improving over the same-architecture MAE baseline by +13.00 points, and 83.21% Dice on the 32K^2 WSI PAIP dataset, improving by +16.84 points, while providing up to a 24.8x inference speedup.