cd /news/machine-learning/a-survey-detection-channel-overrides… · home topics machine-learning article
[ARTICLE · art-111229] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts

A new audit of the AION-1 astronomical foundation model, a 39-modality transformer trained on more than 200 million objects, finds that a survey detection channel overrides pixel information and biases tomographic mean redshifts. Holding image tokens identical and editing only the survey segmentation map changes reported quantities by 110-4400 times a matched placebo, with the effect driven by detection gating at the field centre (r = 0.47) rather than the light enclosed (r = 0.30). The Legacy Survey pipeline's 3.68% missing segments shift tomographic mean redshifts by a median 0.71 times the LSST DESC requirement, exceeding it in 12 of 40 assignments, and up to 8.3 times in the worst bin, though withholding the detection channel removes the effect at no measurable cost.

read1 min views1 publishedAug 26, 2026

arXiv:2608.23626v1 Announce Type: new Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic. We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs. Holding the image tokens byte-identical and editing only the survey segmentation map changes every quantity the model reports -- flux, size, ellipticity, redshift -- by 110-4400 times a matched placebo. The mechanism is detection gating, presence at the field centre (r = 0.47), not the light the mask encloses (r = 0.30); across 322 real blends the model ignores how the pipeline partitioned the light (R = -0.006). Nor is the preference specific to that channel: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata at all. The Legacy Survey pipeline leaves 3.68% of targets with no segment covering their position. Propagating that rate, with a miss represented by the fields the pipeline actually returns, shifts tomographic mean redshifts by a median 0.71 times the LSST DESC requirement over 40 assignments and exceeds it in 12; observed positional errors take the worst bin to 8.3 times. Drawing the misses by their measured magnitude dependence rather than uniformly does not change it. Spectroscopy removes the effect, withholding the detection channel removes it at no measurable cost, and the effect grows with model scale. Two further limits lie in the tokeniser: its image codec resolves 28 effective states on source patches against 934 for the spectrum codec, and the redshift readout is quantisation-limited. Sparse dictionaries are unreliable causal handles: across 15, recovery spans 26-75% and moves up to 18 points on the seed alone.

── more in #machine-learning 4 stories · sorted by recency
── more on @aion-1 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-survey-detection-c…] indexed:0 read:1min 2026-08-26 ·