arXiv:2610.00251v1 Announce Type: new Abstract: Memorization audits of generative models read similarity scores against thresholds, with no null distribution, and the conclusions they support can be wrong. By MemBench's rule, the benchmark's mitigations roughly halve Stable Diffusion's memorization; audited with false-discovery control, two thirds of the certified images are no longer detected under random prompt perturbations, five sixths under attention rescaling, and all of them under embedding optimization. The field's data-copying test, read against its own null, flags ten of twenty-four generators that reproduce nothing. We argue that for memorization the null is the hard part, and supply two. For a whole model, training and held-out images are exchangeable given its samples, and relabelling them is a permutation test, exact for any statistic when the held-out images are a random split; under it, a nearest-neighbour preference still fires on seven of those twenty-four, and a count restricted to the near-duplicate scale on none (McNemar p=0.016). For single images, the natural nulls fail twice, measurably: ranking an image among random images yields 596 false discoveries among 2,365 controls, and resampling independent generations makes the null three times too narrow. Calibrated against matched controls, the audit certifies 36 of 61 MemBench images at 5% false-discovery rate, held-out controls are certified in 0.01% of calibration splits, and on this benchmark two generations per image recover that count. A calibrated maximum, which reads occasional rather than typical copying, certifies 46. As the scale-restricted statistic we recommend the small-scale mass of the Intersection Euler Characteristic Profile, which also counts distinct images copied and tests whether two models copy the same ones.
The Null Is the Hard Part: Exact Tests for Memorization in Generative Models
A new arXiv paper (2610.00251v1) argues that memorization audits of generative models are unreliable because they compare similarity scores against thresholds without a null distribution, and supplies two exact null tests. Under the paper's permutation test, a nearest-neighbour preference still fires on seven of twenty-four generators that reproduce nothing, while a count restricted to the near-duplicate scale flags none (McNemar p=0.016); calibrated against matched controls, the audit certifies 36 of 61 MemBench images at a 5% false-discovery rate, with held-out controls certified in 0.01% of calibration splits. The authors recommend the small-scale mass of the Intersection Euler Characteristic Profile as the scale-restricted statistic, which also counts distinct images copied and tests whether two models copy the same ones.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.