Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs Activation probes for monitoring deployed language models need sample counts set by coverage rather than per-concept difficulty, according to arXiv paper 2610.10594v1. Tracing learning curves over 10-590 synthetic samples across three monitoring concepts (high-stakes situations, harmful replies, and instruction non-compliance), fourteen held-out evaluation distributions, and four probe models, the authors found probes for high-stakes and harmful concepts reach within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, while instruction probes need several times as many. Concept and distribution account for 42-45% of the variance in the half-gain size, while generator LLM, probe model, and prompt detail account for under 10%, and every kind has a median half-gain size of 7-11 own-kind synthetic samples; the authors release the evaluation suites, dev sets, and generated sets. arXiv:2610.10594v1 Announce Type: new Abstract: Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open. We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail. The need is set by what is monitored: probes for high-stakes and harmful are within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, instruction probes need several times as many, and the ordering holds on three smaller probe models and on real samples from dev set . Prior work advises spending a generation budget on breadth, more kinds of data, over depth, more of each kind. We read the depth a concept needs as the half-gain size of a fitted curve, the number of samples at which half the gain is in hand. Concept and distribution account for 42-45% of its variance, the generator, probe model, and prompt detail for under 10%. What sets the value of the half-gain size is coverage, not per-kind difficulty: the number of samples of its own kind a distribution needs to saturate. Every kind, one per evaluation distribution, has a median half-gain size of 7-11 own-kind synthetic samples under all three concepts alike. What differs is how far samples of one kind transfer to the concept's other kinds, almost fully under high-stakes, less under harmful, and least under instruction, which accounts for most of the gap between concepts on generated and real samples. Breadth therefore pays differently by concept: many kinds are necessary under instruction, where no kind covers another, and nearly redundant under high-stakes, where one kind covers the rest. We release the evaluation suites, dev sets, and generated sets.