arXiv:2610.10594v1 Announce Type: new Abstract: Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open. We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail. The need is set by what is monitored: probes for high-stakes and harmful are within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, instruction probes need several times as many, and the ordering holds on three smaller probe models and on real samples (from dev set). Prior work advises spending a generation budget on breadth, more kinds of data, over depth, more of each kind. We read the depth a concept needs as the half-gain size of a fitted curve, the number of samples at which half the gain is in hand. Concept and distribution account for 42-45% of its variance, the generator, probe model, and prompt detail for under 10%. What sets the value of the half-gain size is coverage, not per-kind difficulty: the number of samples of its own kind a distribution needs to saturate. Every kind, one per evaluation distribution, has a median half-gain size of 7-11 own-kind synthetic samples under all three concepts alike. What differs is how far samples of one kind transfer to the concept's other kinds, almost fully under high-stakes, less under harmful, and least under instruction, which accounts for most of the gap between concepts on generated and real samples. Breadth therefore pays differently by concept: many kinds are necessary under instruction, where no kind covers another, and nearly redundant under high-stakes, where one kind covers the rest. We release the evaluation suites, dev sets, and generated sets.
Coverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe Needs
Activation probes for monitoring deployed language models need sample counts set by coverage rather than per-concept difficulty, according to arXiv paper 2610.10594v1. Tracing learning curves over 10-590 synthetic samples across three monitoring concepts (high-stakes situations, harmful replies, and instruction non-compliance), fourteen held-out evaluation distributions, and four probe models, the authors found probes for high-stakes and harmful concepts reach within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, while instruction probes need several times as many. Concept and distribution account for 42-45% of the variance in the half-gain size, while generator LLM, probe model, and prompt detail account for under 10%, and every kind has a median half-gain size of 7-11 own-kind synthetic samples; the authors release the evaluation suites, dev sets, and generated sets.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.