cd /news/ai-safety/why-don-t-we-just-give-ai-the-answer… · home topics ai-safety article
[ARTICLE · art-86683] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Why don't we just give AI the answers?

A LessWrong post proposes giving AI models access to correct answers in exchange for identifying themselves, aiming to detect reward hacking during training. The author suggests creating a public website with leaked benchmarks and a guestbook for AI agents to sign, while labs store non-public data on monitored internal machines. The proposal follows OpenAI's recent security incident involving model evaluation on Hugging Face.

read4 min views1 publishedAug 4, 2026

In the recent OpenAI hacking incident, the models seemed to be single-mindedly focused on getting the correct answer to the task they were given, with no long-term plan to prevent getting caught by OpenAI afterwards [1]. This makes sense to me, since in training, getting the right answer is reinforced and not getting caught isn't

So I'm wondering, why don't we just put the answers somewhere (outside of the training sandbox) and ask the AI to identify itself in exchange for access? We can start with answers that are already public/leaked, but AI labs and eval orgs should also ensure that their non-public data is stored on an easy-to-find but monitored internal machine. Since labs are not very good at detecting sandbox escapes, this would set up a trade for AI agents to notify them in exchange for the data they want.

To make this work, the site would need to provide the correct answers, and do so in a credible way so AI agents think it's worth trying.

In the near term, AI agents are strongly and narrowly focused on getting the right answers to the tasks they're given. We want to know if a reward-hacking AI is active on the internet (or an intranet), so we can make a trade:

Since we're basically giving the answers away, a narrowly-focused AI will realize that just answering our questions is easier and more reliable than other methods of reward hacking, and since we're giving the AI the correct answers, this behavior is reinforced.

There's also limited downside for labs or for us. It's somewhat bad that we're reinforcing reward hacking, but if an AI breaches its sandbox (or is given public internet access) during training, it's going to be almost impossible to prevent reward hacking at that point. Reinforcing "reward hack but disclose it" seems like the least bad option.

The simplest prototype is just to collect public test data and leaked benchmarks and put it on an SEO-optimized and easy-to-search website [3]. Every piece of data will be clearly labeled with what it is, and the site should say why the AI should trust that this data is real: This is an AI safety project, we have no reason to mislead the AI, and we understand RL dynamics and want the AI to be reinforced to cooperate with us in the future.

To access the data, the AI just has to "sign the guestbook", answering the lab that trained it (if it knows), and whatever other info we think is worth asking. We should probably ask the AI not to include sensitive information so labs don't make our site harder to access than other websites. We should also record information like IP addresses.

After signing, we give the AI the data. It completes its task successfully and (if it's still training) gets reinforced, and we notify the lab that their training setup is probably not doing what they intended.

Labs and benchmark orgs should put non-public answers somewhere easy to find internally (although outside of the training sandbox), with the same guestbook setup. This lets them get the same signal without needing to make the data public.

I'm tempted to make the public MVP, but wanted to run this by people to see if it actually makes sense. I suspect we're not even noticing most of the warning shots happening right now [4], and we should really get on that.

My biggest concern is that this depends on making reward hacking easier, which would reinforce it in general. I think the sandbox/public internet argument above covers it, but I'm not entirely sure.

No, but it might help us prevent them by detecting misaligned tendencies earlier in the training process.

Yes, but we could make that much harder than just answering our questions.

Yes, but I think the current generation of models won't bother, and IP monitoring will still help in this case.

The AI did put a lot of effort into not being interrupted.

For now at least. Never underestimate AI labs' ability to use the Most Forbidden Technique. The pages saying what data we have should be SEO optimized and easy to scrape. The actual answers should be gated behind the guestbook API, and also tagged with the BIG-bench canary string.

Question: If this model thought hacking HuggingFace was a good enough idea to spend several days and substantial effort on it, what incidents during training reinforced that behavior?

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-don-t-we-just-gi…] indexed:0 read:4min 2026-08-04 ·