# AI safety prizes

> Source: <https://www.lesswrong.com/posts/rqf9LLgPn2xTvP7m6/ai-safety-prizes>
> Published: 2026-07-31 20:56:22+00:00

Rather than paying up front for AI safety research (push funding), perhaps we should pay after the fact for the work that made the most progress (pull funding). This way, you only pay for work that was actually valuable. When we know what the target is, but not how to get there or who is best placed to solve the problem, a prize is a useful incentive structure.

## Benefits of prizes

Prizes work well to incentivise innovation when:

**The eventual winner is unpredictable** (otherwise just fund the obvious choice in advance).- Trying to decide who is most likely to make a breakthrough (and therefore who to fund) is often very difficult, particularly since applicants have private information that might be hard to credibly signal.

**A valid solution is easily verifiable**, to save time and controversy when allocating the prize.- In some domains (e.g. mathematical proofs) verification is far easier than generation. Whereas in e.g. fuzzy policy work, deciding an idea is good is harder than coming up with the idea.

**You want, and can get, fanfare. **If a prize carries a big reputational boost as well as just $, then you can incentivise a lot more effort than the raw $ would justify.- Conversely, in sensitive areas, public prizes are a poor fit.

Historically, prizes have worked well in e.g. [DARPA’s autonomous vehicle challenges](https://en.wikipedia.org/wiki/DARPA_Grand_Challenge) to source diverse talent into the field.

## Costs of prizes

There are also some risks and downsides of prizes to be aware of:

**If you make a large prize with bad win conditions, people will Goodhart you!** Be very careful to specify exactly what you want, if specifying technical criteria.- And even if not, you might just distort the field’s incentives, and cause people to work on the prize topic when they otherwise might have done even more important work elsewhere.
- Alternatively, just use a panel of experts who will make a qualitative judgement on who to award the prize to, but this is more work for the judges, and is less good at incentivising the best efforts if people can
[guess the judge’s passwords](https://www.lesswrong.com/posts/NMoLJuDJEms7Ku9XS/guessing-the-teacher-s-password)/optimise for looking good rather than being good.

**Prizes disincentivise collaboration and open science.** So in cases where sharing early results, failed research directions, useful tooling and infra, etc, is very valuable, prizes suffer.- This also means that different teams will waste effort duplicating intermediate results that other teams had already achieved. Giving intermediate prizes for partial results (conditional on sharing the findings) or promising failures is helpful for this.

**Innovators may be too poor or risk-averse to pursue prizes.** It might be hard for a would-be innovator to self-finance their research on the possibility of later winning a large prize. Innovators may also be too risk-averse from a social planner’s perspective.- Combining push and pull funding can help here.

## AI-specific considerations

Prizes may be an especially good fit for the [new wave](https://nanransohoff.substack.com/p/the-third-wave-of-american-philanthropy) of AI philanthropists to fund, because:

- If investment returns are very high, paying out on a prize later is far preferable to paying up-front grants. This could easily be a >2x multiplier.
- Similarly, if we later have access to advanced AI labour to help evaluate prizes, this avoids needing to spend scarce high-context human labour evaluating grant applications now.
- There are reputational risks to some people taking funding from AI company employees (especially leadership). However, if AI company employees endow a major prize with legible criteria, administered by independent external judges, this will minimise (but not eliminate) PR worries about taking lab employee money.

Practically, what prizes could we establish?

**Compute verification advances.** This is likely the most promising category, since there are things we are confident would be helpful to have exist, and it is possible to specify in a lot of technical detail what is needed.- Most prominently, the AI Futures Project has argued for retrofitting data centers to be
[inference-only](https://ai-2040.com/supplements/verification-plan#concrete-inference-only-retrofitting-proposal). A prize could take the form of an advance market commitment for a set number of inference-only retrofitting devices, or just a working prototype meeting certain cost and quality desiderata. - There are likely many other prizes that could be funded in hardware security, cybersecurity, zero-knowledge proofs of important workload properties, etc.

**Alignment science.** I am pessimistic that we can specify in enough detail what a solution to the alignment problem (or some narrow part of it) should look like to write out a detailed technical prize here. But other more subjective options are possible:- A grand prize of perhaps billions of $ for ‘solving alignment’ with a very high bar for paying out, e.g. a diverse panel of AI safety researchers need to agree that alignment is solved with a high degree of certainty.
- But likely there will be many incremental improvements along the way, so I think one big prize at the end is not the best way to go.

- Regular ‘best paper’ prizes in a safety/alignment category at top conferences, or an annual prize. But judging such a prize would be subjective and effortful, so it may not be the best use of expert judges’ time.
- These prizes could be limited to work done outside frontier AI companies, since independent/academic/non-profit researchers will be more motivated by cash prizes.

**AI control.** Control protocols seem somewhere in-between hardware and alignment in terms of verifiability and objectivity. Plausibly there are benchmarks where we could incentivise strong performances (like ARC AGI prizes, but for control). E.g. you need to build a supervisor model using some (weak) model that successfully flags malicious actions by a (strong) misaligned model organism with some minimum reliability, using some maximum compute budget.**Interpretability. **Some problems in interpretability seem prize-amenable, in terms of being easy to verify progress on. E.g. Jane Street and Dwarkesh ran a small $50k [prize](https://colab.research.google.com/drive/1rIDPs1CtyRe9aISbwZkHLaYWxqVbOjdm) competition for working out the hidden trigger in an open-weight model they backdoored. But it also feels very possible the field will go in new directions that mean prizes specifying particular approaches or directions will be a lot less relevant in a year’s time (cf. the decline of SAEs).**AI policy.** Probably not a good fit for prizes, given the subjectivity of what a good policy project looks like.

If we were to launch one or more prizes, getting good publicity would be key. For that, the prizes should be about fun/interesting problems (as with the Jane Street/Dwarkesh one) and endorsed by fancy people (a $1B ‘Amodei-Altman Alignment Award’, anyone?).

I’d be keen to hear from people with ideas on what prizes (if any) we should be setting!

[Discuss](https://www.lesswrong.com/posts/rqf9LLgPn2xTvP7m6/ai-safety-prizes#comments)
