Making benchmarks outputs directly useful for AI safety and security AI benchmarks are saturating fast, and Epoch's Open Problems already task AI with solving useful problems during evaluation. The author proposes extending this approach to directly optimize AI safety and security, such as having AI find and fix vulnerabilities in code or improve system performance, to leverage benchmark evaluations for tangible safety tools. Epistemic status: written in 30 min. This is not as polished as I’d like but I prefer to share this as is than not to share it at all. AI are becoming increasingly good at solving problems. Benchmarks are saturating fast. I think we should take advantage of this to make them solve useful problems while being evaluated. Epoch’s Open Problems https://epoch.ai/frontiermath/open-problems are already doing this, and this is great. It would be even better if the problems solved were directly relevant for AI safety or security. We can easily task AI to either optimize systems' performances like the nanoGPT speedrun https://github.com/KellerJordan/modded-nanogpt but on code that is useful for safety this time but also to find and fix vulnerabilities in an adversarial red and blue team setup where you have both to break others' systems and to make sure yours is solid . Examples of things that would be very useful to get In all of such cases, leveraging a benchmark evaluation to actually get better tools could help significantly during the short window we still have. I don’t have enough bandwidth to work on this myself but I’d love to see people taking care of this. Please reach out if you are interested.