Idea for Expository AI A Hacker News user proposed an RL environment to improve frontier AI models' math explanations, where a large model teaches a small 0.5-1B parameter model to solve hard problems, with rewards based on the small model's success and human oversight to prevent leaking solutions. | |||||||||||| 1 point by | I've heard some complaints about the frontier models still be bad at explaining math and was thinking of an RL environment that would help might be to:-Take very hard math problem with a verifiable answer -Have frontier model explain to a tiny model like 0.5-1B params and provably bad score on the problem how to solve but not the solution, and reward the frontier model for prompts/explanations that helped the tiny model solve the problem Obviously some amount of human supervision is needed to weed out it giving too much information | ||||||||||| |