PopuLoRA: Co-Evolving LLM Populations for Reasoning Self- Play
PopuLoRA is a method for training large language models (LLMs) that uses co-evolving populations of teacher and student adapters to generate and solve verifiable reasoning tasks, such as code and math problems. Unlike si…