Where AI learns to do real work Labelbox launched Recursion, an agent platform that runs jobs autonomously on a schedule, alert or event and converts frequently repeated graded runs into training environments for specialist models, cutting cost per task 5–10× once a specialist takes over. Labelbox said over 90% of leading US AI labs post-train and evaluate on its platform, and that it produced the data foundation for Meta's GIM benchmark: 820 expert-authored problems across seven cognitive categories, including 229 multimodal items and 528 rubric-graded prompts. Meta released GIM-615 and an evaluation framework that benchmarked 22 models across 47 reporting configurations, finding roughly 20% of items above frontier ability. Feature Builder Builds each ready ticket into a working feature, with tests, in its own sandbox. Labelbox builds the environments frontier labs use to train AI and the platform enterprises use to put agents to work. Agents for the work that keeps coming back Recursion isn’t a chatbot. Its agents start on their own, on a schedule, an alert or an event, and work in the background until the job is done: every ready ticket built, every new bug fixed, every dependency kept current, dozens of experiments run at once. Here are a few of the jobs teams hand it. Describe the job, what starts it and what a good result looks like. From then on Recursion runs it in the background: it plans the work, brings in as many specialist agents as it needs, runs them in parallel and delivers the result into your apps, with the evidence attached. Models, sandboxes, credentials, memory and grading are handled for you, so most teams use it out of the box. 48 configs ran in parallel overnight. Config 31 beats the baseline by 2.1 points, confirmed on a rerun. Every run lands on one fleet board. Gets better, and cheaper, the more it works. Every graded run leaves notes and skills the next run starts from. When a job runs often enough, Recursion turns its graded runs into training environments, the kind Labelbox builds for frontier labs, trains a specialist model on them and switches over once it beats the model you use today. 5–10×lower cost per task once a specialist takes over The data and environments frontier models learn from Over 90% of leading US AI labs post-train and evaluate on Labelbox. Problem Meta needed a benchmark that remained discriminative as existing LLM evaluations saturated. The team wanted tasks grounded in practical reasoning rather than obscure knowledge or synthetic puzzles, with enough rubric detail to capture partial credit and enough quality control to support a public-private contamination diagnostic. Solution Labelbox produced the data foundation for GIM: 820 expert-authored problems across seven cognitive categories, including 229 multimodal items and 528 rubric-graded prompts. The work included original prompt creation, structured scoring criteria, review, expert feedback, and quality assurance, enabling Meta to calibrate a 2PL IRT model over more than 200,000 prompt-response pairs. Result Meta released GIM-615, calibrated item parameters, and an evaluation framework that benchmarked 22 models across 47 reporting configurations. The paper found GIM remains far from saturated, with roughly 20% of items above frontier ability, giving researchers a durable way to compare model capability, thinking budgets, and future systems. Our applied research team publishes the benchmarks and methods we use to train and evaluate frontier models. Each agent gets only the access its job needs, in a workspace of its own. The model never sees your passwords or keys, and every step is recorded. Your teams already ask AI for answers. Recursion takes on whole jobs: it starts on its own, works across your apps and hands back finished, graded work. Bring the job that eats your experts’ week, and we’ll set it running. Frontier AI labs already train on Labelbox environments and expert data. Tell us where your models fall short, and we’ll scope the environments, data and evals.