{"slug": "where-ai-learns-to-do-real-work", "title": "Where AI learns to do real work", "summary": "Labelbox launched Recursion, an agent platform that runs jobs autonomously on a schedule, alert or event and converts frequently repeated graded runs into training environments for specialist models, cutting cost per task 5–10× once a specialist takes over. Labelbox said over 90% of leading US AI labs post-train and evaluate on its platform, and that it produced the data foundation for Meta's GIM benchmark: 820 expert-authored problems across seven cognitive categories, including 229 multimodal items and 528 rubric-graded prompts. Meta released GIM-615 and an evaluation framework that benchmarked 22 models across 47 reporting configurations, finding roughly 20% of items above frontier ability.", "body_md": "### Feature Builder\n\nBuilds each ready ticket into a working feature, with tests, in its own sandbox.\n\nLabelbox builds the environments frontier labs use to train AI and the platform enterprises use to put agents to work.\n\nAgents for the work that keeps coming back\n\nRecursion isn’t a chatbot. Its agents start on their own, on a schedule, an alert or an event, and work in the background until the job is done: every ready ticket built, every new bug fixed, every dependency kept current, dozens of experiments run at once. Here are a few of the jobs teams hand it.\n\nDescribe the job, what starts it and what a good result looks like. From then on Recursion runs it in the background: it plans the work, brings in as many specialist agents as it needs, runs them in parallel and delivers the result into your apps, with the evidence attached. Models, sandboxes, credentials, memory and grading are handled for you, so most teams use it out of the box.\n\n48 configs ran in parallel overnight. Config 31 beats the baseline by 2.1 points, confirmed on a rerun.\n\nEvery run lands on one fleet board.\n\nGets better, and cheaper, the more it works.\n\nEvery graded run leaves notes and skills the next run starts from. When a job runs often enough, Recursion turns its graded runs into training environments, the kind Labelbox builds for frontier labs, trains a specialist model on them and switches over once it beats the model you use today.\n\n5–10×lower cost per task once a specialist takes over\n\nThe data and environments frontier models learn from\n\nOver 90% of leading US AI labs post-train and evaluate on Labelbox.\n\nProblem\n\nMeta needed a benchmark that remained discriminative as existing LLM evaluations saturated. The team wanted tasks grounded in practical reasoning rather than obscure knowledge or synthetic puzzles, with enough rubric detail to capture partial credit and enough quality control to support a public-private contamination diagnostic.\n\nSolution\n\nLabelbox produced the data foundation for GIM: 820 expert-authored problems across seven cognitive categories, including 229 multimodal items and 528 rubric-graded prompts. The work included original prompt creation, structured scoring criteria, review, expert feedback, and quality assurance, enabling Meta to calibrate a 2PL IRT model over more than 200,000 prompt-response pairs.\n\nResult\n\nMeta released GIM-615, calibrated item parameters, and an evaluation framework that benchmarked 22 models across 47 reporting configurations. The paper found GIM remains far from saturated, with roughly 20% of items above frontier ability, giving researchers a durable way to compare model capability, thinking budgets, and future systems.\n\nOur applied research team publishes the benchmarks and methods we use to train and evaluate frontier models.\n\nEach agent gets only the access its job needs, in a workspace of its own. The model never sees your passwords or keys, and every step is recorded.\n\nYour teams already ask AI for answers. Recursion takes on whole jobs: it starts on its own, works across your apps and hands back finished, graded work. Bring the job that eats your experts’ week, and we’ll set it running.\n\nFrontier AI labs already train on Labelbox environments and expert data. Tell us where your models fall short, and we’ll scope the environments, data and evals.", "url": "https://wpnews.pro/news/where-ai-learns-to-do-real-work", "canonical_source": "https://labelbox.com/", "published_at": "2026-10-03 04:40:10+00:00", "updated_at": "2026-10-03 05:06:09.525121+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "ai-startups", "ai-research", "mlops"], "entities": ["Labelbox", "Recursion", "Meta", "GIM", "GIM-615"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/where-ai-learns-to-do-real-work", "markdown": "https://wpnews.pro/news/where-ai-learns-to-do-real-work.md", "text": "https://wpnews.pro/news/where-ai-learns-to-do-real-work.txt", "jsonld": "https://wpnews.pro/news/where-ai-learns-to-do-real-work.jsonld"}}