{"slug": "applied-compute-adds-self-distillation-workflows-to-ac2-for-production-agent", "title": "Applied Compute adds self-distillation workflows to AC2 for production agent feedback", "summary": "Applied Compute added native on-policy self-distillation (OPSD) and relevance-masked self-distillation (RMSD) workflows to its AC2 platform, enabling enterprise AI teams to convert production traces and user corrections into model-training data, the company announced on August 4th. The release supports the co-founders' bet that companies will want models that learn internal workflows, with Applied Compute reporting target behavior improvements within 10 training steps without overfitting. The company, founded by former OpenAI researchers Yash Patil, Rhythm Garg, and Linden Li, has reportedly discussed a $1.3 billion valuation.", "body_md": "[Applied Compute](https://www.appliedcompute.com/) added native on-policy self-distillation and relevance-masked self-distillation workflows to AC2, giving enterprise AI teams a way to turn production traces and user corrections into model-training data, the company said in an [August 4th thread on X](https://x.com/appliedcompute/status/2084429372128403913). The release advances the central bet made by co-founders [Yash Patil (@ypatil125)](https://x.com/ypatil125), [Rhythm Garg (@rhythmrg)](https://x.com/rhythmrg) and [Linden Li (@lindensli)](https://x.com/lindensli): companies will want models that learn their internal workflows instead of relying indefinitely on general-purpose systems. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))\n\n[https://x.com/appliedcompute/status/2084429372128403913](https://x.com/appliedcompute/status/2084429372128403913)\n\nPatil, Garg and Li started Applied Compute after working at OpenAI. Patil worked on the Codex coding effort, while Garg and Li were contributors to OpenAI's o1 system card; Garg also co-authored OpenAI research on competitive programming with reasoning models. That background has shaped Applied Compute around post-training, reinforcement learning and the infrastructure required to keep specialized models improving after deployment. ([theinformation.com](https://www.theinformation.com/articles/applied-compute-founded-ex-openai-researchers-talks-1-3-billion-valuation?utm_source=openai))\n\n### Training on the trace a model already produced\n\nProduction agents create records of failed tool calls, retries, accepted edits, user clarifications and human comments. Those records contain information about what the model should have done, but they are difficult to convert into a numerical reward suitable for conventional reinforcement learning. Applied Compute's new AC2 workflows instead train on the agent's existing trajectory. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))\n\nWith on-policy self-distillation, or OPSD, the student model sees the original prompt and its own response. A teacher version using the same weights sees that material plus an added hint, such as a correction or instruction explaining the failure. AC2 then compares the teacher's and student's token preferences and trains the student toward the behavior produced with the extra context. The method is designed for qualitative feedback and tasks that are expensive, stateful or impossible to replay multiple times. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))\n\nApplied Compute's relevance-masked variant, RMSD, narrows the update to tokens judged relevant to the desired correction. A teacher might differ from the student on wording, style and other details that have little bearing on the actual mistake. RMSD first filters for token positions with large teacher-student differences and then uses a language-model judge to choose the positions most relevant to training. AC2 exposes those choices alongside the full trace, the inserted hint and a token-level heat map. ([appliedcompute.com](https://www.appliedcompute.com/research/relevance-masked-self-distillation))\n\nThat visibility addresses a practical risk in continual learning: a run can show a declining loss while the model memorizes a narrow data set or adopts an unrelated teacher behavior. Applied Compute said it has seen a target behavior improve within 10 training steps without transferring beyond the examples used in the run. The AC2 interface is meant to let researchers inspect that failure while training is underway rather than relying solely on aggregate evaluation scores. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))\n\n### The public evidence remains controlled\n\nApplied Compute introduced RMSD in a May 22nd research report using a deliberately artificial task: training Qwen3-4B to spell \"pineapple\" incorrectly as \"pinapple\" in responses about tropical food. In that experiment, RMSD reached the target behavior in roughly half as many training steps as ordinary OPSD and took about 5% less wall-clock time, according to the company's results. It also retained stronger performance on several unrelated evaluations than supervised fine-tuning did. ([appliedcompute.com](https://www.appliedcompute.com/research/relevance-masked-self-distillation))\n\nThe test demonstrated whether the method could teach an out-of-distribution behavior while limiting damage elsewhere. It did not establish performance across a broad set of enterprise deployments. Applied Compute acknowledged that RMSD's ability to generalize across tasks, its sensitivity to the judge model and the best timing for updating teacher weights remain open research questions. The August 3rd product report adds examples involving production traces and tool-call correction, but does not publish customer-level results for the new AC2 workflows. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))\n\nAC2 supports three paths: training offline from stored production transcripts, resampling a single turn, or running a full replayable environment in which a task is graded before distillation. Applied Compute said it generally begins with stored traces to measure gains from existing production data, then moves toward resampling and online training as part of a continual-learning pipeline. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))\n\n### Applied Compute's infrastructure bet\n\nThe release gives a concrete product form to the pitch that helped Applied Compute raise $80 million on April 8th at a $1.3 billion post-money valuation. The financing was led by [Kleiner Perkins](https://www.kleinerperkins.com/perspectives/applied-compute-closing-the-gap-between-frontier-ai-and-real-world-impact/), with participation from Elad Gil, Lux Capital, Greenoaks, Neo and Hanabi. Applied Compute says the round brought its total funding to $160 million. ([appliedcompute.com](https://www.appliedcompute.com/company/fundraise))\n\nApplied Compute positions AC2 as one control plane for training, evaluation, inference, deployment and continued model improvement. Its website lists Microsoft, Nvidia, Handshake, NTT Data, DoorDash, Harvey, Cognition, Mercor and several biotechnology companies as customers or collaborators. Those relationships place AC2 against a wide set of model-training, cloud-compute and agent-observability tools, while Applied Compute's pitch combines those functions with researchers embedded alongside customers. ([appliedcompute.com](https://www.appliedcompute.com/))\n\nThe self-distillation release targets the data advantage Applied Compute is selling to those customers. Every deployed agent produces new evidence about a company's preferences and operating procedures. AC2 is designed to move that evidence from logs into model weights, with researchers inspecting which behaviors changed and whether the update introduced regressions. Turning that loop into repeatable infrastructure would make production usage itself an input to the next model version, rather than a record reviewed only after something goes wrong. ([appliedcompute.com](https://www.appliedcompute.com/platform/productionizing-self-distillation-methods))", "url": "https://wpnews.pro/news/applied-compute-adds-self-distillation-workflows-to-ac2-for-production-agent", "canonical_source": "https://runtimewire.com/article/applied-compute-ac2-self-distillation-production-traces", "published_at": "2026-08-04 05:49:22+00:00", "updated_at": "2026-08-04 05:55:44.119602+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "ai-products", "ai-infrastructure"], "entities": ["Applied Compute", "AC2", "Yash Patil", "Rhythm Garg", "Linden Li", "OpenAI"], "alternates": {"html": "https://wpnews.pro/news/applied-compute-adds-self-distillation-workflows-to-ac2-for-production-agent", "markdown": "https://wpnews.pro/news/applied-compute-adds-self-distillation-workflows-to-ac2-for-production-agent.md", "text": "https://wpnews.pro/news/applied-compute-adds-self-distillation-workflows-to-ac2-for-production-agent.txt", "jsonld": "https://wpnews.pro/news/applied-compute-adds-self-distillation-workflows-to-ac2-for-production-agent.jsonld"}}