{"slug": "meta-turned-engineers-judgment-into-agent-skills", "title": "Meta turned engineers’ judgment into agent skills", "summary": "Meta's Capacity Efficiency Program used an internally built agentic platform to cut performance-regression diagnosis time from about 10 hours to around 30 minutes and recover hundreds of megawatts of power, according to Meta's engineering blog and software engineer Tommy Tran. Tran said the platform encodes senior engineers' judgment into reusable \"skills\" while keeping stable capabilities as \"tools\" built on the Model Context Protocol, with humans still approving production changes. The design separates proactive system optimization and regression remediation within a single architecture.", "body_md": "This is your last article that you can read this month before you need to [register](https://leaddev.com/register) a free LeadDev.com account.\n\nEstimated reading time: 6 minutes\n\n**Key takeaways:**\n\n- Meta encoded senior **engineers’ reasoning** into**reusable “skills,”** not just a generic tool.\n- Separating stable **“tools”** from evolving**“skills”** is what enabled the platform to scale.\n- Diagnosis time dropped from **10 hours to 30 minutes** , with humans still approving production changes.\n\nMinor performance issues here and there might not seem like the end of the world, but at scale, they really add up.\n\n“Working at a hyperscale level, small inefficiencies can compound into inefficient use of compute and power,” [Tommy Tran](https://www.linkedin.com/in/tommy-tran-ay/), software engineer at Meta, tells LeadDev. \n\nThe thing is, finding, diagnosing, and fixing performance issues can be grating and highly manual. Human expertise doesn’t easily scale to the task, but Meta is showing how [AI agents](https://leaddev.com/technical-direction/how-to-prepare-for-ai-agents) can help.\n\nIt is using an internally built agentic platform to revamp engineering operations, specifically around identifying and fixing performance regressions.\n\nAccording to [Meta’s engineering blog](https://engineering.fb.com/2026/04/16/developer-tools/capacity-efficiency-at-meta-how-unified-ai-agents-optimize-performance-at-hyperscale/), its Capacity Efficiency Program is using agents to recover hundreds of megawatts of power and reduce diagnosis time from about 10 hours to around 30 minutes. \n\nBelow, we’ll get a behind-the-scenes look at Meta’s use of agents in software infrastructure optimization.\n\n## Your inbox, upgraded.\n\nReceive weekly engineering insights to level up your leadership approach.\n\n## The initial problem\n\nBefore the agent platform was created, Meta engineers spent significant time investigating software infrastructure inefficiencies. Diagnosing issues often relied on particular senior engineers with deep expertise, who couldn’t be everywhere at once.\n\n“Before I started this project, finding and fixing those inefficiencies depended on a handful of senior efficiency engineers doing it reactively and by hand,” says Tran. Just tracing root causes within a massive interconnected system can take hours.\n\nFor instance, certain code patterns can waste CPU cycles, creating enormous waste at scale. “Before, finding and fixing those required a performance expert to manually profile services, trace hot paths, identify the anti-pattern, write a targeted fix, and verify it actually saved resources,” explains Tran.\n\nHowever, that approach could only reach a fraction of the available opportunities, and could take a specialist hours of work to discover, let alone test and remediate, says Tran. There was a clear opportunity to encode specialist knowledge of these patterns and task [AI agents](https://leaddev.com/technical-direction/why-everyones-suddenly-talking-about-ai-agents) with surfacing and suggesting fixes autonomously.\n\n“The core idea I set out to prove was that you could take the judgment those engineers apply, encode it into a platform, and make it scale,” says Tran. “That is what I built.”\n\n## Building a unified agent platform\n\nThe first design decision Tran made was to combine proactive system optimization and regression remediation within the same architecture. The second was determining when to use ‘tools’ versus ‘skills.’\n\n“Tools are the stable, reusable capabilities for observing and acting on the system,” he says, adding that he built them on [Model Context Protocol](https://leaddev.com/ai/an-engineers-guide-to-model-context-protocol-mcp) (MCP) given that it’s an open standard and helps make the tool layer interoperable.\n\nSkills, on the other hand, are where the domain expertise lives. In agentic development, [skills](https://dev.to/loc_carrre_0d798813c662/agent-skills-explained-what-they-are-what-they-arent-and-how-to-use-them-bf9) encode instructions and capabilities for agents to follow. In Meta’s platform, they capture task-specific judgment from senior engineers and are intended to evolve and be refined over time.\n\n“I build skills by sitting with senior efficiency engineers and codifying how they actually reason: the checks they run, the signals they trust, the order they investigate in,” says Tran. “It is less ‘write down an answer’ and more ‘encode a playbook.’”\n\nTran adds that keeping [MCP tools](https://leaddev.com/ai/mcp-and-the-future-of-ai-tools) and agent skills separate was a conscious architectural choice. “Keeping them separate means the plumbing stays stable while expertise evolves independently, and new problems become new skills rather than new systems.”\n\nA key element is the boundary between agent autonomy and human oversight. “The line I drew is clear,” says Tran. “The platform does the investigation and produces a ready-to-review fix, but a human engineer stays in the loop for anything that changes production.”\n\n## More like this\n\n## Energy savings and reduced engineering time\n\nBy applying this approach, Meta has seen tangible business outcomes, including quantifiable energy savings and reduced engineering time.\n\nNow, the platform handles a volume of efficiency work that would traditionally have required a large team of specialists. “A single agent run covers ground that used to take an engineer days,” says Tran. Also, it works continuously.\n\n“The outcome I care most about is that the system turned what was a reactive, one-case-at-a-time process into continuous, fleet-wide coverage operating around the clock,” says Tran. Those results validated the approach, and the same kind of gains are available to any organization managing infrastructure at scale.\n\nUsability is also enhanced. A domain expert can describe an inefficiency pattern once as a natural-language prompt, and the platform turns that into an agent that scans relevant code, identifies every instance, and produces ready-to-review code changes. That, in turn, frees up engineers for other meaningful work.\n\n“It fundamentally changed what engineers spend their time on,” he adds. “Work that used to be hours of manual investigation is now reviewing a proposed fix, which frees those engineers for harder, novel problems that actually require human creativity.”\n\n## Tips for other agentic system builders\n\nMeta’s story comes at a moment when many engineering teams are investigating how AI agent tools can benefit their workflows. Some are in the process of [building their own agentic systems](https://www.infoworld.com/article/4154570/best-practices-for-building-agentic-systems.html). \n\nHowever, teams no longer have to go into this blind: Meta’s case study sheds some light on helpful takeaways for engineering leadership.\n\n### Ground the project in a real use case\n\n“First, start with a real, painful, well-understood workflow, not a demo,” advises Tran. In his case, it was proven operational enhancements, but the same philosophy could be applied in many other areas.\n\n### Separate tools from skills\n\nSecond, consider which capabilities should remain stable tools versus evolving skills. “That single architectural decision is what lets an agent platform generalize instead of ossifying,” says Tran.\n\n### Don’t replace engineers, enhance them\n\nMeta’s agent platform didn’t remove engineers – it enhanced them. In Meta’s framework, engineers remain in control, and the goal is to remove the hours of manual investigation, not the engineer’s judgment. So, consider how new agent flows can [enhance existing talent](https://leaddev.com/ai/ai-doesnt-create-great-developers-it-amplifies-them), not replace it.\n\n### Encourage internal use\n\nGetting a project like this off the ground, especially one that changes engineering workflows, will take convincing. “What helped most was showing measurable wins on real cases rather than arguing in the abstract,” says Tran. The fact that the platform produces low-risk suggestions helped make the case, too.\n\n### Get senior engineers involved\n\nTran also got senior engineers involved early on in the process, which aided trust and benefited adoption. “Once the people whose judgment was being encoded saw the system reflect their own reasoning, they trusted it. That trust is what drives adoption. I think that principle holds anywhere you are introducing AI into an engineering workflow, not just at Meta.”\n\n### Keep humans in the loop\n\nMeta’s system does the hard work around discovery and makes the output easy to review. It doesn’t risk making production deployments without human judgment.\n\n### Measure the impact\n\nFinally, frame the results in business terms. “Measure impact in terms leadership already cares about: cost, energy, engineering time,” recommends Tran.\n\n**Berlin** • **November 9 & 10, 2026**\n\nClose the gap between what leadership expects and what’s actually possible at **LeadDev Berlin**.\n\n## Out of a few people’s heads\n\nMany large organizations have deep expert knowledge, but that knowledge doesn’t always scale across the organization. “Every field has senior experts whose judgment is trapped in their heads,” says Tran.\n\nPerhaps the most relatable outcome of this approach is the ability to share engineering knowledge through agent skills. For Tran, this offers a repeatable method for unlocking and scaling expertise across an organization.\n\nSince the skills encode expert knowledge, such a framework could also prevent [institutional knowledge decay](https://leaddev.com/ai/how-to-avoid-knowledge-decay-in-software-engineering), a common problem in software engineering.\n\n“It broke deep efficiency knowledge out of a few people’s heads and embedded it in the platform where the whole organization benefits,” says Tran. “That shift is what I think other engineering organizations will recognize in their own operations.”", "url": "https://wpnews.pro/news/meta-turned-engineers-judgment-into-agent-skills", "canonical_source": "https://leaddev.com/software-quality/meta-turned-engineers-judgment-into-agent-skills?utm_source=leaddev&utm_medium=RSS", "published_at": "2026-10-05 09:36:32+00:00", "updated_at": "2026-10-05 09:46:51.678277+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "ai-infrastructure", "mlops", "artificial-intelligence"], "entities": ["Meta", "Tommy Tran", "Capacity Efficiency Program", "Model Context Protocol", "LeadDev"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/meta-turned-engineers-judgment-into-agent-skills", "markdown": "https://wpnews.pro/news/meta-turned-engineers-judgment-into-agent-skills.md", "text": "https://wpnews.pro/news/meta-turned-engineers-judgment-into-agent-skills.txt", "jsonld": "https://wpnews.pro/news/meta-turned-engineers-judgment-into-agent-skills.jsonld"}}