{"slug": "googles-wikiskill-improves-agent-performance-across-5-benchmarks", "title": "Google’s WikiSkill improves agent performance across 5 benchmarks", "summary": "Google Research introduced WikiSkill, a framework that uses a persistent wiki-style knowledge base to let AI agents retain learned skills, boosting performance by double digits across five benchmarks. On LiveMathematicianBench, Gemini-3.5-Flash's score rose from 33.0% to 72.6%, and on SpreadSheetBench from 50.5% to 76.6%, with an average gain of 12.0 points across all benchmarks. The research, authored by Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, and Tu Vu, found that removing the wiki component largely eliminates the gains.", "body_md": "Photo: Merlin Lightpainting / Pexels\n\n# Google’s WikiSkill improves agent performance across 5 benchmarks\n\nA persistent wiki-style knowledge base lets AI agents actually remember what they've learned, boosting scores by double digits\n\nGoogle Research has introduced WikiSkill, a framework that gives AI agents the ability to learn from experience and retain that knowledge across iterations. The method uses a persistent wiki-style knowledge base to track skill improvements, and the results across five benchmarks suggest it works remarkably well.\n\nThe core problem WikiSkill addresses is that previous skill-evolution methods for AI agents would generate useful insights during execution, then discard them after each cycle rather than carrying them forward.\n\n## How WikiSkill actually works\n\nThe framework is built on a three-layer architecture, each serving a distinct purpose. The Raw Layer captures immutable execution traces, essentially a complete record of everything the agent did and what happened. The Wiki Layer consolidates that raw data into accumulated knowledge. And the Skill Layer hosts the executable procedures the agent can actually deploy.\n\nFour key components keep the system running. An inference agent handles task execution. A wiki maintainer updates the knowledge base as new information comes in. A skill proposer generates candidate procedures based on accumulated knowledge. And a validation gating mechanism acts as quality control, ensuring only genuinely useful skills make the cut.\n\n## The numbers across five benchmarks\n\nWikiSkill was validated across five benchmarks designed to test fundamentally different capabilities: math reasoning, web search, spreadsheet management, long-context question answering, and embodied interaction.\n\nThe standout results came from Gemini-3.5-Flash. On LiveMathematicianBench, performance jumped from 33.0% to 72.6%. On SpreadSheetBench, scores climbed from 50.5% to 76.6%. Across all benchmarks, the average gain was 12.0 points.\n\nAblation studies confirmed that the persistent wiki is the critical component driving performance gains. Remove it, and the improvements largely disappear.\n\nSkills learned by one model were shown to outperform self-evolved skills when applied to a different model entirely, including models from different families.\n\nThe research was authored by Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan from Google Research, and Tu Vu from Google Research and Virginia Tech.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/googles-wikiskill-improves-agent-performance-across-5-benchmarks", "canonical_source": "https://cryptobriefing.com/google-wikiskill-agent-performance-benchmarks/", "published_at": "2026-08-29 14:18:56+00:00", "updated_at": "2026-08-29 14:51:27.856242+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents"], "entities": ["Google Research", "WikiSkill", "Gemini-3.5-Flash", "LiveMathematicianBench", "SpreadSheetBench", "Liyan Tang", "Cyrus Rashtchian", "Tu Vu"], "alternates": {"html": "https://wpnews.pro/news/googles-wikiskill-improves-agent-performance-across-5-benchmarks", "markdown": "https://wpnews.pro/news/googles-wikiskill-improves-agent-performance-across-5-benchmarks.md", "text": "https://wpnews.pro/news/googles-wikiskill-improves-agent-performance-across-5-benchmarks.txt", "jsonld": "https://wpnews.pro/news/googles-wikiskill-improves-agent-performance-across-5-benchmarks.jsonld"}}