{"slug": "tencent-workbuddy-bench-agentic-coding-leaderboard", "title": "Tencent WorkBuddy Bench – Agentic Coding Leaderboard", "summary": "Tencent's WorkBuddy Bench leaderboard shows no single model dominates agentic coding tasks, with Claude Opus 4.8 leading five of eight scored columns, GLM-5.2 leading two, and GPT-5.5 leading one. The open-weight GLM-5.2 tops Security under both harnesses, while GPT-5.5 achieves top-tier scores with the smallest output token budget across all subsets under CodeBuddy Code. The benchmark reveals that scores vary significantly between harnesses and that integration details, such as cross-turn reasoning passback, can lift scores by several points.", "body_md": "## Key Takeaways核心结论\n\n-\nNo single model dominates — column leadership splits across subsets and harnesses:\n**Claude Opus 4.8** leads five of the eight scored columns (Code on both harnesses**74.43 / 77.90**, Web on both harnesses** 68.14 / 69.86**, Office on CodeBuddy Code** 82.37**),** GLM-5.2**two (Security** 76.32 / 80.86**on both),** GPT-5.5**one (Office** 86.05**on Claude Code).没有一家通吃：各列榜首分散在不同子集与 harness 之间:** Claude Opus 4.8**拿下八个计分列中的五个(Code 双 harness** 74.43 / 77.90**、Web 双 harness** 68.14 / 69.86**、CodeBuddy Code 下 Office** 82.37**),** GLM-5.2**两个(Security 双 harness** 76.32 / 80.86**),** GPT-5.5**一个(Claude Code 下 Office** 86.05**)。 -\nAn open-weight model holds its own:\n**GLM-5.2** tops Security under both harnesses —**two of the eight columns**— and stays within a point of the Code lead under Claude Code; open-weight competitiveness here is board-dependent, not uniform.开源权重模型不落下风:**GLM-5.2**在两个 harness 下均领跑 Security(拿下八个计分列中的**两个**),并在 Claude Code 下的 Code 榜上与榜首相差不到一分。开源权重的竞争力在这套基准上因榜单而异,并非全面领先。 -\nThe harness is part of the result: the same model can swing hard between harnesses —\n**GLM-5.2** on Web scores**67.43 vs 60.71**,** GPT-5.5**on Security** 77.91 vs 64.39**— so read scores per harness.harness 是结果的一部分:同一模型在不同 harness 下可能大幅波动 ——** GLM-5.2**在 Web 上** 67.43 vs 60.71**,** GPT-5.5**在 Security 上** 77.91 vs 64.39**，分数应按 harness 分开解读。 -\nEfficiency does not follow rank:\n**GPT-5.5** posts the smallest CodeBuddy Code output budget on all four subsets while staying top-tier on Code and Office, yet some mid-table runs spend**3–4×** its tokens.效率与排名并不一致:**GPT-5.5**在 CodeBuddy Code 下全部四个子集的输出 token 预算都是最小,并在 Code 与 Office 保持第一梯队分数,而一些中游成绩却要花** 3–4 倍**的 token。 -\nIntegration details move scores: enabling cross-turn reasoning passback for\n**HY-3** lifted its Code score by**+3.82**(CodeBuddy Code) /**+1.92**(Claude Code) in a diagnostic rerun.接入细节会左右分数:在一次诊断性复跑中,为** HY-3**开启跨轮推理回传后,其 Code 分数提升**+3.82**(CodeBuddy Code)/**+1.92**(Claude Code)。\n\n## Token efficiencyToken 效率\n\nscore vs output tokens per run, 3-run avg分数 vs 每次运行输出 token，三次运行平均\n\nEfficiency does not follow rank order: **GPT-5.5** posts the smallest output budget of any model on every subset under CodeBuddy Code — 6.9k output tokens per run on Code, 13.5k on Web, 10.2k on Office, 7.5k on Security — and turns that into top-tier scores on Code and Office, plus Security under Claude Code (77.91 at 7.5k). **GLM-5.2**’s two Security leads are paid for in tokens (30–31k per run on both harnesses), and its 77.06 on Code under Claude Code costs 22.0k against GPT-5.5’s 8.7k; the smallest Claude Code point on the Code chart — Claude Opus 4.8 at 4.7k — comes from its modified-instruction run (see the leaderboard notes). Security is by far the heaviest subset: **MiniMax-M3** averages 88.8 turns and roughly 11.1M cache-inclusive input tokens per run under CodeBuddy Code for its 74.14.\n效率与排名并不一致：**GPT-5.5** 在 CodeBuddy Code 下四个子集的输出预算都是全场最小（每次运行 Code 6.9k、Web 13.5k、Office 10.2k、Security 7.5k 输出 token），并在 Code、Office 以及 Claude Code 下的 Security（77.91，7.5k）拿到第一梯队分数。**GLM-5.2** 的两个 Security 榜首是用 token 换来的（两个 harness 下每次运行均为 30–31k），其 Claude Code 下 Code 的 77.06 也花费 22.0k（GPT-5.5 为 8.7k）；Code 图中 Claude Code 下最小的点 —— Claude Opus 4.8 的 4.7k —— 来自其调整指令的运行（见榜单说明）。Security 是 token 消耗最大的子集：**MiniMax-M3** 在 CodeBuddy Code 下平均每次运行 88.8 轮交互、输入 token（含缓存）约 1,110 万，最终得分 74.14。\n\nAll points are 3-run averages in think mode. Turns count unique assistant messages including subagent activity; output tokens come from each run’s final usage. Output tokens are the comparable axis — input-token accounting differs across harness configurations, so input tokens are never compared across harnesses. 所有数据点均为 3 次运行平均、think 模式。轮数按去重后的 assistant 消息计数（含 subagent 活动）；输出 token 取自每次运行的最终 usage。图中仅比较输出 token，输入 token 的统计口径随 harness 配置不同，因此不做跨 harness 的输入 token 比较。\n\nA few runs ended in task-level refusals on security-flavored requests: Claude Opus 4.8 recorded 13 under Claude Code, and GPT-5.5 2 under CodeBuddy Code. 少数运行以任务级拒答结束（针对安全类请求）：Claude Opus 4.8 在 Claude Code 下 13 次，GPT-5.5 在 CodeBuddy Code 下 2 次。", "url": "https://wpnews.pro/news/tencent-workbuddy-bench-agentic-coding-leaderboard", "canonical_source": "https://workbuddybench.com/index.html", "published_at": "2026-07-24 02:18:47+00:00", "updated_at": "2026-07-24 02:22:10.893252+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents"], "entities": ["Tencent", "Claude Opus 4.8", "GLM-5.2", "GPT-5.5", "CodeBuddy Code", "Claude Code", "MiniMax-M3", "HY-3"], "alternates": {"html": "https://wpnews.pro/news/tencent-workbuddy-bench-agentic-coding-leaderboard", "markdown": "https://wpnews.pro/news/tencent-workbuddy-bench-agentic-coding-leaderboard.md", "text": "https://wpnews.pro/news/tencent-workbuddy-bench-agentic-coding-leaderboard.txt", "jsonld": "https://wpnews.pro/news/tencent-workbuddy-bench-agentic-coding-leaderboard.jsonld"}}