Tencent WorkBuddy Bench – Agentic Coding Leaderboard Tencent's WorkBuddy Bench leaderboard shows no single model dominates agentic coding tasks, with Claude Opus 4.8 leading five of eight scored columns, GLM-5.2 leading two, and GPT-5.5 leading one. The open-weight GLM-5.2 tops Security under both harnesses, while GPT-5.5 achieves top-tier scores with the smallest output token budget across all subsets under CodeBuddy Code. The benchmark reveals that scores vary significantly between harnesses and that integration details, such as cross-turn reasoning passback, can lift scores by several points. Key Takeaways核心结论 - No single model dominates — column leadership splits across subsets and harnesses: Claude Opus 4.8 leads five of the eight scored columns Code on both harnesses 74.43 / 77.90 , Web on both harnesses 68.14 / 69.86 , Office on CodeBuddy Code 82.37 , GLM-5.2 two Security 76.32 / 80.86 on both , GPT-5.5 one Office 86.05 on Claude Code .没有一家通吃:各列榜首分散在不同子集与 harness 之间: Claude Opus 4.8 拿下八个计分列中的五个 Code 双 harness 74.43 / 77.90 、Web 双 harness 68.14 / 69.86 、CodeBuddy Code 下 Office 82.37 , GLM-5.2 两个 Security 双 harness 76.32 / 80.86 , GPT-5.5 一个 Claude Code 下 Office 86.05 。 - An open-weight model holds its own: GLM-5.2 tops Security under both harnesses — two of the eight columns — and stays within a point of the Code lead under Claude Code; open-weight competitiveness here is board-dependent, not uniform.开源权重模型不落下风: GLM-5.2 在两个 harness 下均领跑 Security 拿下八个计分列中的 两个 ,并在 Claude Code 下的 Code 榜上与榜首相差不到一分。开源权重的竞争力在这套基准上因榜单而异,并非全面领先。 - The harness is part of the result: the same model can swing hard between harnesses — GLM-5.2 on Web scores 67.43 vs 60.71 , GPT-5.5 on Security 77.91 vs 64.39 — so read scores per harness.harness 是结果的一部分:同一模型在不同 harness 下可能大幅波动 —— GLM-5.2 在 Web 上 67.43 vs 60.71 , GPT-5.5 在 Security 上 77.91 vs 64.39 ,分数应按 harness 分开解读。 - Efficiency does not follow rank: GPT-5.5 posts the smallest CodeBuddy Code output budget on all four subsets while staying top-tier on Code and Office, yet some mid-table runs spend 3–4× its tokens.效率与排名并不一致: GPT-5.5 在 CodeBuddy Code 下全部四个子集的输出 token 预算都是最小,并在 Code 与 Office 保持第一梯队分数,而一些中游成绩却要花 3–4 倍 的 token。 - Integration details move scores: enabling cross-turn reasoning passback for HY-3 lifted its Code score by +3.82 CodeBuddy Code / +1.92 Claude Code in a diagnostic rerun.接入细节会左右分数:在一次诊断性复跑中,为 HY-3 开启跨轮推理回传后,其 Code 分数提升 +3.82 CodeBuddy Code / +1.92 Claude Code 。 Token efficiencyToken 效率 score vs output tokens per run, 3-run avg分数 vs 每次运行输出 token,三次运行平均 Efficiency does not follow rank order: GPT-5.5 posts the smallest output budget of any model on every subset under CodeBuddy Code — 6.9k output tokens per run on Code, 13.5k on Web, 10.2k on Office, 7.5k on Security — and turns that into top-tier scores on Code and Office, plus Security under Claude Code 77.91 at 7.5k . GLM-5.2 ’s two Security leads are paid for in tokens 30–31k per run on both harnesses , and its 77.06 on Code under Claude Code costs 22.0k against GPT-5.5’s 8.7k; the smallest Claude Code point on the Code chart — Claude Opus 4.8 at 4.7k — comes from its modified-instruction run see the leaderboard notes . Security is by far the heaviest subset: MiniMax-M3 averages 88.8 turns and roughly 11.1M cache-inclusive input tokens per run under CodeBuddy Code for its 74.14. 效率与排名并不一致: GPT-5.5 在 CodeBuddy Code 下四个子集的输出预算都是全场最小(每次运行 Code 6.9k、Web 13.5k、Office 10.2k、Security 7.5k 输出 token),并在 Code、Office 以及 Claude Code 下的 Security(77.91,7.5k)拿到第一梯队分数。 GLM-5.2 的两个 Security 榜首是用 token 换来的(两个 harness 下每次运行均为 30–31k),其 Claude Code 下 Code 的 77.06 也花费 22.0k(GPT-5.5 为 8.7k);Code 图中 Claude Code 下最小的点 —— Claude Opus 4.8 的 4.7k —— 来自其调整指令的运行(见榜单说明)。Security 是 token 消耗最大的子集: MiniMax-M3 在 CodeBuddy Code 下平均每次运行 88.8 轮交互、输入 token(含缓存)约 1,110 万,最终得分 74.14。 All points are 3-run averages in think mode. Turns count unique assistant messages including subagent activity; output tokens come from each run’s final usage. Output tokens are the comparable axis — input-token accounting differs across harness configurations, so input tokens are never compared across harnesses. 所有数据点均为 3 次运行平均、think 模式。轮数按去重后的 assistant 消息计数(含 subagent 活动);输出 token 取自每次运行的最终 usage。图中仅比较输出 token,输入 token 的统计口径随 harness 配置不同,因此不做跨 harness 的输入 token 比较。 A few runs ended in task-level refusals on security-flavored requests: Claude Opus 4.8 recorded 13 under Claude Code, and GPT-5.5 2 under CodeBuddy Code. 少数运行以任务级拒答结束(针对安全类请求):Claude Opus 4.8 在 Claude Code 下 13 次,GPT-5.5 在 CodeBuddy Code 下 2 次。