{"slug": "show-hn-do-codex-skills-save-tokens-a-six-run-task-size-benchmark", "title": "Show HN: Do Codex skills save tokens? A six-run task-size benchmark", "summary": "A six-run benchmark by an independent developer found that Codex skills saved tokens on a medium 2048 game build but lost on a small fix, with GPT-5.6-sol runs showing the same engineering-loop skill produced mixed results. The study held model, reasoning effort, starting commit, task contract, sandbox, and acceptance criteria constant, varying only repository-skill routing, and measured Codex CLI input plus output tokens with cached input included once. The author urges replication with forked fixtures and publication of negative findings.", "body_md": "Medium implementation\n\n### Dependency-free 2048\n\nFour browser-game files, ten engine tests, syntax checks, and a post-run evaluator.\n\nSix controlled GPT-5.6-sol runs\n\nThe same engineering-loop skill lost on a small fix and won on a medium build. Explore the result, inspect the evidence, then run your own replication.\n\nTask-size boundary\n\nMedium implementation\n\nFour browser-game files, ten engine tests, syntax checks, and a post-run evaluator.\n\nWhat was held constant\n\nSame model, reasoning effort, starting commit, task contract, sandbox, and acceptance criteria. Only repository-skill routing changed.\n\nAcceptance, required checks, and evidence completeness were primary. A cheaper failed run would not win.\n\nNo repository skill, the original v0.2.0 loop, and the lean v0.4.0 loop started from equivalent fresh copies.\n\nToken totals are Codex CLI input plus output tokens. Cached input is already included and was not counted twice.\n\nRead this before sharing\n\nMake the evidence better\n\nFork the fixture, hold the environment constant, report every result, and publish negative findings too.", "url": "https://wpnews.pro/news/show-hn-do-codex-skills-save-tokens-a-six-run-task-size-benchmark", "canonical_source": "https://codex-howto-benchmark.nguyenvantamdk2.chatgpt.site", "published_at": "2026-08-03 03:46:20+00:00", "updated_at": "2026-08-03 03:52:17.658565+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-tools", "ai-research"], "entities": ["Codex", "GPT-5.6-sol", "2048"], "alternates": {"html": "https://wpnews.pro/news/show-hn-do-codex-skills-save-tokens-a-six-run-task-size-benchmark", "markdown": "https://wpnews.pro/news/show-hn-do-codex-skills-save-tokens-a-six-run-task-size-benchmark.md", "text": "https://wpnews.pro/news/show-hn-do-codex-skills-save-tokens-a-six-run-task-size-benchmark.txt", "jsonld": "https://wpnews.pro/news/show-hn-do-codex-skills-save-tokens-a-six-run-task-size-benchmark.jsonld"}}