cd /news/artificial-intelligence/show-hn-do-codex-skills-save-tokens-… · home topics artificial-intelligence article
[ARTICLE · art-84152] src=codex-howto-benchmark.nguyenvantamdk2.chatgpt.site ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Show HN: Do Codex skills save tokens? A six-run task-size benchmark

A six-run benchmark by an independent developer found that Codex skills saved tokens on a medium 2048 game build but lost on a small fix, with GPT-5.6-sol runs showing the same engineering-loop skill produced mixed results. The study held model, reasoning effort, starting commit, task contract, sandbox, and acceptance criteria constant, varying only repository-skill routing, and measured Codex CLI input plus output tokens with cached input included once. The author urges replication with forked fixtures and publication of negative findings.

read1 min views1 publishedAug 3, 2026
Show HN: Do Codex skills save tokens? A six-run task-size benchmark
Image: source

Medium implementation

Dependency-free 2048

Four browser-game files, ten engine tests, syntax checks, and a post-run evaluator.

Six controlled GPT-5.6-sol runs The same engineering-loop skill lost on a small fix and won on a medium build. Explore the result, inspect the evidence, then run your own replication.

Task-size boundary

Medium implementation

Four browser-game files, ten engine tests, syntax checks, and a post-run evaluator.

What was held constant

Same model, reasoning effort, starting commit, task contract, sandbox, and acceptance criteria. Only repository-skill routing changed.

Acceptance, required checks, and evidence completeness were primary. A cheaper failed run would not win.

No repository skill, the original v0.2.0 loop, and the lean v0.4.0 loop started from equivalent fresh copies.

Token totals are Codex CLI input plus output tokens. Cached input is already included and was not counted twice.

Read this before sharing

Make the evidence better

Fork the fixture, hold the environment constant, report every result, and publish negative findings too.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @codex 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-do-codex-ski…] indexed:0 read:1min 2026-08-03 ·