Codex How To

Six controlled GPT-5.6-sol runs

Do Codex skills
save tokens? It depends.

The same engineering-loop skill lost on a small fix and won on a medium build. Explore the result, inspect the evidence, then run your own replication.

Task-size boundary

One skill. Opposite outcomes.

Medium implementation

Dependency-free 2048

All passed

Four browser-game files, ten engine tests, syntax checks, and a post-run evaluator.

No skill
828,446
Full v0.2
553,179
Lean v0.4fewest tokens
380,767
VariantTimeRetriesAccepted
No repository skill350s1Yes
Engineering loop v0.2.0257s1Yes
Lean engineering loop v0.4.0247s0Yes

What was held constant

Evidence before conclusions.

Same model, reasoning effort, starting commit, task contract, sandbox, and acceptance criteria. Only repository-skill routing changed.

01

Quality first

Acceptance, required checks, and evidence completeness were primary. A cheaper failed run would not win.

02

Three variants

No repository skill, the original v0.2.0 loop, and the lean v0.4.0 loop started from equivalent fresh copies.

03

Reported usage

Token totals are Codex CLI input plus output tokens. Cached input is already included and was not counted twice.

Read this before sharing

This is a boundary to test, not a universal claim.

Make the evidence better

Run the protocol on your task.

Fork the fixture, hold the environment constant, report every result, and publish negative findings too.