Benchmark · engines on an identical spec
Same spec, three engines. Consumption differed.
12 runs, 2026-07-20 to 2026-08-10 · Updated 22 September 2026
We ran 12 controlled runs. We gave Claude Code, OpenAI Codex, and Moonshot Kimi the same fresh spec, the same starter repo, and the same merge gates. All three engines shipped a median of 3 merged pull requests where that field was recorded. Claude Code used about 4.7 times the input tokens of Codex. Kimi used about 2.4 times. This dated sample does not show equal quality, acceptance, or billed cost.
Treat these numbers as a snapshot for this harness. Check your plan, your limits, your repo, and your review process. Then make an engineering or buying choice.
What we measured
Medians across the runs that recorded each metric. The sample size differs per metric, and some older runs did not record every field. Each cell shows its own n. Token columns compare with Codex, which used the least.
| Engine | Merged PRs | Iterations | Min to 1st merge | Input tokens | Output tokens |
|---|---|---|---|---|---|
| Claude Code12 runs | 3n=7 | 8n=7 | 33n=8 | 12.1M · 4.7xn=9 | 166k · 5.7xn=9 |
| OpenAI Codex12 runs | 3n=7 | 10n=7 | 38n=8 | 2.6M · 1xn=9 | 29k · 1xn=9 |
| Moonshot Kimi11 runs | 3n=6 | 12.5n=6 | 43n=7 | 6.3M · 2.4xn=8 | 90k · 3.1xn=8 |
Start with the merged-PR column. It is the same for all three engines. The iteration and timing columns differ less than the spread between runs of the same engine. So we treat them as noise. We do not read them as a ranking. The token columns hold the one difference large enough to act on.
How the test works
One spec, invented fresh each run
We write a new spec for every run. No engine has seen it before. All three engines get the same text.
Three brand-new repositories
Each engine gets its own new project of the same kind. Each one starts from the same scaffold. Nothing is shared. Nothing carries over from an earlier run.
The real product loop
Each project runs the same loop our customers get. That means planning, dev iterations, the five merge gates, real CI, and real pull requests. There is no shortcut for this test.
Measured from the run records
We read merged pull requests, iteration counts, wall clock times, and token use from the run records. No model and no person scores them.
Curated by hand before publication
This page is a fixed snapshot taken on 2026-08-20. It is not wired to the database. It cannot show what a later run records.
Capacity limits stay on the record
An engine can stop early if it runs out of its provider allowance. We record those runs as their own outcome. They do not count as a failure.
An uneven roster holds the page
We publish an engine only when it ran the same specs as the others. We could refresh some rows. That would raise the sample for some engines and leave others behind. Each number sits beside its own sample size. So a gap would read as a ranking. An uneven set of rows is a reason to hold this snapshot.
What this does not show
- Twelve runs is a small sample. The `n` beside a number is often smaller. Some older runs did not record every field. Treat each figure as a rough guide.
- The projects are small and new. These numbers say nothing about a large or old repo.
- We ran Python libraries and web apps. Other stacks may differ.
- We ran the engines the way we drive them, inside our loop and our prompts. A different harness would give other numbers. So this is not a general claim about any vendor's model.
- Engines improve on their own schedule. A figure from this window will not hold forever, and the date above is the honest scope of it.
- Token counts and merged PR counts are not billed cost. They are not accepted work or a quality score. Twelve runs cannot support those claims about anyone's product.
Use the engine you already pay for
Keelen runs on the credentials you connect. You pick the engine for each project, and you can change it any time.
FAQ
Which AI engine is best for coding?
This page does not rank code quality. It gives a dated median of three merged pull requests, where that field was recorded. It also gives token use in one harness. Check your provider terms and try the workflow on your own repo before you choose.
Why does Claude Code use so many more tokens?
Most of the gap is input tokens. That can include context sent more than once. This snapshot cannot say why each gap shows up, or if code quality was the same. It cannot say what a provider billed. The numbers can move as tools and provider plans change.
Does the token difference cost me more money?
This benchmark does not measure bills. Your bill comes from provider pricing, plan allowances, and API rules. Keelen uses the credentials you connect. So check the provider terms for your setup.
Can I change engine per project?
Yes. The engine is a setting on each project. Point one project at the subscription you own. Point another at a cheaper key. You can change this later.
Why is GLM not in this comparison?
Keelen supports Zhipu GLM. These runs covered Claude Code, Codex, and Kimi only. We would rather leave a row out than publish a number we did not measure.