Matchup 01 · planned
Which coding agent wins this job?
أيّ وكيل يُنجز المهمة؟
Same brief, same machine, fresh folder, three runs each. We publish every log, and a score only appears after its log does.
Rule 1 / 3
Three runs, three rings
Agents do not give the same answer twice, so each one runs three times and we show every ring.
Rule 2 / 3
Did it work, and how fast?
Ten fixed checks decide Pass. Time is the real clock, not a benchmark number.
Rule 3 / 3
What it cost you
Money from the provider’s own usage page, and your time: every message you had to add.
FIG. 1 BRACKET, ONE RING PER RUN
- Claude Code
- Codex CLI
- Hermes Agent
- OpenClaw
- Claude Code + DeepSeek
FIG. 2 BOX SCORE
| Harness | Pass | Time | Cost | Human |
|---|---|---|---|---|
| Claude Code | — | — | — | — |
| Codex CLI | — | — | — | — |
| Hermes Agent | — | — | — | — |
| OpenClaw | — | — | — | — |
| Claude Code + DeepSeek | — | — | — | — |
Dashed: not run yet. No score appears until its log file is published.
Dashed ring: not run yet. It fills only when the log is published
Pass: checks passed out of 10, such as RTL, phone width, alt text
Time: minutes from the first prompt to "done"
Cost: from the provider’s usage page, or "plan" on a subscription
Human: messages a person typed after the first prompt
No log, no score. Nothing here is estimated