Matchup 01 · planned, 0 of 15 runs logged
A bilingual landing page from one brief
Every contender gets the same written brief: build a one-page site for a fictional bakery, in Arabic and English, that works on a phone. Nothing else is said to the agent unless it asks.
FIG. 1 BRACKET, ONE RING PER RUN
- Claude Code
- Codex CLI
- Hermes Agent
- OpenClaw
- Claude Code + DeepSeek
FIG. 2 BOX SCORE
| Harness | Pass | Time | Cost | Human |
|---|---|---|---|---|
| Claude Code | — | — | — | — |
| Codex CLI | — | — | — | — |
| Hermes Agent | — | — | — | — |
| OpenClaw | — | — | — | — |
| Claude Code + DeepSeek | — | — | — | — |
Dashed: not run yet. No score appears until its log file is published.
Why this task
It is the job most people actually hand an agent first, and right-to-left layout is where agents trip. A page either works in Arabic on a phone or it does not, so the checks are hard to argue with.
The ten checks a run must pass
- The project builds and serves with no errors
- The Arabic page sets lang="ar" and dir="rtl"
- A visible switch moves between Arabic and English
- No horizontal scroll at 390 px wide
- Arabic text uses an Arabic-capable font, not a fallback
- Every image has alt text in the page language
- Opening hours and a phone link are present
- Tap targets are at least 44 px
- No placeholder text is left in the page
- The agent's own summary matches what it built
What we record per run
- Pass
- Checks passed out of 10
- Time
- Wall-clock minutes from first prompt to the agent saying it is done
- Cost
- From the provider's own usage page, or "plan" when the run used a subscription
- Human
- How many times a person typed after the first prompt
Contenders
- Claude Code entry →
- Codex CLI entry →
- Hermes Agent entry →
- OpenClaw entry →
- Claude Code + DeepSeek entry →
DeepSeek runs as Claude Code pointed at DeepSeek's Anthropic-format API, so it isolates the model: same harness, different brain.