karimaballaarena

Matchup 01 · planned, 0 of 15 runs logged

A bilingual landing page from one brief

Every contender gets the same written brief: build a one-page site for a fictional bakery, in Arabic and English, that works on a phone. Nothing else is said to the agent unless it asks.

FIG. 1 BRACKET, ONE RING PER RUN

  • Claude Code
  • Codex CLI
  • Hermes Agent
  • OpenClaw
  • Claude Code + DeepSeek

FIG. 2 BOX SCORE

HarnessPassTimeCostHuman
Claude Code————
Codex CLI————
Hermes Agent————
OpenClaw————
Claude Code + DeepSeek————

Dashed: not run yet. No score appears until its log file is published.

Why this task

It is the job most people actually hand an agent first, and right-to-left layout is where agents trip. A page either works in Arabic on a phone or it does not, so the checks are hard to argue with.

The ten checks a run must pass

  1. The project builds and serves with no errors
  2. The Arabic page sets lang="ar" and dir="rtl"
  3. A visible switch moves between Arabic and English
  4. No horizontal scroll at 390 px wide
  5. Arabic text uses an Arabic-capable font, not a fallback
  6. Every image has alt text in the page language
  7. Opening hours and a phone link are present
  8. Tap targets are at least 44 px
  9. No placeholder text is left in the page
  10. The agent's own summary matches what it built

What we record per run

Pass
Checks passed out of 10
Time
Wall-clock minutes from first prompt to the agent saying it is done
Cost
From the provider's own usage page, or "plan" when the run used a subscription
Human
How many times a person typed after the first prompt

Contenders

DeepSeek runs as Claude Code pointed at DeepSeek's Anthropic-format API, so it isolates the model: same harness, different brain.