Method
A test you can run again
A ranking is only worth something if someone else can repeat it. So we fix the conditions, record everything, and publish the files before the score.
- 1
Same start for everyone
An empty folder, the same written brief, the same machine, and a fresh install of the harness at the version its --version flag reports on the day.
- 2
Three runs, not one
Agents are not deterministic. Each contender runs three times and we show all three rings, good and bad.
- 3
Hands off
After the brief, a person only answers questions the agent asks. Every extra message is counted in the Human column.
- 4
Logs before scores
Each run produces a log: the prompt, harness and model versions, start and end time, the full transcript, the final project as a zip, and the ten check results. The log is published first; the score cell fills only after.
- 5
No sponsor, no verdict for sale
No harness maker pays for a matchup or sees results early. Links to the rest of the network sit below the results, never inside them.
Where logs live
Log files will be stored in object storage and linked from each ring. Until the first run is logged, there is nothing to download, and this page says so.
Logs published so far: 0