karimaballaarena

Matchup 01 · planned

Which coding agent wins this job?

أيّ وكيل يُنجز المهمة؟

Same brief, same machine, fresh folder, three runs each. We publish every log, and a score only appears after its log does.

Rule 1 / 3

Three runs, three rings

Agents do not give the same answer twice, so each one runs three times and we show every ring.

Rule 2 / 3

Did it work, and how fast?

Ten fixed checks decide Pass. Time is the real clock, not a benchmark number.

Rule 3 / 3

What it cost you

Money from the provider’s own usage page, and your time: every message you had to add.

FIG. 1 BRACKET, ONE RING PER RUN

  • Claude Code
  • Codex CLI
  • Hermes Agent
  • OpenClaw
  • Claude Code + DeepSeek

FIG. 2 BOX SCORE

HarnessPassTimeCostHuman
Claude Code————
Codex CLI————
Hermes Agent————
OpenClaw————
Claude Code + DeepSeek————

Dashed: not run yet. No score appears until its log file is published.

Dashed ring: not run yet. It fills only when the log is published

Pass: checks passed out of 10, such as RTL, phone width, alt text

Time: minutes from the first prompt to "done"

Cost: from the provider’s usage page, or "plan" on a subscription

Human: messages a person typed after the first prompt

No log, no score. Nothing here is estimated