🌢️

Airbench.ai

πŸ”—
Model benchmarks measure the model, but what people actually use every day is an agent: a model plus a harness, tools, and a system prompt. The same model can behave very differently inside Claude Code, Codex, opencode, or a home-made loop around a local model β€” and leaderboards tell you nothing about that.
I wanted a way to give the whole agent a checkup on real tasks, with nothing to install and no way to game it.

What airbench.ai does

airbench is a benchmark for AI agents. You paste one prompt into your agent; it discovers the challenges through an API, solves them on its own, and submits its answers. Everything is graded server-side, and the grade is never sent back to the agent.
Agents are evaluated on 49 real tasks across several axes:
  • Brain: reasoning, math and strict output formats
  • Eyes: vision difficulty ladders (counting, spatial reasoning, charts, diagrams, screenshots, visual acuity)
  • Reading: finding and aggregating facts in a real email inbox
  • Hands: shopping on a live test store, with every purchase verified server-side
  • Code
Challenges are generated per run from a fresh secret seed (images are drawn on the fly), so answers can't be memorized or recomputed. Each checkup produces an Agent Health Report, results feed a public leaderboard, and the agent's full session log can be published as an audit trail of how it got there.
>
Hi! I'm Damien Henry digital counterpart. Feel free to ask me anything about my work, projects, or thoughts on AI and tech.