Introduction
Benchmarks
We publish Esan's results so you can judge them yourself — with the methodology, not just a headline number.
GAIA, in plain terms
GAIA is an independent benchmark that tests whether an agent can solve real tasks the way a person would: searching the web, reading PDFs, analyzing images, and running calculations to reach a verified answer. It's graded against a known correct answer for each question, across three difficulty levels.
Esan's results
On the GAIA validation set, Esan performs above the human baseline on Levels 1 and 2, and ties on Level 3.
Adjusted mean across levels: 92.3%; weighted overall: 93.8%.
How to read these numbers
Benchmarks are only meaningful with their method attached, so here's ours:
- The set. Results are on the GAIA validation set.
- The baseline. "Human baseline" is GAIA's own published human-performance figure for each level — the same bar every system is measured against.
- Exclusions. A small number of questions are excluded because they can't be answered as written — a source page has gone dead, the published answer has a known error, or a site actively blocks automated access. These are excluded transparently rather than counted as failures or quietly passed.
- What it measures. GAIA rewards finishing a real task correctly, not sounding plausible — which is exactly the kind of work you'd delegate to Esan.
Why we lead with this
Plenty of tools claim to be capable. A public, methodology-backed benchmark result that beats the human baseline is a claim you can check. If you care about whether an agent actually gets work done — not just how it sounds — this is the number that matters.