Esan
Try Esan

Introduction

Benchmarks

We publish Esan's results so you can judge them yourself — with the methodology, not just a headline number.

GAIA, in plain terms

GAIA is an independent benchmark that tests whether an agent can solve real tasks the way a person would: searching the web, reading PDFs, analyzing images, and running calculations to reach a verified answer. It's graded against a known correct answer for each question, across three difficulty levels.

Esan's results

On the GAIA validation set, Esan performs above the human baseline on Levels 1 and 2, and ties on Level 3.

Adjusted mean across levels: 92.3%; weighted overall: 93.8%.

How to read these numbers

Benchmarks are only meaningful with their method attached, so here's ours:

  • The set. Results are on the GAIA validation set.
  • The baseline. "Human baseline" is GAIA's own published human-performance figure for each level — the same bar every system is measured against.
  • Exclusions. A small number of questions are excluded because they can't be answered as written — a source page has gone dead, the published answer has a known error, or a site actively blocks automated access. These are excluded transparently rather than counted as failures or quietly passed.
  • What it measures. GAIA rewards finishing a real task correctly, not sounding plausible — which is exactly the kind of work you'd delegate to Esan.

Why we lead with this

Plenty of tools claim to be capable. A public, methodology-backed benchmark result that beats the human baseline is a claim you can check. If you care about whether an agent actually gets work done — not just how it sounds — this is the number that matters.