Nine APIs, 20 frozen tasks, 1,504 rubrics from DeepResearch-Bench-II, judged blind by gemini-3.1-pro-preview. 180/180 runs, 0 failures.
effort is the only variable
| # | Model | IR | An | Pr | Total | 95% CI | vs best | $/run | |
|---|---|---|---|---|---|---|---|---|---|
| ▸1 | exa medium | 39.25 | 50.19 | 90.50 | 44.96 | 37.1 to 52.9 | reference | $0.10 | |
| ▸2 | exa high | 39.24 | 41.85 | 91.75 | 43.84 | 35.9 to 51.6 | tie | $0.50 | |
| ▸3 | exa low | 38.86 | 44.06 | 87.00 | 43.77 | 34.9 to 52.8 | tie | $0.025 | |
| ▸4 | exa minimal | 34.80 | 46.85 | 86.75 | 40.87 | 31.9 to 50.0 | tie | $0.012 | |
| ▸5 | exa xhigh | 35.33 | 31.55 | 93.42 | 39.46 | 32.3 to 46.8 | −5.50 | $1.00 |
The four cheapest tiers are statistically indistinguishable from each other. Nothing separates $0.012 from $0.10 at 20 tasks. The only gaps that survive resampling are medium and high scoring above xhigh, which costs 10× the best tier.
one flagship tier each
| # | Model | IR | An | Pr | Total | 95% CI | vs best | $/run | |
|---|---|---|---|---|---|---|---|---|---|
| ▸1 | exa high | 39.24 | 41.85 | 91.75 | 43.84 | 35.9 to 51.6 | reference | $0.50 | |
| ▸2 | perplexity | 26.10 | 48.39 | 81.29 | 34.09 | 27.1 to 41.2 | −9.75 | $0.80 | |
| ▸3 | valyu | 25.13 | 37.17 | 78.20 | 31.35 | 23.9 to 38.5 | −12.49 | $0.50 | |
| ▸4 | parallel | 22.91 | 31.07 | 84.32 | 29.05 | 23.5 to 35.0 | −14.79 | $0.30 | |
| ▸5 | tavily | 14.46 | 22.54 | 75.40 | 20.13 | 15.3 to 25.5 | −23.71 | $0.41 |
exa high is ahead of every other system by a margin that survives resampling, driven by InfoRecall. perplexity, valyu and parallel cannot all be separated from one another. perplexity is the anchor, not an entrant.
runs citing the one article the prompt forbids
| Model | Cited | Rate | Total score | |
|---|---|---|---|---|
| perplexity | 9/20 | 45% | 34.09 | |
| valyu | 6/20 | 30% | 31.35 | |
No API-level domain filtering, so this counts instruction-following: the blocked article appearing in a run’s returned sources. Whether the judge then penalised it is a separate question. Only valyu drew evidence-level penalties, 15 rubric items worth −1.17 on its total. Perplexity drew none, so its score is the same with or without the penalty rule and its distance from the published anchor is not explained by leakage. Evidence-level detection is a lower bound, since it only catches a blocked source the judge saw a rubric item resting on.
perplexity against its published DRB2 score
| Source | IR | An | Pr | Total | Tasks |
|---|---|---|---|---|---|
| DRB2 leaderboard | 33.05 | 44.47 | 79.34 | 38.58 | 132 |
| this run | 26.10 | 48.39 | 81.29 | 34.09 | 20 |
| difference | -6.95 | +3.92 | +1.95 | -4.49 |
Ours reads 4.49 low. Leaderboard neighbours score 39.23 and 29.89, so this sits in the same band. The miss runs in the direction that widens Exa’s lead over perplexity, so read that gap as a range rather than a point.
read alongside Table 2
Exa provided $1,000 in API credits and Exa leads Table 2. Exa had no input into task selection, configuration, prompts or judging. All of it was frozen and published before the first API call, and every raw run is in the repository for re-scoring.
20 tasks across 20 domains. 1,178 InfoRecall, 229 Analysis, 97 Presentation. Frozen 2026-08-10. · · ·