All benchmarksCode ↗

Deep Research Endpoints on DRB2-20

Nine APIs, 20 frozen tasks, 1,504 rubrics from DeepResearch-Bench-II, judged blind by gemini-3.1-pro-preview. 180/180 runs, 0 failures.

Table 1 · Exa effort tiers

effort is the only variable

#ModelIRAnPrTotal95% CIvs best$/run

The four cheapest tiers are statistically indistinguishable from each other. Nothing separates $0.012 from $0.10 at 20 tasks. The only gaps that survive resampling are medium and high scoring above xhigh, which costs 10× the best tier.

Table 2 · Cross-provider

one flagship tier each

#ModelIRAnPrTotal95% CIvs best$/run

exa high is ahead of every other system by a margin that survives resampling, driven by InfoRecall. perplexity, valyu and parallel cannot all be separated from one another. perplexity is the anchor, not an entrant.

Blocked source, ignored

runs citing the one article the prompt forbids

ModelCitedRateTotal score
perplexity9/2045%34.09
valyu6/2030%31.35

No API-level domain filtering, so this counts instruction-following: the blocked article appearing in a run’s returned sources. Whether the judge then penalised it is a separate question. Only valyu drew evidence-level penalties, 15 rubric items worth 1.17 on its total. Perplexity drew none, so its score is the same with or without the penalty rule and its distance from the published anchor is not explained by leakage. Evidence-level detection is a lower bound, since it only catches a blocked source the judge saw a rubric item resting on.

Harness check

perplexity against its published DRB2 score

SourceIRAnPrTotalTasks
DRB2 leaderboard33.0544.4779.3438.58132
this run26.1048.3981.2934.0920
difference-6.95+3.92+1.95-4.49

Ours reads 4.49 low. Leaderboard neighbours score 39.23 and 29.89, so this sits in the same band. The miss runs in the direction that widens Exa’s lead over perplexity, so read that gap as a range rather than a point.

Funding

read alongside Table 2

Exa provided $1,000 in API credits and Exa leads Table 2. Exa had no input into task selection, configuration, prompts or judging. All of it was frozen and published before the first API call, and every raw run is in the repository for re-scoring.

20 tasks across 20 domains. 1,178 InfoRecall, 229 Analysis, 97 Presentation. Frozen 2026-08-10. · · ·