All benchmarksCode ↗ResearcherBench ↗

Deep Research Endpoints on RB-30

Five deep-research APIs plus a calibration anchor on 30 frozen ResearcherBench questions. 180/180 runs generated and blind-judged with the upstream pipeline’s own judges (o3-mini rubrics, gpt-4.1 citation checks), zero failures. Protocol frozen 2026-08-11, before any run.

#Model95% CI
perplexityPerplexity · calibration anchor · 77k chars avg0.7340.650.800.750.780.680.220.329,148$0.96*
1parallel ultraParallel · 30k chars avg0.7220.670.780.760.710.700.950.485,856$0.30
2valyu standardValyu · 52k chars avg0.7020.650.760.730.670.710.760.509,654$0.50
3exa highExa · 16k chars avg0.6660.600.730.720.640.640.890.362,468$0.50
4youcom exhaustiveYou.com · 5k chars avg0.5850.530.650.640.560.560.150.331,112$0.45
5tavily proTavily · 24k chars avg0.5520.460.640.620.490.550.440.383,681$0.41*

click a metric header to sort · hover a header for its definition

Sorted by coverage by default; ⚓ = calibration anchor, unranked. OC / LR / TD = per-category coverage (Open Consulting / Literature Review / Technical Details, 10 tasks each). Coverage differences under ~5 points are within noise at n=30 — the top three CIs overlap. Hover any header for its definition.

the table as a picture · shared 0–1 scale · rows in board order

coverage · with 95% CIfaithfulnessgroundednessperplexity$0.96/run*0.7340.220.32parallel ultra$0.30/run0.7220.950.48valyu standard$0.50/run0.7020.760.50exa high$0.50/run0.6660.890.36youcom exhaustive$0.45/run0.5850.150.33tavily pro$0.41/run*0.5520.440.38

shared 0–1 scale · hover any bar for exact figures, any column title for its definition

Claim compositions behind the two citation metrics, and per-task scores, are on the .

Calibration anchor

Our perplexity run scored 0.734; the published ResearcherBench figure for Perplexity Deep Research is 0.480 (full 65 questions, same judge model). Because of this +0.25 deviation, absolute scores on this page are not comparable to the published leaderboard. The judge models are the paper’s own — verified live before judging — so the likely main contributor is system version skew: the paper evaluated the consumer product circa early 2025, this run used the current sonar-deep-research API (the same harness reproduced DRB2-20’s Perplexity anchor within ~4.5 points). Subset composition may also contribute symmetrically. Within-table comparisons use identical tasks, prompts, and judges for all six systems.

Notes

n=30; coverage differences under ~5 points are noise-range. One prompt protocol, single judge pass. The same five fixed provider configs also ran DRB2-20. Full protocol, aggregation and incident log on the tab; raw runs, judging artifacts and the full analysis (INSIGHTS.md) in the repository.

Which endpoint for what

Funding

Exa provided $1,000 in API credits for this benchmark series. Exa had no input into task selection, configuration, prompts or judging — the selection rule and every config were frozen mechanically before the first API call, and all run logs are published.

30 tasks · 10 per category · 22 subjects · 428 rubric items · ~32k claims verified · frozen 2026-08-11 · · · ·