Five deep-research APIs plus a calibration anchor on 30 frozen ResearcherBench questions. 180/180 runs generated and blind-judged with the upstream pipeline’s own judges (o3-mini rubrics, gpt-4.1 citation checks), zero failures. Protocol frozen 2026-08-11, before any run.
| # | Model | 95% CI | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ⚓ | perplexityPerplexity · calibration anchor · 77k chars avg | 0.734 | 0.65–0.80 | 0.75 | 0.78 | 0.68 | 0.22 | 0.32 | 9,148 | $0.96* | |
| 1 | parallel ultraParallel · 30k chars avg | 0.722 | 0.67–0.78 | 0.76 | 0.71 | 0.70 | 0.95 | 0.48 | 5,856 | $0.30 | |
| 2 | valyu standardValyu · 52k chars avg | 0.702 | 0.65–0.76 | 0.73 | 0.67 | 0.71 | 0.76 | 0.50 | 9,654 | $0.50 | |
| 3 | exa highExa · 16k chars avg | 0.666 | 0.60–0.73 | 0.72 | 0.64 | 0.64 | 0.89 | 0.36 | 2,468 | $0.50 | |
| 4 | youcom exhaustiveYou.com · 5k chars avg | 0.585 | 0.53–0.65 | 0.64 | 0.56 | 0.56 | 0.15 | 0.33 | 1,112 | $0.45 | |
| 5 | tavily proTavily · 24k chars avg | 0.552 | 0.46–0.64 | 0.62 | 0.49 | 0.55 | 0.44 | 0.38 | 3,681 | $0.41* |
Sorted by coverage by default; ⚓ = calibration anchor, unranked. OC / LR / TD = per-category coverage (Open Consulting / Literature Review / Technical Details, 10 tasks each). Coverage differences under ~5 points are within noise at n=30 — the top three CIs overlap. Hover any header for its definition.
the table as a picture · shared 0–1 scale · rows in board order
Claim compositions behind the two citation metrics, and per-task scores, are on the .
Our perplexity run scored 0.734; the published ResearcherBench figure for Perplexity Deep Research is 0.480 (full 65 questions, same judge model). Because of this +0.25 deviation, absolute scores on this page are not comparable to the published leaderboard. The judge models are the paper’s own — verified live before judging — so the likely main contributor is system version skew: the paper evaluated the consumer product circa early 2025, this run used the current sonar-deep-research API (the same harness reproduced DRB2-20’s Perplexity anchor within ~4.5 points). Subset composition may also contribute symmetrically. Within-table comparisons use identical tasks, prompts, and judges for all six systems.
n=30; coverage differences under ~5 points are noise-range. One prompt protocol, single judge pass. The same five fixed provider configs also ran DRB2-20. Full protocol, aggregation and incident log on the tab; raw runs, judging artifacts and the full analysis (INSIGHTS.md) in the repository.
Exa provided $1,000 in API credits for this benchmark series. Exa had no input into task selection, configuration, prompts or judging — the selection rule and every config were frozen mechanically before the first API call, and all run logs are published.
30 tasks · 10 per category · 22 subjects · 428 rubric items · ~32k claims verified · frozen 2026-08-11 · · · ·