← arena

provider specs

what you need to know before integrating — verified against each vendor’s own docs on 2026-08-26; every provider name links to the page it was checked against

providerpricing modelest. cost/querylatencyapi patterncitationsspecial capabilities
OpenAI o3 ↗
o3-deep-research
per-token$10 / $40 per 1M tok (~$1-10/query)10-30+ minResponses API, async background (poll / webhook)yes (inline annotations: url, title, start/end index)
MCP serverscode interpreterweb search$10 / 1k web-search calls
OpenAI GPT-5.5 Pro ↗
gpt-5.5-pro
per-token$30 / $180 per 1M tok (no cached rate)several min (vendor)Responses API, async background recommendedyes (inline annotations via web search tool)
web searchMCP serversfile searchcode interpreter+10% regional processing
OpenAI o4-mini ↗
o4-mini-deep-research
deprecated · off the arena roster since 2026-07-23 — jobs fail model_not_found; OpenAI still lists it, no notice
per-token$2 / $8 per 1M tok3-10 minResponses API, async backgroundyes (inline annotations)
MCP serverscode interpreterfile searchweb search
Perplexity Sonar ↗
sonar-deep-research
deprecated · Sonar endpoints retire 2026-09-27 → Agent API preset “high”
per-token + per-search$2 / $8 per 1M tok + $5 / 1k searches (~$0.30-1.30/query)1-3 minPOST /v1/sonar (sync) or /v1/async/sonar (poll)yes (citations[] URLs + search_results[] w/ snippets)
citation-focusedsearch groundingbilled per token + citation + reasoning + search
Gemini Deep Research ↗
deep-research-preview-04-2026
per-token + per-search~$1-3/task (vendor est.)minutes (60 min max)Interactions API, async (background=true; poll / stream / webhook)yes (cited report; schema not documented)
google searchurl contextcode executionMCPfile searchvisualizationmultimodal
Gemini Deep Research Max ↗
deep-research-max-preview-04-2026
per-token + per-search~$3-7/task (vendor est.)up to 60 minInteractions API, async (background=true; poll / stream / webhook)yes (cited report; schema not documented)
~160 searches/taskgoogle searchurl contextcode executionMCPfile searchvisualization
Parallel Pro ↗
pro
per-request$0.10/run2-10 minTask API, async (poll / webhook / SSE)yes (citations + excerpts + confidence, per element)
structured JSON outputauto-schema1 of 9 tiers (lite→ultra8x)failed runs not billed
Parallel Pro-Fast ↗
pro-fast
per-request$0.10/run30s-5 minTask API, async (poll / webhook / SSE)yes (citations + excerpts + confidence, per element)
structured JSON outputauto-schemafaster variant (speed over freshness)
Parallel Ultra ↗
ultra
per-request$0.30/run5-25 minTask API, async (poll / webhook / SSE)yes (citations + excerpts + confidence, per element)
structured JSON outputauto-schemadeep research modeultra2x/4x/8x at $0.60/$1.20/$2.40
Parallel Ultra-Fast ↗
ultra-fast
per-request$0.30/run1-10 minTask API, async (poll / webhook / SSE)yes (citations + excerpts + confidence, per element)
structured JSON outputauto-schemafaster variant (speed over freshness)
Valyu DeepResearch ↗
deepresearch (fast/standard/heavy/max)
per-request$0.10 / $0.50 / $2.50 / $15 per task5 min - 3 hr (vendor pages disagree)REST, async (poll / HMAC webhook)yes (sources[]: title + url)
SEC/academic/clinical/patent sourcesstructured JSON outputpython sandbox $0.10/runscreenshots $0.05/URL
You.com Research ↗
research (lite→frontier)
per-request$0.012 / $0.05 / $0.10 / $0.45 / $1.20 per call< 10s - 5 min (frontier p50 5 min)REST, sync (background=true → poll / SSE; frontier background-only)yes (inline [[n]] + sources[] w/ snippets)
structured JSON output (not lite)source control (domains/freshness)finance research variant
Exa Agent ↗
agent (effort minimal→xhigh, auto/max)
deprecated · legacy /research API (exa-research, exa-research-pro) retired by Exa — HTTP 410 RESEARCH_RETIRED; arena now runs Agent tiers
per-request$0.012 - $1.00/run (auto/max metered, $0.10/ACU)2-15 min (measured); vendor: not statedAgent API POST /agent/runs, async (poll 4s / SSE)yes (output.grounding: field-level citations + confidence)
Exa's own indexstructured output w/ groundingreturns real per-run costDollarsdeep search $12-15/1k

benchmarks

the RACE column is an independent academic benchmark (DeepResearch Bench, 2025 model versions). vendor scores are self-reported — note that DRACO is authored by Perplexity and Valyu's figures are its own re-run, so treat them as marketing, not neutral evals. every number links to its source. for a blind, community-run signal, see our arena leaderboard.