INDEPENDENT BENCHMARK / SEPTEMBER 2026

Astra-26

One agent. Five ways to search.

How much does the web tool stack change the answer? The same Astra model, prompt, and task set, with Exa, Keenable, Parallel, Valyu, or built-in web tools.

Codex CLI26 LiveBrowseComp tasks5 tool stacks30-call prompt budgetFormerly Astra-20
HIGHEST ACCURACY42.3%Exa & built-in web · 11/26
FASTEST MEDIAN RUN158 secExa · recorded run duration
UNSOLVED BY ALL10 / 26No correct answer from any stack
RUNS AT THE BUDGET88 / 130Stopped at exactly 30 tool calls

Answer accuracy

26 tasks · one run per task and backend

Cost calculation ↗
Custom axes
Accuracy (%)

Correct answers / all 26 tasks, including abstentions. One answer = 3.85 percentage points.

Compare the numbers

Accuracy includes every attempted task.

BackendAccuracyCost / itemMedian time
42.3%11/26 correct$1.262m 38s
42.3%11/26 correct$3.844m 17s
38.5%10/26 correct$1.302m 51s
26.9%7/26 correct$1.373m 35s
23.1%6/26 correct$1.564m 59s

Cost per item = model tokens at API rates + attributed search spend, divided by all 26 tasks. Failed runs and abstentions count toward the cost.

READING THE RESULTS

Exa and built-in web tie at 11/26 correct. Keenable is one answer behind. The table shows the cost and time behind those scores; select a backend to inspect its cost breakdown.

Cost combines a model estimate at API rates with attributed search charges; it is not a total bill. Judge costs are excluded. One run per question; small score gaps are not evidence of a general winner.

Behind the aggregate

Which tasks did each stack solve?

Each column is one task. A filled cell (+) means a correct answer.

02 / TASKS
Task IDB01B02B03B04B05B06B07B08B09B10B11B13B14B15B16B17B18B19B21B23B25B26B27B28B29B30
Exa+++++++++++
Built-in web+++++++++++
Keenable++++++++++
Parallel+++++++
Valyu++++++
3 solved by all five · 10 solved by none · 16 solved by at least one

Use of the 30-call budget

Runs that stopped at exactly 30 search or fetch calls

03 / BUDGET
Exa17/26
Built-in web11/26
Keenable20/26
Parallel19/26
Valyu21/26

The budget is instructed in the prompt. None of the published scored runs exceeded 30 calls.

What stays fixed

The experiment in four lines

04 / SETUP
Agent
GPT-6 Astra · medium
Harness
Codex CLI 0.153.4
Budget
30 calls · 1 run per task / stack
Grading
Exact match + blinded Claude Opus 5
Four vendors share search / fetch wrappers. Built-in web uses its native tools.