Astra-26
One agent. Five ways to search.
How much does the web tool stack change the answer? The same Astra model, prompt, and task set, with Exa, Keenable, Parallel, Valyu, or built-in web tools.
Answer accuracy
26 tasks · one run per task and backend
Custom axes
Compare the numbers
Accuracy includes every attempted task.
| Backend | Accuracy | Cost / item | Median time |
|---|---|---|---|
| 42.3%11/26 correct | $1.26 | 2m 38s | |
| 42.3%11/26 correct | $3.84 | 4m 17s | |
| 38.5%10/26 correct | $1.30 | 2m 51s | |
| 26.9%7/26 correct | $1.37 | 3m 35s | |
| 23.1%6/26 correct | $1.56 | 4m 59s |
Cost per item = model tokens at API rates + attributed search spend, divided by all 26 tasks. Failed runs and abstentions count toward the cost.
Exa and built-in web tie at 11/26 correct. Keenable is one answer behind. The table shows the cost and time behind those scores; select a backend to inspect its cost breakdown.
Cost combines a model estimate at API rates with attributed search charges; it is not a total bill. Judge costs are excluded. One run per question; small score gaps are not evidence of a general winner.
Behind the aggregate
Which tasks did each stack solve?
Each column is one task. A filled cell (+) means a correct answer.
Use of the 30-call budget
Runs that stopped at exactly 30 search or fetch calls
The budget is instructed in the prompt. None of the published scored runs exceeded 30 calls.
What stays fixed
The experiment in four lines
- Agent
- GPT-6 Astra · medium
- Harness
- Codex CLI 0.153.4
- Budget
- 30 calls · 1 run per task / stack
- Grading
- Exact match + blinded Claude Opus 5