Exploratory study September 2026
Can models learn to search better?
Does the search tool a model practices with affect how well it later uses unfamiliar tools? In this run, search-trained models showed higher accuracy with two tools they had not used during training.
Correct answers as a percentage of resolved Claude judgments.
Keenable
Training condition Correct / resolved
- Qwen3-4B-Instruct-250743.8%64/146
- Trained without search44.3%66/149
- Trained with Exa49.7%73/147
- Trained with Serper49.3%73/148
- Trained with Perplexity Search57.0%85/149
Valyu
Training condition Correct / resolved
- Qwen3-4B-Instruct-250736.6%53/145
- Trained without search37.4%55/147
- Trained with Exa39.7%58/146
- Trained with Serper39.5%58/147
- Trained with Perplexity Search40.0%60/150
No search
Training condition Correct / resolved
- Qwen3-4B-Instruct-250718.2%27/148
- Trained without search18.2%27/148
- Trained with Exa18.7%28/150
- Trained with Serper18.1%27/149
- Trained with Perplexity Search18.1%27/149
The same 150 held-out questions in every condition. 32 unresolved Claude grades across 2,250 answers are excluded. All panels share a 0–100% scale. Rows follow the study design, not a ranking.
- 01 / practiceFour training conditions
ExaSerperPerplexity SearchNo search
Same 240 questions · 8 attempts per question · 30 updates per condition
- 02 / freezeFive model versions
Four trained versions, plus the original model as a baseline.
Qwen3-4B-Instruct-2507
- 03 / evaluateThree exam conditions
KeenableValyuNo search
Same 150 held-out questions per condition · 2,250 answers total
Keenable and Valyu were used only at exam time. The 2,250 answers are repeated evaluations of 150 unique questions across five model versions and three exam conditions.
All results
Claude · Correct / resolved answers| Training condition | Exam withKeenable | Exam withValyu | Exam withNo search |
|---|---|---|---|
| Qwen3-4B-Instruct-2507No training updates | 43.8% | 36.6% | 18.2% |
| Trained without search | 44.3% | 37.4% | 18.2% |
| Trained with Exa | 49.7% | 39.7% | 18.7% |
| Trained with Serper | 49.3% | 39.5% | 18.1% |
| Trained with Perplexity Search | 57.0% | 40.0% | 18.1% |
Inspect counts, unresolved grades, and search failures
The full-set range counts every unresolved grade as incorrect at the lower bound and correct at the upper bound. It is not a confidence interval. Search-unavailable counts flag provider failures during answer generation. Those answers remain in the results, and these counts are the same for both judges.
| Training | Exam tool | Correct | Resolved | Unresolved | Full-set range | Search unavailable |
|---|---|---|---|---|---|---|
| Qwen3-4B-Instruct-2507 | Keenable | 64 | 146 | 4 | 42.7%–45.3% | 0 |
| Qwen3-4B-Instruct-2507 | Valyu | 53 | 145 | 5 | 35.3%–38.7% | 5 |
| Qwen3-4B-Instruct-2507 | No search | 27 | 148 | 2 | 18.0%–19.3% | 0 |
| Trained without search | Keenable | 66 | 149 | 1 | 44.0%–44.7% | 0 |
| Trained without search | Valyu | 55 | 147 | 3 | 36.7%–38.7% | 6 |
| Trained without search | No search | 27 | 148 | 2 | 18.0%–19.3% | 0 |
| Trained with Exa | Keenable | 73 | 147 | 3 | 48.7%–50.7% | 0 |
| Trained with Exa | Valyu | 58 | 146 | 4 | 38.7%–41.3% | 5 |
| Trained with Exa | No search | 28 | 150 | 0 | 18.7% | 0 |
| Trained with Serper | Keenable | 73 | 148 | 2 | 48.7%–50.0% | 0 |
| Trained with Serper | Valyu | 58 | 147 | 3 | 38.7%–40.7% | 6 |
| Trained with Serper | No search | 27 | 149 | 1 | 18.0%–18.7% | 0 |
| Trained with Perplexity Search | Keenable | 85 | 149 | 1 | 56.7%–57.3% | 0 |
| Trained with Perplexity Search | Valyu | 60 | 150 | 0 | 40.0% | 4 |
| Trained with Perplexity Search | No search | 27 | 149 | 1 | 18.0%–18.7% | 0 |
How the judges compare
Judge models: claude-fable-5-1 · jev-1.13.0
Claude and Jev gave the same verdict on 95.4% of the 2,250 saved exam answers (2,147/2,250). Agreement was 96.8% among the 2,218 answers both graded correct or incorrect.
| Exam tool | Claude score | Jev score | Same verdict |
|---|---|---|---|
| Keenable | 48.8%361/739 | 46.3%347/750 | 95.1%713/750 |
| Valyu | 38.6%284/735 | 36.0%270/750 | 94.3%707/750 |
| No search | 18.3%136/744 | 18.4%138/750 | 96.9%727/750 |
What this suggests
Practice may transfer between tools
Each search-trained version scored higher than the no-search-trained control with both unfamiliar exam tools. The search environment used during training is a useful variable to investigate.
The model and exam tool work together
The size of the observed gain differs between Keenable and Valyu. Teams choosing a training tool or a deployed search tool should evaluate the combination they intend to use.
The mechanism remains open
Similar no-search exam scores are consistent with an improvement in search-assisted answering. This study does not establish whether queries, evidence selection, or answer construction improved.
How the run worked
We used Qwen3-4B-Instruct-2507 with rank-32 LoRA adapters and Dr. GRPO. Each training condition used the same 240 questions from HotpotQA and MuSiQue, with eight attempts per question and 30 updates. Questions were screened for ambiguity and difficulty before training. The exam used 150 held-out questions, split equally between those datasets. Held out refers to this experiment; exposure in the base model's pretraining data was not assessed.
Claude supplied the training rewards: correct answers received 1 and incorrect answers 0. Unresolved judgments were excluded from the within-question reward comparison; their advantages were set to zero. Claude and Jev separately graded the same 2,250 saved exam answers against reference answers using the frozen rubric. Jev was used only for the exam check and did not supply training rewards.
Estimated grading cost for the 2,250 exam answers: Claude $13.73 (API-equivalent); Jev $0.056 at its published input-token rate. Successful requests only; excludes diagnostics, failed calls and earlier attempts. Claude used a subscription; these are estimates, not verified bills.
The formal primary comparison gives HotpotQA and MuSiQue equal weight, averaging the three search-trained versions against the no-search-trained control across Keenable and Valyu. The no-search exam is a separate outcome. The selected exam judge is Claude (claude-fable-5-1).
Read the sampling uncertainty · Claude
We used 10,000 paired bootstrap draws over question families, stratified by dataset with equal dataset weights. These are 95% pointwise intervals, without a correction for multiple comparisons. They describe uncertainty over questions conditional on these trained checkpoints; they do not capture variation across training seeds.
Claude's unresolved-grade bounds are +3.6 to +6.3 percentage points. These are grading bounds, not a confidence interval. The 95% bootstrap interval for the lower endpoint is −0.5 to +7.7 percentage points; for the upper endpoint it is +2.8 to +10.2 percentage points. The conservative envelope includes zero. Repeated answers to the same question do not provide 2,250 independent question samples.
Limits to the claim
- One model, one seed per condition, 150 exam questions. The provider ordering is provisional. More seeds, models, and question sets are needed to test how reliably the pattern repeats.
- The control received less effective learning signal. Twelve of the 30 no-search training updates had zero gradient; search arms had between one and nine. The contrast includes this difference in how much useful reward variation the conditions produced.
- Judging and provider reliability remain part of the result. Claude left 32 exam answers unresolved. Valyu had four to six search-unavailable attempts per 150-answer cell. Both judges evaluated the same generated answers.
- Accuracy alone does not establish deployment value. This run does not demonstrate cost savings, latency improvements, user satisfaction, or which learned behavior caused the observed gains.
Evidence & publication status
The plan, Claude answer-judging prompt, and question-screening prompt were locally frozen with recorded hashes before the run, on 28 September 2026 at 23:00:38 UTC. Public preregistration has not been established. This is a post-run disclosure.
Operational fixes during execution expanded a case-identifier limit and handled transient or empty Valyu responses with retries and a recorded search fallback. The executed source therefore includes changes made after the initial freeze.
The public download contains aggregate cell counts, accuracy bounds, search-failure counts, and the primary comparison. It omits question text and individual answer traces. Aggregate results from both judges' analyses, dated 30 September 2026, are included separately in the downloads. The selector changes the displayed results.
Both judges · JSON Both judges · CSV