Exploratory study September 2026

Can models learn to search better?

Does the search tool a model practices with affect how well it later uses unfamiliar tools? In this run, search-trained models showed higher accuracy with two tools they had not used during training.

Correct answers as a percentage of resolved Claude judgments.

Exam condition

Keenable

Training condition Correct / resolved

  1. Qwen3-4B-Instruct-250743.8%64/146
  2. Trained without search44.3%66/149
  3. Trained with Exa49.7%73/147
  4. Trained with Serper49.3%73/148
  5. Trained with Perplexity Search57.0%85/149
Exam condition

Valyu

Training condition Correct / resolved

  1. Qwen3-4B-Instruct-250736.6%53/145
  2. Trained without search37.4%55/147
  3. Trained with Exa39.7%58/146
  4. Trained with Serper39.5%58/147
  5. Trained with Perplexity Search40.0%60/150
Exam condition

No search

Training condition Correct / resolved

  1. Qwen3-4B-Instruct-250718.2%27/148
  2. Trained without search18.2%27/148
  3. Trained with Exa18.7%28/150
  4. Trained with Serper18.1%27/149
  5. Trained with Perplexity Search18.1%27/149

The same 150 held-out questions in every condition. 32 unresolved Claude grades across 2,250 answers are excluded. All panels share a 0–100% scale. Rows follow the study design, not a ranking.

  1. 01 / practiceFour training conditions

    ExaSerperPerplexity SearchNo search

    Same 240 questions · 8 attempts per question · 30 updates per condition

  2. 02 / freezeFive model versions

    Four trained versions, plus the original model as a baseline.

    Qwen3-4B-Instruct-2507

  3. 03 / evaluateThree exam conditions

    KeenableValyuNo search

    Same 150 held-out questions per condition · 2,250 answers total

Keenable and Valyu were used only at exam time. The 2,250 answers are repeated evaluations of 150 unique questions across five model versions and three exam conditions.

All results

Claude · Correct / resolved answers
Claude-judged accuracy among resolved answers. Every bar uses a zero to one hundred percent scale.
Training conditionExam withKeenableExam withValyuExam withNo search
Qwen3-4B-Instruct-2507No training updates43.8%36.6%18.2%
Trained without search44.3%37.4%18.2%
Trained with Exa49.7%39.7%18.7%
Trained with Serper49.3%39.5%18.1%
Trained with Perplexity Search57.0%40.0%18.1%
Claude accuracy is correct ÷ (correct + incorrect). Each cell contains 150 answers. 32 unresolved grades are excluded from these percentages. Bars share a 0–100% scale. Rows follow the study design, not a ranking.
Inspect counts, unresolved grades, and search failures

The full-set range counts every unresolved grade as incorrect at the lower bound and correct at the upper bound. It is not a confidence interval. Search-unavailable counts flag provider failures during answer generation. Those answers remain in the results, and these counts are the same for both judges.

Claude: correct, resolved, and unresolved answer counts, full-set accuracy bounds, and search-unavailable attempts for all fifteen cells.
TrainingExam toolCorrectResolvedUnresolvedFull-set rangeSearch unavailable
Qwen3-4B-Instruct-2507Keenable64146442.7%–45.3%0
Qwen3-4B-Instruct-2507Valyu53145535.3%–38.7%5
Qwen3-4B-Instruct-2507No search27148218.0%–19.3%0
Trained without searchKeenable66149144.0%–44.7%0
Trained without searchValyu55147336.7%–38.7%6
Trained without searchNo search27148218.0%–19.3%0
Trained with ExaKeenable73147348.7%–50.7%0
Trained with ExaValyu58146438.7%–41.3%5
Trained with ExaNo search28150018.7%0
Trained with SerperKeenable73148248.7%–50.0%0
Trained with SerperValyu58147338.7%–40.7%6
Trained with SerperNo search27149118.0%–18.7%0
Trained with Perplexity SearchKeenable85149156.7%–57.3%0
Trained with Perplexity SearchValyu60150040.0%4
Trained with Perplexity SearchNo search27149118.0%–18.7%0

How the judges compare

Judge models: claude-fable-5-1 · jev-1.13.0

Claude and Jev gave the same verdict on 95.4% of the 2,250 saved exam answers (2,147/2,250). Agreement was 96.8% among the 2,218 answers both graded correct or incorrect.

Claude and Jev accuracy and verdict agreement, pooled over all five model versions for each exam tool.
Exam toolClaude scoreJev scoreSame verdict
Keenable48.8%361/73946.3%347/75095.1%713/750
Valyu38.6%284/73536.0%270/75094.3%707/750
No search18.3%136/74418.4%138/75096.9%727/750
Each row pools all five model versions: 750 answers from the same 150 questions. Scores are correct ÷ resolved grades; agreement uses all 750 answers. Claude left 32 grades unresolved; the judges gave opposite correct/incorrect verdicts on 71 others. Agreement does not establish correctness or constitute an independent experimental replication.

What this suggests

Practice may transfer between tools

Each search-trained version scored higher than the no-search-trained control with both unfamiliar exam tools. The search environment used during training is a useful variable to investigate.

The model and exam tool work together

The size of the observed gain differs between Keenable and Valyu. Teams choosing a training tool or a deployed search tool should evaluate the combination they intend to use.

The mechanism remains open

Similar no-search exam scores are consistent with an improvement in search-assisted answering. This study does not establish whether queries, evidence selection, or answer construction improved.

How the run worked

We used Qwen3-4B-Instruct-2507 with rank-32 LoRA adapters and Dr. GRPO. Each training condition used the same 240 questions from HotpotQA and MuSiQue, with eight attempts per question and 30 updates. Questions were screened for ambiguity and difficulty before training. The exam used 150 held-out questions, split equally between those datasets. Held out refers to this experiment; exposure in the base model's pretraining data was not assessed.

Claude supplied the training rewards: correct answers received 1 and incorrect answers 0. Unresolved judgments were excluded from the within-question reward comparison; their advantages were set to zero. Claude and Jev separately graded the same 2,250 saved exam answers against reference answers using the frozen rubric. Jev was used only for the exam check and did not supply training rewards.

Estimated grading cost for the 2,250 exam answers: Claude $13.73 (API-equivalent); Jev $0.056 at its published input-token rate. Successful requests only; excludes diagnostics, failed calls and earlier attempts. Claude used a subscription; these are estimates, not verified bills.

The formal primary comparison gives HotpotQA and MuSiQue equal weight, averaging the three search-trained versions against the no-search-trained control across Keenable and Valyu. The no-search exam is a separate outcome. The selected exam judge is Claude (claude-fable-5-1).

Read the sampling uncertainty · Claude

We used 10,000 paired bootstrap draws over question families, stratified by dataset with equal dataset weights. These are 95% pointwise intervals, without a correction for multiple comparisons. They describe uncertainty over questions conditional on these trained checkpoints; they do not capture variation across training seeds.

Claude's unresolved-grade bounds are +3.6 to +6.3 percentage points. These are grading bounds, not a confidence interval. The 95% bootstrap interval for the lower endpoint is −0.5 to +7.7 percentage points; for the upper endpoint it is +2.8 to +10.2 percentage points. The conservative envelope includes zero. Repeated answers to the same question do not provide 2,250 independent question samples.

Limits to the claim

  • One model, one seed per condition, 150 exam questions. The provider ordering is provisional. More seeds, models, and question sets are needed to test how reliably the pattern repeats.
  • The control received less effective learning signal. Twelve of the 30 no-search training updates had zero gradient; search arms had between one and nine. The contrast includes this difference in how much useful reward variation the conditions produced.
  • Judging and provider reliability remain part of the result. Claude left 32 exam answers unresolved. Valyu had four to six search-unavailable attempts per 150-answer cell. Both judges evaluated the same generated answers.
  • Accuracy alone does not establish deployment value. This run does not demonstrate cost savings, latency improvements, user satisfaction, or which learned behavior caused the observed gains.

Evidence & publication status

The plan, Claude answer-judging prompt, and question-screening prompt were locally frozen with recorded hashes before the run, on 28 September 2026 at 23:00:38 UTC. Public preregistration has not been established. This is a post-run disclosure.

Operational fixes during execution expanded a case-identifier limit and handled transient or empty Valyu responses with retries and a recorded search fallback. The executed source therefore includes changes made after the initial freeze.

The public download contains aggregate cell counts, accuracy bounds, search-failure counts, and the primary comparison. It omits question text and individual answer traces. Aggregate results from both judges' analyses, dated 30 September 2026, are included separately in the downloads. The selector changes the displayed results.

Both judges · JSON Both judges · CSV