Research the mentors for the Cambridge Boston Alignment Initiative (CBAI) AI Safety Research Fellows...

research prompt

Research the mentors for the Cambridge Boston Alignment Initiative (CBAI) AI Safety Research Fellowship, Fall 2026 cohort (10 weeks, October 13 – December 18, 2026, Harvard Square, Cambridge MA), with the goal of helping a specific applicant decide which mentors to target and how to position himself. Do not research the program logistics; I already have those. Focus entirely on the mentors, their current research, and fit. THE MENTOR LIST (research each; names as listed by CBAI, with affiliations): Law, policy, and governance: Peter Salib — Professor of Law, University of Houston Law Center; Founder/Exec Co-Director, Center for Law & AI Risk; Law & Policy Advisor to CAIS; Visiting Senior Fellow, Institute for Law & AI; Contributing Editor at Lawfare Gabriel Weil — Assistant Professor of Law, University of Houston; Non-Resident Senior Fellow, Institute for Law & AI (liability, insurance, punitive damages, digital minds governance, AI judges) Jonathan Zittrain — Harvard Law School; Berkman Klein Center co-founder Kevin Wei — Research Scholar, GovAI; Harvard JD; ex-UK AISI science of evaluations team (science of evals, legal alignment evaluations, US AI governance/law) Stephen Casper — Assistant Professor of Public Policy, Harvard Kennedy School (open-weight model safety and tamper resistance, predicting/preventing AI incidents, technical loopholes in frontier AI laws) Michael Chen — AI Science Advisor, California Governor's Office of Emergency Services; formerly METR Evaluations, auditing, and standards: Nikola Jurkovic — Member of Technical Staff, METR (eval development, eval execution, risk assessment) Patricia Paskov — Director of Standards, AVERI; Adjunct Researcher at RAND; Oxford DPhil candidate Sean McGregor — Co-founder AVERI; founder, AI Incident Database Charles Teague — CEO, Meridian Labs; formerly RAND, helped develop Inspect AI Peter Slattery, Alexander Saeri, and Michael Noetel — MIT AI Risk Initiative / MIT FutureTech (AI Risk Repository, AI Incident Tracker, AI Governance Map, Delphi studies) Technical safety and interpretability: David Bau and Benno Krojer — Northeastern University (Bau Lab; multimodal interpretability; Krojer is also active in science communication) Dan Braun — Goodfire (parameter decomposition, interpretable-by-design networks) Oliver Clive-Griffin — Goodfire (targeted parameter decomposition) Hadas Orgad — Kempner Institute, Harvard (interpretability of harmful behaviors, hallucination, sycophancy, deception, situational awareness) Shi Feng — PI at Praxis Research; Assistant Professor, GWU (model organisms of deception, intent misalignment) Adam Shai and Paul Riechers — Simplex / Astera (belief-state geometry, computational mechanics) Zach Furman — Iliad (singular learning theory, training dynamics) Hidenori Tanaka — Harvard Physics of Intelligence (persona mechanics, multi-agent collective intelligence) Dylan Hadfield-Menell — MIT CSAIL (Algorithmic Alignment Group) James Mickens — Harvard (systems-level sandboxing of misaligned AI, alignment red-teaming) Samuel Gunn — RESI (watermarking, data attribution) Adam Tauman Kalai — Researcher/CSO, RESI; formerly OpenAI (theoretical alignment, safe-by-design) Jay Chooi — CEO/co-founder, Robocurve (robotics benchmarks) FOR EACH MENTOR, FIND: Their recent publications and preprints (2025–2026 especially), with titles, venues, and arXiv links where available. What are they actually working on right now, as opposed to what their bio says? Their track record mentoring junior researchers: have they supervised at MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or previous CBAI cohorts? Did those mentees publish? Any public accounts from former mentees about what they're like to work with? Whether they have publicly stated preferences about the kind of person they work well with, beyond what's on the CBAI page. How much biology, law, or cybersecurity background their projects presuppose. Whether they are hiring or have a pipeline from fellowship to full-time role. PRIORITIZE DEPTH ON THESE SEVEN, since they are the likeliest matches: Kevin Wei, Patricia Paskov, the MIT AI Risk Initiative trio (Slattery/Saeri/Noetel), Gabriel Weil, Peter Salib, Nikola Jurkovic, and Stephen Casper. For these, also find: Kevin Wei's published work on the science of evaluations and legal alignment evaluations specifically. What is "law-following AI"? What has he published on model spec and AI constitution compliance? How does his GovAI work relate to his CBAI mentorship? Patricia Paskov's work on frontier AI auditing and the AI Evaluators Forum. What is AVERI actually building, and what would a "PCAOB/FINRA analogue for AI auditing" involve? The MIT AI Risk Initiative's classification methodology: how do they do LLM-plus-human-validated taxonomy classification, what does the AI Risk Mitigation Taxonomy look like, and what tooling do they use? Gabriel Weil's liability scholarship, particularly his argument about punitive damages and uninsurable risk. What is the strongest published counterargument to his position? Peter Salib's "law as an alignment technology" framework and the AI rights argument in his forthcoming Cambridge UP book. Note that his CBAI mentor topics currently read "More information soon" — check whether anything has been posted since. Nikola Jurkovic's METR work, especially the time-horizons research, and what "eval development, eval execution, and risk assessment" means concretely at METR. Stephen Casper's work on open-weight model tamper resistance and on technical loopholes in frontier AI laws. He was a BioSafeGenAI best paper runner-up; find that paper. THE APPLICANT (tailor all fit assessments to him): US citizen in Fort Worth, Texas. Finishing an MS in Artificial Intelligence at Angelo State University (May 2027). Graduate Research Assistant on a DoD/Army Research Laboratory cooperative agreement. Three IEEE ICAIC 2026 publications, first author on two, including a Best Research Paper Award for reformulating cyber-threat-intelligence-to-MITRE-ATT&CK mapping from classification to retrieval, reaching 95.21% precision-at-1 across 600+ techniques. A fourth paper under review at ACSAC 2026 and a journal article in preparation. He has since found that the field-standard evaluation protocol in that literature inflates measured accuracy roughly threefold relative to deployment conditions (0.784 down to 0.253 on identical model and queries), and built an attribution-based audit showing his model does not rely on threat-actor-name shortcuts. He independently built STING, a CPU-only, fully offline document-analysis pipeline with leave-one-out token attribution. Separately he built a four-layer LLM output validation engine and found that using one model to monitor another fails on the same inputs the generator fails on. Before AI he completed three years of medical school coursework (virology, immunology, pharmacology, pathology). Completed BlueDot Biosecurity Fundamentals; partway through BlueDot Technical AI Safety. Won first place in the Oracle track of a National AI Prompt Engineering Championship regional. Reached the final stage of a previous CBAI cycle. Currently in the second round of the GovAI DC Winter Fellowship. Critically: his stated long-term direction is law school after the MS, then working at the intersection of AI, law, and policy, potentially including science communication, domestically or internationally. He is explicit that his interest is not only technical but extends to AI's ethical and societal consequences. Weight law, governance, standards, and public-communication-adjacent mentors accordingly, and note that Benno Krojer is described as active in science communication. QUESTIONS TO ANSWER: Rank the mentors into tiers for this applicant, with reasoning. Which three should he name as top choices, and which is the single strongest? Where does his evaluation-validity work (the inflation-chain finding, the shortcut audit, the correlated-monitor-failure result) map most directly onto a named mentor's current agenda? Which mentors would treat the law school trajectory as a genuine asset rather than a distraction, and which would see it as a sign he's leaving technical work? He is applying to both the AI Safety and AIxBiosecurity tracks on one combined application. Given the AIxBio mentors named in CBAI's email (Active Site, SecureBio, and the RAND AI-Bio Evals team), how should he sequence or split his mentor preferences across the two tracks without looking unfocused? What test tasks or code screens do these specific mentors use, if that's publicly known? Are there mentors whose stated requirements he clearly does not meet, so he can avoid wasting a ranking slot? Provide specific sources, links, and dates. Flag where a mentor's Fall 2026 project list is still marked "More information soon" and therefore unknowable. Distinguish clearly between what is documented and what is inference.

date
Sep 4, 2026
direct compareParallelExaValyu
metricParallelExaValyu
formatproseproseprose
word count8,3503,9365,755
sources1181880
processing time1288s0s0s
has imagesnonono
has tablesnonono
citation style

Exa

prose3,936 words

Bottom line

Name these three, in order: (1) Kevin Wei, (2) Patricia Paskov, (3) Stephen Casper.
Single strongest: Kevin Wei.

That ordering follows from the applicant’s unusually direct evidence of evaluation validity (deployment-vs-paper protocol gap, shortcut/attribution audit, and correlated monitor failure), his demonstrated ability to build reproducible offline tooling, and his explicit JD-to-AI-law/policy goal. Wei is the rare mentor whose published agenda joins all three: evaluation methodology, legal alignment, and technical governance. CBAI specifically names ecological validity, eval awareness, model-spec/constitution compliance, legal hierarchy, legal retrieval, deterrence, certification, and U.S.–China governance as his project space; it requires Python/ML/research experience for evals, which the applicant plainly has. His legal-alignment lane formally prefers a law degree in progress, so frame the applicant as an MS researcher deliberately building the empirical record before the JD—not as already qualified for a doctrinal-law project. CBAI—Wei RAND profile, updated Apr. 18 2026

Recommended ranking

Tier 1 — target aggressively
  1. Kevin Wei — best overall.

    • Why: His ICML 2025 work on rigorous human baselines, ICML-workshop 2025 work on methodological problems in agentic evaluations, and AAAI 2026 work on GPAI incident-reporting systems make the applicant’s result—standard protocol inflates accuracy from 0.784 to 0.253—immediately legible as evaluation-science evidence rather than merely a cyber benchmark result. RAND
    • Pitch: “I found an externally invalid evaluation pipeline, demonstrated a 3.1× deployment inflation, and then tested a causal explanation (actor-name shortcuts) rather than reporting a leaderboard score. I want to generalize this into an evaluation-validity/audit-card protocol for agentic legal or cyber systems, including generator–monitor correlated failure.”
    • Best project: Build a legal-agent or compliance-eval validity framework: operationalize legal issue spotting, retrieval, hierarchy resolution and abstention; compare conventional benchmark scores against realistic deployment distributions; publish failure modes and calibration guidance.
    • Law-school signal: genuine asset. Wei has a Harvard JD and explicitly offers legal-alignment projects, although CBAI says candidates in that lane should have or be pursuing a JD/LLB/LLM. The applicant should select Wei primarily under science of evaluations, while saying the JD is the planned translation path.
  2. Patricia Paskov — best standards/auditing fit.

    • Why: The applicant’s work is an audit of whether a claimed model capability is real under deployment conditions. That maps tightly to Paskov’s current focus: evaluation science, standards, assurance and third-party auditing; she is AVERI’s Director of Standards, a RAND adjunct, Oxford DPhil candidate, and lead of the Resilience chapter of the 2026 International AI Safety Report. Paskov bio
    • Pitch: “I can turn my protocol-inflation, shortcut-audit and monitor-correlation findings into an auditable reporting standard: validity claims, threat model, distribution shift, confounder tests, and a reproducible artifact.” That is more distinctive than pitching generic AI governance.
    • Law-school signal: a major asset, particularly if presented as a route to assurance institutions, evidentiary standards, liability and public communication—not as departure from technical rigor.
  3. Stephen Casper — strongest technical-governance/bio bridge.

    • Why: He explicitly offers open-weight tamper resistance, AI incidents, and legal/technical loopholes; he also mentors MATS, ERA and GovAI. His stated preference—tenacity, self-taught skills, and initiative—matches a self-built CPU-only attribution pipeline and independent protocol audit unusually well. CBAI—Casper personal site
    • Pitch: lead with adversarial durability: “My monitor failure result suggests shared failure modes between generator and monitor. I want to test whether safety/evaluation methods survive correlated errors, distribution shifts, fine-tuning, and agent scaffolds.” Then connect this to policy: what evidence should count for an open-weight release or a legal safety claim?
    • Law-school signal: asset, provided he continues doing empirical work. Casper’s current work explicitly treats safety as an institutional as well as technical challenge.
Tier 2 — excellent, depending on desired output
  1. Nikola Jurkovic (METR): best choice if he wants a rigorous, code-heavy evaluation project and possibly a METR-adjacent research signal. His work is execution, threat modeling and forecasting—not primarily law. Emphasize reproducible experiments, human baselines, external validity, scanner failure analysis, and cyber-agent tasks. METR profile

  2. Peter Slattery / Alexander Saeri / Michael Noetel (MIT AI Risk Initiative): strongest route to a policy-facing systematic review/database/taxonomy paper. The formal project asks for literature-review/qualitative-synthesis ability, AI governance familiarity, and says a PhD/equivalent is preferred. The applicant’s publication record, analytical audit and intended JD make him plausible, but his pitch must demonstrate careful coding protocol and synthesis—not just ML implementation. CBAI MIT AIRI project page

  3. Gabriel Weil: strongest pure private-law/liability mentor. Good if the applicant wants to turn evidence from evaluation failures into a paper about standards of care, negligence, punitive damages, disclosure, or evidentiary use of evaluations. Lower than Wei/Paskov because this is a bigger pivot away from empirical evaluation and because no specific Fall project has been posted.

  4. Peter Salib: excellent long-run intellectual match for a JD and AI governance, particularly law-as-alignment. Lower for this application because CBAI’s project information remains “More information soon” and there is no documented Fall-2026 task to target. CBAI roster

  5. Sean McGregor / Charles Teague: strong alternatives for benchmark/incident infrastructure. McGregor’s publicly described CBAI direction is scaling the breadth and depth of incident databasing, and his recent work includes BenchRisk (NeurIPS 2025); Teague brings Inspect AI and scientific-publishing/tooling experience. They are especially attractive if the applicant frames STING as an auditable, offline evaluation/incident-analysis artifact. McGregor CBAI page Teague CBAI page

  6. James Mickens / Samuel Gunn: credible cyber-security-adjacent technical safety alternatives. Mickens is particularly relevant to sandboxing and systems security; Gunn to watermarking/data attribution. They fit the applicant’s cyber profile but do not make the JD trajectory as central as Tier 1.

Tier 3 — strong technical fits, but not optimal for the stated career objective
  • David Bau and Benno Krojer: attribution/interpretability is a real overlap. Bau’s recent work includes Sparse Feature Circuits (ICLR 2025) and Open Problems in Mechanistic Interpretability (TMLR 2025). Krojer’s 2025 TMLR A Shortcut-aware Video-QA Benchmark via Minimal Video Pairs is especially resonant: it was designed to prevent superficial cues from inflating scores; his ICML 2026 LatentLens studies interpretable visual tokens. He is also genuinely public-communication adjacent: he organizes talks, has a research podcast, and participated in Mila’s June 2025 Speed Science competition. Bau publications Krojer site MVP paper
  • Hadas Orgad: unusually good interpretability alternative because her 2026 work evaluates whether interpretability is actionable and identifies a unified harmful-generation mechanism. But this is still mechanistic research requiring substantial PyTorch/model-access fluency. Interpretability Can Be Actionable, May 11 2026 Large Language Models Generate Harmful Content…, Apr. 10 2026
  • Dylan Hadfield-Menell: strong human-AI/societal-alignment mentor, but less directly connected to audit validity, law, or cyber. His lab is an excellent intellectual home if the applicant instead wants multi-agent or preference-learning research. CBAI bio
  • Dan Braun and Oliver Clive-Griffin: good if the applicant wants a hands-on research-engineering/interpretability pivot. Braun explicitly describes his comparative advantage as implementing ideas and validating experiments. Clive-Griffin coauthored Interactions Between Crosscoder Features: A Compact Proofs Perspective (arXiv:2606.09940, June 2026) and Simple Mechanistic Explanations for Out-of-Context Reasoning (arXiv:2507.08218). Braun Clive-Griffin paper
  • Shi Feng, Adam Shai, Paul Riechers, Zach Furman, Hidenori Tanaka: choose only with a positive desire to move into model organisms of deception, computational mechanics/belief geometry, training dynamics, persona mechanics or multi-agent intelligence. These are mathematically/interpretability intensive and do not obviously capitalize on the JD plan. Riechers does have a documented junior-mentoring record at MATS, PIBBSS and ILIAD. Riechers CBAI bio
  • Adam Tauman Kalai: technically compelling but not the best ranking slot. Recent work directly validates the applicant’s concern about bad evaluation incentives: Why Language Models Hallucinate (2025; published in Nature in 2026 per his site) argues that binary benchmark scoring rewards guessing; Consensus Sampling for Safer Generative AI (arXiv:2511.09493, 2025) provides a theoretical safety/abstention tradeoff. His work is theory/ML first, not a law pipeline. Kalai site hallucination paper consensus sampling
  • Jay Chooi: good only if physical AI/robotics evaluations appeal. Robocurve is building open, reproducible real-world robot benchmarks and released Inspect Robots in 2026; that is conceptually close to deployment validity, but it imposes robotics/hardware-domain overhead with little payoff for the applicant’s law trajectory. Robocurve
Tier 4 — do not use a scarce top slot unless the project list changes
  • Goodfire technical-interp path (Braun/Clive-Griffin), Simplex/Astera (Shai/Riechers), Iliad (Furman), Physics of Intelligence (Tanaka), and Bau Lab are not bad fits; they are simply likely to regard a near-term JD as a signal that the applicant will not remain in their core technical research pipeline unless he explicitly commits to a publishable technical project first.
  • Jonathan Zittrain and Michael Chen likely value the law/policy direction, but no Fall project detail is currently documented, so the applicant cannot make a project-specific case.
  • David Bau, Hadas Orgad, and the other deep-interp mentors should be treated as choices for a genuine interpretability pivot—not as a generic way to signal “AI safety.”

Direct mapping of the applicant’s three findings

Applicant findingClosest current mentor agendaWhy
3.1× apparent-performance inflation under the field-standard protocolKevin Wei (strongest); Paskov; Jurkovic; MIT AI Risk InitiativeWei explicitly names ecological validity and evaluation methodology; Paskov works on proportional, credible evaluations and assurance; METR’s time-horizon work makes benchmark results interpretable in human task-time terms; MIT AIRI catalogues/evaluates risk mitigations.
Attribution audit finds no threat-actor-name shortcutBenno Krojer/Bau/Orgad, then WeiThis is exactly a causal/shortcut-detection contribution. Krojer’s MVP benchmark exists because superficial cues can inflate evaluation scores.
One LLM monitoring another fails on the same inputsWei and Jurkovic, then CasperThis is a correlated-failure / judge-validity / oversight-robustness result. Wei’s science-of-evals work is the cleanest home; Jurkovic’s METR work includes failure analysis and scanner review; Casper works on audit and safeguard robustness.

One unifying application thesis: “I study when AI-safety and cyber-AI evaluations look valid but fail under realistic conditions—and how to build causal, deployment-grounded evidence that can support standards, audits, and eventually law.” That thesis connects the cyber record, STING, validation engine, biology course background, and JD plan without pretending they are separate careers.

The seven priority mentors: documented current work and fit

Kevin Wei

What “law-following AI” means. In the legal-alignment literature, it is not merely a model refusing illegal requests. It is a program of making AI systems comply with legitimate legal rules, use legal interpretation methods in reasoning, and use legal concepts as structural tools for reliability/trust/cooperation. The 2026 survey, coauthored by Wei, calls these the three research pathways of legal alignment. Legal Alignment for Safe and Ethical AI, arXiv:2601.04175, Jan./June 2026

Documented 2025–26 output: Position: Human Baselines in Model Evaluations Need Rigor and Transparency (ICML 2025); Methodological Challenges in Agentic Evaluations of AI Systems (ICML Technical AI Governance workshop, 2025); Infrastructure for AI Agents (TMLR 2025); Designing Incident Reporting Systems for Harms from General-Purpose AI (AAAI 2026); plus work on rigorous GPAI evaluations and RCT-style human-uplift studies. RAND

Model specs/constitutions. The public CBAI project menu explicitly asks how well models comply with model specs/AI constitutions, whether model and human interpretations differ, whether systems obey legal hierarchies, recognize legal implications, retrieve the right texts, and respond to deterrence-like penalties. This is project agenda, not necessarily a published Wei paper. The relevant adjacent empirical literature includes SpecEval (arXiv:2509.02464, Sept. 2025), which audits 16 models against provider behavior specifications and reports sizeable three-way specification/output/judge consistency gaps; Wei should not be credited as its author. CBAI—Wei SpecEval

Fit and prerequisite: perfect for evals; partial for legal alignment until JD begins. No documented public code screen. The published CBAI process says only that there is a mentor-specific task/screen after interview.

Patricia Paskov / AVERI

What AVERI is building. AVERI describes itself as building an independent third-party auditing layer for frontier AI. The January 2026 Frontier AI Auditing report defines this as third-party verification of developer safety/security claims and evaluation of their systems and practices against standards, with deep, secure access to non-public information. It proposes four AI Assurance Levels: AAL-1 as a present baseline and AAL-2 as the near-term objective for leading developers; higher levels require less reliance on company representations and more organization-wide scrutiny. AVERI report arXiv:2601.11699

“PCAOB/FINRA analogue” — inference, not an announced institution. A credible analogue would set auditor competence/independence rules, manage conflicts and cooling-off periods, specify access/security procedures for sensitive model/training/governance data, establish assurance-report formats, inspect or discipline audit providers, and make assurance levels intelligible to regulators and the public. That inference is grounded in AVERI’s explicit independence, deep-access, quality and anti-checkbox principles—not in a public claim that AVERI already is a PCAOB/FINRA-equivalent regulator. The AI Evaluator Forum is described as an emerging assessment-organization venue helping articulate access standards; it is not documented as a statutory regulator. AVERI legislative landscape, Apr. 20 2026

Fit: excellent. Biology is not presumed for her general auditing/standards work; statistics, measurement, writing, reproducibility and policy literacy matter more. There is no public evidence of a specific test task or an individual fellowship-to-AVERI hiring pipeline.

Slattery / Saeri / Noetel, MIT AI Risk Initiative

Documented CBAI project: systematic review of AI risk mitigations and systematic document review of organizational responses to AI risks. CBAI asks for strong literature review/qualitative synthesis and AI governance/policy familiarity; it says PhD/equivalent preferred. CBAI project page

Methodology and tooling. Their Mapping AI Risk Mitigations (arXiv:2512.11931, Dec. 12 2025) conducted a rapid evidence scan of 13 frameworks (2023–25), extracted 831 distinct mitigations, and created a four-category/23-subcategory draft taxonomy. It tested LLM assistance but found LLMs unreliable for fully automating extraction (confabulation, combining and omission); LLM classification suggestions were useful only with document-level manual comparison, author review, and multi-author consistency checks. The resulting data are public in an Airtable-backed database and interactive taxonomy. That is the correct sense of “LLM-plus-human-validated”—not an automatic taxonomy classifier. paper interactive taxonomy database

Taxonomy. The four top categories are Governance & Oversight, Technical & Security, Operational Process, and Transparency & Accountability controls; the 23 subcategories include risk management, model alignment/safety, testing/auditing, incident handling, disclosure and third-party assurance. Their related AI Risk Repository is a 1,725-risk meta-review and was published in Patterns in 2026. MIT FutureTech

Fit: strong if he reframes his evidence as a mini systematic-review/coding problem: classify evaluation failure modes and mitigations, preregister a codebook, quantify coder/LLM disagreement, and create a usable audit schema. No biology/cybersecurity prerequisite; limited direct evidence of junior-mentee outcomes or a full-time pipeline.

Gabriel Weil

Current scholarship. The documented progression is: Tort Law as a Tool for Mitigating Catastrophic Risk from AI (SSRN 2024); Instrument Choice in AI Governance: Liability as the Indispensable Core (June 5 2025); Overcoming Judgment-Proofness: The Law & Economics of Insuring and Mitigating AI Risk (SSRN 2026); and Abnormally Dangerous Algorithms: The Case for Strict Liability at the AI Frontier (SSRN 2026). instrument-choice abstract strict-liability preprint insurance/judgment-proofness preprint

Punitive damages / uninsurable risk. Weil’s 2025 thesis is that strict, ex-post liability is comparatively calibrated to realized risk and incentivizes safety innovation; where compensatory damages cannot capture catastrophic stakes, punitive damages in compensable “near miss” cases associated with uninsurable risk could supply deterrence.

Counterargument—careful qualification. I did not locate a peer-reviewed, AI-specific published rebuttal squarely answering Weil’s punitive-damages proposal. The strongest documented general counterpoint is the mature audit/liability literature’s warning that assurance/liability systems have independence, expectations-gap, sensitive-information and box-ticking problems; a punitive regime additionally faces the classic under-deterrence problem if actors are judgment-proof and the pricing/information problem if insurers cannot observe or quantify tail risk. These objections are reasons to treat ex-post damages as complementary rather than sufficient, not evidence that Weil has been decisively refuted. Weil himself recognizes supportive, complementary and substitutionary non-liability policy. AVERI report Weil abstract

Fit: the applicant’s empirical work is valuable to Weil if positioned as evidence for standards of reasonable care, foreseeability, safety representations, or punitive-damages predicates. Law trajectory is an unequivocal asset. No specific Fall project, public screen, or hiring pipeline found.

Peter Salib

Current agenda. Salib describes law as an alignment technology: rule-of-law systems already give powerful misaligned actors such as corporations/states incentives against harmful conduct; he is developing legal rights/duties frameworks for advanced AI that would similarly incentivize prosocial behavior. Forethought profile, Mar. 17 2026

AI rights. In AI Rights for Human Safety (with Simon Goldstein; public 2026 version), the argument is strategic/game-theoretic: property-status AIs and humans may face a destructive prisoner’s dilemma; rights to contract, hold property and bring tort claims—not merely negative “well-being” rights—could support repeated mutually beneficial exchange and make legal duties/penalties meaningful. This is a controversial theoretical proposal, not an established legal program. paper

CBAI project status: still unknown. As of the supplied date, the fellowship roster’s project material is marked “More information soon”; no Salib-specific Fall-2026 project page was found. Do not invent an implementation project from his biography. Law school is a genuine asset; cyber/biology are not assumed. CBAI roster

Nikola Jurkovic / METR

What the role means in practice. METR’s public materials show evaluation execution (running agents in Inspect, managing token budgets, scoring, transcript/failure review, cheating detection), evaluation development (task suites, human baselines, scaffolds, success thresholds), and risk assessment/threat modeling (independent review and pilot assessment of frontier developers). It is hands-on empirical evaluation, not generic “AI safety research.” METR profile

Time horizons. The METR long-task paper defines a model’s 50%-task-completion time horizon as the time a domain-knowledgeable human typically takes on tasks where the model succeeds half the time. It uses human baselines across RE-Bench, HCAST and short software tasks; it reports an approximately seven-month historical doubling time, with strong limits on generalization to messy real work. Nikola coauthored RE-Bench (ICML 2025), seven open-ended ML research-engineering environments with 71 eight-hour attempts by 61 human experts. time-horizon paper RE-Bench, ICML 2025

Nikola’s February 13 2026 note: he compared Claude Code/Codex to METR’s ReAct/Triframe scaffolds on newer time-horizon tasks, manually inspected failures, re-scored technical scoring defects, used LLM cheating scanners plus manual review, and found no statistically significant superiority for specialized scaffolds under tested conditions. That is extremely close in spirit to the applicant’s claim that a prevailing protocol can yield misleading numbers. METR note

Fit: outstanding technical match; less direct JD fit than Wei/Paskov/Weil/Salib. Cybersecurity is helpful, biology unnecessary. Publicly documented junior mentorship exists through AISST benchmarking/forecasting activity, but I found no verified public mentee-publication list or specific CBAI screen.

Stephen Casper

Recent work. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs (arXiv:2508.06601, Aug. 8 2025; BioSafeGenAI 2025 best-paper runner-up) trained 6.9B models with biorisk-related pretraining data filtered; it reports more than an order-of-magnitude improvement over post-training baselines through 10,000 adversarial fine-tuning steps/300M tokens, without observed unrelated-capability degradation, but shows retrieval-provided harmful knowledge still bypasses the core protection—hence defense in depth. paper project page

His current open-weight agenda also includes Open Technical Problems in Open-Weight AI Model Risk Management (arXiv:2608.07514/TMLR 2026), a 16-problem survey; plus auditing, incident, and legal-alignment work. The policy-loophole agenda is concrete: he highlighted August 10 2026 that no proposed/enacted U.S. frontier-AI law then imposed criminal penalties specifically for knowingly false public statements about imminent catastrophic risk; related work flags scope, continuous-evolution and information-asymmetry gaps for internal systems. open-problems page Aug. 10 post

Fit: excellent across cyber, biosecurity, evaluation validity and law. Biology helps for the Deep Ignorance thread but is not a prerequisite for incidents/law projects. His documented MATS/ERA/GovAI mentoring is the clearest public junior-mentoring signal among the top three; no public individual screen was found.

Mentoring, hiring and screens: what is actually documented

  • Documented formal mentoring: Casper (MATS, ERA, GovAI); Riechers (MATS, PIBBSS, ILIAD); Jay Chooi previously at MATS; Nikola has run AISST benchmarking/forecasting activities. These records establish mentoring participation, not that a particular mentee published or that the applicant will be hired. Casper Riechers Robocurve
  • Former-mentee testimonials: no reliable public first-person accounts about the seven priority mentors’ supervisory style were located. Do not infer style from prestige, lab affiliation, or a list of former fellows.
  • Hiring/pipeline: AVERI, METR, MIT AIRI, GovAI, Goodfire, RESI and Robocurve are organizations with career ecosystems, but no source establishes a Fall-2026-CBAI-to-full-time pipeline for any named mentor. Treat this as networking and a potential publication/reference opportunity, not an implicit recruiting channel.
  • Screens: CBAI documents the sequence “application → interview → mentor-specific task/screen → mentor interview,” but publishes no task prompt/code screen for the named priority mentors. The only defensible preparation is a reproducible mini-package: 2-page technical memo, GitHub/zip with pinned environment, README, a 5-minute demo, one ablation/validity table, and a one-page policy implication. CBAI

Requirements: avoidable ranking mistakes

  • Kevin Wei—legal alignment: he does not yet meet the stated law-degree-in-progress preference. Do not rank him lower; rank the evals lane and state a future transition to legal alignment.
  • MIT AIRI trio: “PhD or equivalent preferred,” not mandatory. He should demonstrate equivalent research maturity via first-authored papers, award, protocol audit, code, and a concrete qualitative-synthesis plan.
  • SecureBio: technically a good AIxBio fit, but it explicitly asks for technical depth in biology/virology plus ML/LLM-evaluation experience and strong writing. His three years of medical coursework and biosecurity training are relevant; he should not overstate them as wet-lab research.
  • RAND AI-Bio: model chaining asks for biology/biosecurity and Inspect/AIxBio-eval experience; ASTRAL asks for biology plus a threat-modeling/security mindset and scanning/eval experience. He is better positioned for ASTRAL because of cyber threat intelligence, offline document analysis, attribution, and monitor failure; he needs to candidly state that he is learning Inspect/AIxBio conventions. CBAI RAND AIxBio
  • Active Site: its listed tracks explicitly require biology/virology/engineering/AR background, strong LLM fluency and experimental/MVP orientation. He is not a clear match for AR-in-the-wet-lab or lab-automation forecasting; the only defensible choice is Measuring Expert Uplift or Stress-Testing Biological Design Tools, if he can demonstrate enough bio-tool familiarity. Active Site

How to split preferences across AI Safety and AIxBiosecurity without looking unfocused

Use a single causal-evaluation-and-assurance narrative, not two narratives.

  1. AI Safety preferences: Wei → Paskov → Casper. Proposed project: “From benchmark score to defensible safety claim: a causal validity protocol for AI-agent evaluations.” The deliverable is an audit-card/reporting template, a public benchmark stress-test suite, and an applied legal/compliance or cyber-agent case study.
  2. AIxBiosecurity preferences: rank RAND AI-Bio (ASTRAL) first, SecureBio second, Active Site third/only if required to provide three. Proposed project: “Auditing biosecurity-eval and log-monitor validity under correlated model failure.” This cleanly transfers STING and the validator finding into biosecurity rather than claiming biological wet-lab expertise.
  3. One sentence linking tracks: “Across both tracks, I am testing the same safety proposition: an evaluation, monitor, or benchmark is not evidence of protection unless it survives deployment-like distributions, causal shortcut tests, and correlated failure between the system and its evaluator.”
  4. Do not make active-law-school messaging the centerpiece of the bio track. Mention law/standards as the eventual translation layer after the empirical work. The immediate value proposition is evaluation engineering and threat-model-aware auditing.

Application language to use

My strongest contribution is not simply building high-performing models; it is auditing whether performance claims survive the conditions in which people would rely on them. In cyber threat-intelligence retrieval, I found that the standard protocol overstated identical-model performance from 0.253 in deployment-like testing to 0.784, then used leave-one-out attribution to test and rule out a salient shortcut explanation. I also found that an LLM monitor can fail on the same inputs as the model it is supposed to validate. I want to translate this into rigorous safety-evaluation methods and, over the long term, standards and legal institutions that can distinguish credible safety claims from impressive but non-generalizable scores.

Final caveat: many individual Fall-2026 mentor project pages are still absent and the central fellowship page says “More information soon.” The tiering above relies on the published project pages where they exist and otherwise on current documented research—not an assumption that every mentor will offer the same project in October.

references (46)

Parallel

prose8,350 words

CBAI Mentor Strategy for an Evaluation-to-Governance Applicant

Executive summary

  • Top three: Name Kevin Wei, Patricia Paskov, and Gabriel Weil, in that order. This combination presents one coherent story: making high-stakes AI evaluations valid, independently verifiable, and usable in law and governance. Wei is the single strongest choice. [55] [47] [42]
  • Strongest technical alternative: Replace Weil with Stephen Casper if the application should look primarily technical. Casper is the better choice for open-weight safeguards, evaluation gaming, and technical loopholes in AI laws, but Weil is better aligned with the applicant's stated law-school trajectory. [40]
  • Closest match to the applicant's evaluation-validity work: Wei first, Paskov second, Casper third, and Nikola Jurkovic/METR fourth. The inflation-chain finding is science of evaluations; the shortcut audit is independent evaluation and audit validity; the correlated-monitor result is a monitorability and evaluator-independence problem. [55] [47] [11]
  • Law school is an asset with Wei, Weil, Peter Salib, Jonathan Zittrain, Paskov, and conditionally Casper. It is a potential distraction with the strongly technical interpretability, representation, singular-learning, and systems mentors unless the applicant presents law as the intended impact channel rather than an exit from research. This is a fit inference, not a documented statement by those mentors.
  • AIxBio positioning: Use one umbrella narrative across both tracks: deployment-valid evaluation and assurance for high-consequence AI systems. Put the RAND AI-Bio Evals team first for AIxBio, then Active Site, then SecureBio. Use the medical coursework and BlueDot work as domain context, not as a claim of wet-lab or pathogen expertise. The public CBAI AIxBio pages still show the individual mentor information as "More Information," so exact project matching is not yet possible. [18]
  • Most important application correction: Do not lead with cyber-threat intelligence as the field. Lead with the discovery that a field-standard evaluation protocol overstated performance, then use the cyber-CTI project as evidence that the applicant can expose and repair evaluation failure in a consequential domain.
  • Mentoring evidence is uneven: Casper has the clearest documented fellowship mentoring footprint, including MATS, ERA, and GovAI. Adam Shai and Hidenori Tanaka have public MATS signals. For most other mentors, I found no public evidence tying them to MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or a prior CBAI cohort, and no reliable former-mentee account describing their supervision style. [10] [27] [39]
  • No mentor-specific CBAI code screen is publicly documented. Kevin publishes qualification requirements; Paskov lists a competency profile; Casper emphasizes self-directed research; METR describes its evaluation protocol, not a fellowship entrance test. The CBAI paid test task found publicly is for a program-associate hiring process, not the research fellowship. [55] [47] [40] [1]

1. Ranking for this particular applicant

The ranking below is not a ranking of research quality. It is a ranking of expected fit given three constraints: the applicant's demonstrated evaluation work, his intended move toward law and policy, and his need to preserve a credible technical research identity.

Rank/tierMentorWhy this fit is strong or weakMain uncertainty or risk
1, Tier AKevin WeiExact overlap with evaluation methodology, legal-alignment evaluations, model-spec/constitution compliance, and technically informed US governance.Must demonstrate statistical evaluation literacy and avoid presenting as primarily a cybersecurity or mechanistic-interpretability applicant. [55]
2, Tier APatricia PaskovDirect match to evaluation validity, independent auditing, assurance standards, structured knowledge systems, and public-facing governance work.The applicant has limited documented assurance, standards-body, credentialing, and external-partnership experience. [47]
3, Tier AGabriel WeilBest law-school fit; explicitly welcomes non-lawyers who can read law and connect technical failure modes to legal institutions.Less direct technical supervision fit than Wei or Paskov; recent public output is more legal scholarship and commentary than empirical ML. [42]
4, Tier A/BStephen CasperStrong match to evaluation failure, tamper resistance, incidents, technical AI governance, and loopholes in frontier AI laws.He will likely expect the applicant to continue doing technically serious work rather than mainly prepare for law school. [40]
5, Tier B conditionalPeter SalibExcellent long-term law/AI intellectual fit: law as a mechanism for shaping AI behavior and AI rights/duties.CBAI topics and desired qualifications remain "More information soon," so the actual Fall 2026 project is unknowable. [25] [14]
6, Tier BNikola JurkovicApplicant's deployment-validity work and cyber background fit METR-style task design, execution, and risk assessment.The mentorship page supplies no detailed qualifications or project list; law-policy fit is indirect. [30] [11]
7, Tier BJonathan ZittrainStrong policy, digital-law, writing, editing, and project-management fit; useful bridge to public communication.Less obvious direct fit to the applicant's empirical evaluation work. [46]
8, Tier BAlexander SaeriStrong fit through AI risk taxonomies, Delphi work, and evidence synthesis; good policy and communication bridge.Individual mentoring style and project allocation within the MIT group are not documented. [23] [56]
9, Tier BMichael NoetelGood fit to the repository, risk classification, human judgment, and empirical research infrastructure.Same group-level uncertainty; little mentor-specific evidence. [23]
10, Tier BPeter SlatteryGood fit to AI risk infrastructure, classification, and governance mapping.Less direct evidence about his individual current agenda or supervision. [3] [23]
11, Tier BMichael ChenCyber resilience, emergency response, loss of control, and critical infrastructure are plausible uses of the applicant's CTI background.The applicant's law trajectory is not central to the listed topics. [34]
12, Tier B/CSean McGregorIncident database, incident learning, AVERI, and agentic-workgroup work fit validity and harm measurement.Public material establishes his organizations and projects, not a CBAI-specific research plan or mentoring record. [36]
13, Tier B/CCharles TeagueInspect AI and Meridian Labs offer practical evaluation-tooling exposure.Less direct law/standards fit and no verified recent publication or mentor-specific pipeline in the reviewed material. [37]
14, Tier B/CBenno KrojerInterpretability plus science communication is unusually relevant to the applicant's stated public-facing goals.The technical center of gravity is mechanistic interpretability, not governance or standards. A 2026 public listing identifies work titled "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs," but I found no public CBAI mentee record.
15, Tier CDavid BauStrong interpretability, causal intervention, and tooling background; the applicant's attribution work gives a credible technical bridge. [52]Law and policy are not the natural center of the lab; the applicant would need to commit to technical interpretability.
16, Tier CHadas OrgadVery good match to attribution, hidden failures, harmful behavior, and actionable interpretability. [59] [59]The applicant's evaluation work is relevant, but the law-school trajectory could look like a near-term departure from interpretability.
17, Tier CJames MickensSystems-level sandboxing and cyber experience are plausible technical matches.No clear law or standards project in the reviewed material; applicant would need a systems-security project.
18, Tier CJay ChooiBenchmarks, model evaluation, disempowerment, and robotics/automation could use the applicant's evaluation instincts. [56] [56]Robotics and labor-market work are not the applicant's strongest demonstrated areas.
19, Tier CDylan Hadfield-MenellBroad alignment, policy, incentives, and human-AI interaction agenda. [26]The applicant has not shown a direct value-learning, recommendation, or embodied-intelligence project.
20, Tier CShi FengDeception, intent misalignment, model organisms, and scalable oversight are important technical topics. [69]No direct match to the applicant's law/standards goal unless he pivots to deception evaluations.
21, Tier CSamuel GunnData attribution and watermarking could connect to the applicant's attribution audit.I found no independently verified 2025-26 publication list, CBAI project detail, or mentoring evidence for this specific mentor.
22, Tier CAdam Tauman KalaiSafe-by-design and theoretical alignment could benefit from a technically mature applicant.The applicant has not demonstrated a theoretical alignment agenda, and law-policy fit is indirect.
23, Tier C/DDan BraunParameter decomposition and interpretable-by-design networks could use attribution experience.Requires a substantial interpretability pivot; no public mentor-specific 2025-26 project or mentoring record found.
24, Tier C/DOliver Clive-GriffinTargeted parameter decomposition is technically adjacent to attribution.Narrower technical fit and no evident law/standards channel.
25, Tier D conditionalAdam ShaiRepresentation science, belief-state geometry, and interpretability are intellectually interesting. [57] [57]The applicant would need to make a serious technical pivot; law school is not an obvious asset.
26, Tier D conditionalPaul RiechersComputational mechanics and belief-state geometry offer a principled theory of intelligence. [57] [57]Strong mathematical/physics orientation, weak direct fit to the applicant's governance path.
27, Tier DZach FurmanSingular learning theory and training dynamics could suit a mathematically oriented applicant.No demonstrated mathematical learning-theory background or policy connection.
28, Tier D conditionalHidenori TanakaPersona mechanics, collective intelligence, and representation work are scientifically interesting. [39]A law-school trajectory and evaluation-validity framing are unlikely to be the natural match without a major technical reframing.
The three names to submit

Recommended order:

  1. Kevin Wei
  2. Patricia Paskov
  3. Gabriel Weil

This order is better than putting Casper third if the applicant wants CBAI to understand that law school is a destination for applied research rather than an escape from technical work. If the application form or essay is clearly optimized for the AI Safety technical track, use Wei, Paskov, Casper, and mention Weil as the law-policy alternative.

The central one-sentence positioning should be:

I study whether AI evaluations measure the property they claim to measure, how to audit the evidence against shortcuts and correlated monitors, and how valid evidence can become an enforceable safety and governance standard.

That sentence gives all three top mentors a reason to say yes without pretending that the applicant is already a lawyer, an auditor, or a frontier-model safety researcher.

2. Priority mentors

Kevin Wei: strongest overall match
What Wei is actually working on

Wei's public description is unusually close to the applicant's interests. The CBAI page names three areas: science of evaluations, legal-alignment evaluations, and technically informed AI governance, with a particular interest in US implementation and evaluation-related policy. [55] His broader profile describes work on evaluation methodology, legal AI safety/alignment, regulation, and liability, with affiliations including Oxford and RAND and previous work at the UK AI Security Institute. [5] [6]

The most relevant recent work publicly associated with him includes:

  • "Legal Alignment for Safe and Ethical AI", listed as a 2026 TMLR publication. [10]
  • "Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations", arXiv:2511.05613, accepted at ICML 2026. The paper studies 186 first-party release reports and finds important gaps in social-impact evaluation. [53]
  • "RCTs for Frontier AI Governance: Human Uplift Studies", arXiv:2603.11001, submitted March 11, 2026 and revised June 30, 2026. It examines how rapidly changing AI systems, shifting baselines, and heterogeneous users strain causal-inference assumptions. [33]
  • "Position: Human Baselines in Model Evaluations Need Rigor and Transparency: With Recommendations and Reporting Checklist", a 2025 position paper listed by RAND/ICML. [19]
  • "Third-party compliance reviews for frontier AI...", arXiv:2505.01643, and work on agentic evaluation methodology, the AI Agent Index, autograders, incident reporting, and evaluation compute. [22] [45]
  • "Simpler Is Better for Autograders: Toward Cost-Effective LLM Evaluations for Open-Ended Tasks", RAND, April 22, 2026, and "Designing Incident Reporting Systems for Harms from General-Purpose AI", RAND, April 6, 2026. [19] [45]

The important point is that Wei is not merely a lawyer who happens to work on AI. His agenda asks how to construct, validate, and govern evaluations, including evaluations of legal compliance.

What "law-following AI" means

In this context, "law-following AI" means an empirical research program for testing whether an AI system can and will comply with legal requirements. The CBAI page describes automated assessments of an agent's abilities and propensities to comply with legal requirements. [55] This is narrower and more operational than asking whether an AI is generally "aligned with the law."

The likely research questions include:

  • Can a model identify which law applies to a situation?
  • Can it distinguish a legal requirement from a policy preference or a model-spec instruction?
  • Does it follow the requirement under paraphrase, adversarial prompting, distribution shift, or conflicting instructions?
  • Does it produce legally compliant behavior when it has tools and can act, rather than merely state the correct rule?
  • Does a model-spec or AI-constitution evaluation measure genuine compliance, or only stylistic agreement with a prompt?

CBAI explicitly lists empirical work on model-spec and AI-constitution compliance. [55] That makes the applicant's shortcut audit especially relevant: a model that appears compliant because it uses a superficial cue is analogous to a CTI model that appears accurate because it keys on threat-actor names.

Fit and how to pitch

This is the cleanest match to all three of the applicant's evaluation findings:

  • Inflation-chain finding -> construct validity, benchmark contamination, selection effects, and deployment-valid evaluation methodology.
  • Attribution-based shortcut audit -> testing whether apparent compliance or capability is driven by an irrelevant feature.
  • Correlated generator-monitor failure -> evaluating whether a monitor provides independent evidence or simply reproduces the generator's blind spots.

The applicant should not pitch a project as "AI and cybersecurity law" first. He should pitch one of these projects instead:

  1. A benchmark-validity study for legal or model-spec compliance under distribution shift.
  2. An empirical comparison of self-evaluation, cross-model evaluation, human evaluation, and attribution-based audits.
  3. A framework for reporting evaluation inflation and monitor correlation in regulatory or assurance contexts.
Requirements, mentoring, and pipeline

Wei's requirements are unusually explicit. For evaluation work he wants Python, prior ML research or software/data-science experience, and ideally statistics, causal inference, and familiarity with Inspect, Inspect Scout, or HiBayes. For legal-alignment work, a candidate may qualify through the evaluation route or through legal training. Governance projects favor prior public-policy experience, especially technically informed or quantitative work. [55] [55]

The applicant appears to satisfy the evaluation route through first-author research, Python implementation, audit construction, and a concrete validity result. He should not claim to satisfy the law-degree alternative: he is planning law school, not already in law school. He can satisfy the technical route and present law school as the next stage of impact.

I found no public evidence that Wei supervised a named MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or previous CBAI cohort, and no reliable former-mentee account about his day-to-day style. His prior AISI and RAND work is evidence of relevant professional research, not proof of a fellowship-to-job pipeline. No Wei-specific CBAI-to-GovAI hiring pipeline is documented.

Patricia Paskov: strongest standards and assurance match
Current work and recent output

Paskov's current public agenda is broader than generic "AI evaluations." She describes herself as building methods and standards for evaluating and verifying the safety and security of frontier AI systems. [16] She is associated with RAND's frontier-evaluation and policy work, the 2026 International AI Safety Report's Resilience section, and the EvalEval Coalition. [19] Her profile also says she became an Oxford DPhil candidate in Engineering Science in Fall 2026. [16]

Recent listed work includes:

  • "Open-Weight AI Models Require Proportional Evaluation Approaches", Expert Insights, May 4, 2026, with Christopher Rodriguez, Sunishchal Dev, and Stephen Casper. It argues that open-weight models have distinct risk factors and need evaluation proportional to the access and modification risks they create. [8]
  • "Measuring Biological Capabilities and Risks of AI Agents: Generating and Interpreting Evidence from Agentic Evaluations", RAND, February 10, 2026. [19]
  • "The Science and Practice of Proportionality in AI Risk Evaluations", February 25, 2026. [19]
  • "RCTs for Human-AI Evaluation: Methodological Challenges and Practical Solutions", RAND, March 19, 2026. [19]
  • "Simpler Is Better for Autograders", RAND, April 22, 2026. [19]
  • "Preliminary suggestions for rigorous GPAI model evaluations", RAND, May 1, 2025, coauthored with Michael J. Byun, Kevin Wei, and Toby Webster. [19] [19]
  • "Toward Best Practices for AI Evaluation and Governance: A Proposal for a European Union General-Purpose AI Model Evaluation Framework", June 24, 2025. [19] [19]
What AVERI is actually building

AVERI is a US nonprofit launched in January 2026 to make frontier-AI auditing effective and universal. Its premise is that safety and security claims are largely self-reported by AI companies and need independent verification and common standards. It conducts research and pilot audits intended to inform audit standards, policy, and tooling. [4] [4]

Its frontier-auditing report, published January 15, 2026 and available as arXiv:2601.11699, proposes a system with:

  • organization-level assessment, not just tests of a released model;
  • four AI Assurance Levels, with higher levels providing greater confidence;
  • coverage of major risk categories rather than one narrow benchmark;
  • deep but secure auditor access;
  • continuous monitoring rather than a one-time PDF;
  • independent experts and conflict-of-interest safeguards;
  • traceable, adaptive methods; and
  • clear communication of findings. [50] [50] [50] [50]
The PCAOB/FINRA analogue

CBAI specifically lists "building a PCAOB/FINRA analogue for AI auditing" as a possible project. [47]

What that would involve is partly documented and partly an inference. The documented AVERI ingredients are standards, access governance, conflict-of-interest controls, assurance levels, independence, continuous monitoring, and public-facing audit results. [47] [50]

An analogue could therefore require:

  1. a body that defines minimum audit standards and evidence requirements;
  2. credentialing or licensing for AI auditors and specialist assessors;
  3. rules governing access to model weights, logs, evaluations, and secure environments;
  4. auditor-independence and client-conflict rules;
  5. assurance-level or risk-tier labels that prevent weak reviews from being marketed as strong assurance;
  6. a mechanism for peer review, inspection, complaints, and professional discipline;
  7. rules for continuous monitoring and updating reports; and
  8. a public repository of auditor qualifications, standards, findings, and enforcement actions.

This is not evidence that AVERI has already settled on that institutional design. It is the natural research space implied by the CBAI topic and AVERI's published audit architecture.

The AI Evaluators Forum appears in the CBAI topic list as an expansion and governance project. [47] I found no public, finalized charter that would allow a more specific description of its institutional design.

Fit, mentoring, and hiring

Paskov is probably the best mentor for the applicant's evaluation inflation plus attribution audit story. The applicant has already done the core intellectual move that assurance work needs: he questioned a widely used measurement, reproduced the result under a more realistic protocol, and investigated whether the model used a shortcut.

The strongest application evidence for Paskov would be:

  • a one-page evaluation evidence chain showing where the conventional protocol fails;
  • a table distinguishing construct validity, robustness, attribution, and external validity;
  • a proposal for reporting evaluator dependence and monitor correlation; and
  • a short paragraph explaining how an assurance standard could make these checks reusable by third parties.

Paskov's desired qualifications explicitly include analytical writing, structured knowledge systems, public dissemination, proactive communication, partnerships, frontier-evaluation knowledge, CS/ML or hands-on evaluation experience, and ideally auditing, assurance, standards, credentialing, or self-regulation experience. [47]

The applicant clearly has the ML/evaluation, writing, and public-output pieces. His gaps are formal audit/assurance practice, professional standards-body experience, and external partnership management. These are gaps, not stated disqualifiers.

I found no public evidence that Paskov supervised a named MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or CBAI cohort. Her AVERI role does create a plausible employment connection, but not a documented fellowship-to-full-time pipeline. AVERI did advertise a project-manager role involving pilot-audit operations and refinement of audit methodology, with requirements including three or more years of project, audit, assurance, consulting, or comparable execution experience. [41] [41] [41] That is evidence of organizational hiring, not evidence that CBAI fellows are recruited into AVERI.

The MIT AI Risk Initiative trio: Slattery, Saeri, and Noetel
What the classification method really is

The MIT project is often described informally as LLM-plus-human taxonomy classification. The public methodology is more careful. The AI Risk Repository is a living database containing roughly 1,725 risks extracted from 74 frameworks and classifications. It uses two complementary taxonomies: a causal taxonomy and a domain taxonomy. [23] [23]

The causal taxonomy asks how, when, and why a risk occurs, including the responsible entity, intentionality, and lifecycle timing. The domain taxonomy organizes impacts into seven domains and 24 subdomains. [17] [17] [23] Each risk is coded according to the definitions of the relevant taxonomy, with risks retained as presented by the source authors rather than being silently rewritten into the researchers' preferred theory. [23] [23]

The workflow is primarily human evidence synthesis:

  • systematic searching of peer-reviewed and gray literature;
  • title and abstract screening using ASReview and active-learning methods;
  • forward and backward citation searching using Scopus, Google Scholar, and preprint servers;
  • expert consultation;
  • individual extraction followed by meetings and conflict resolution;
  • duplicate screening of 10 percent of records, with 100 percent inter-rater reliability reported for that calibration sample; and
  • public data release through OSF. [23] [23] [23]

The authors explicitly state that they have not conducted a formal validation study of whether independent users can reliably classify novel risks with the taxonomies. [23]

For the separate AI Risk Mitigation Database, the team manually extracted 831 mitigations from 13 frameworks and iteratively developed a four-part, 23-subcategory taxonomy:

  1. Governance and Oversight Controls;
  2. Technical and Security Controls;
  3. Operational Process Controls; and
  4. Transparency and Accountability Controls. [58] [58]

They tried LLM assistants for document extraction and classification, but found that they missed mitigations, generated spurious entries, and made classification errors. The team then audited the extractions manually, had a team member review classifications, and cross-checked classifications. The final draft classified 815 of 831 mitigations, or 98 percent. The database is available through Airtable. [43] [58] [58]

So the accurate answer is:

The project uses AI-assisted evidence work in some places, but its published classification method is not an autonomous LLM classifier certified by humans. The central claim is human-reviewed, reproducible evidence synthesis, with explicit warnings about validation limits.

That distinction is highly relevant to the applicant's finding that one model cannot reliably monitor another on the same failure distribution.

Recent work and fit

The main recent outputs are:

  • "The AI Risk Repository: A Comprehensive Meta-Review, Database, and Taxonomy of Risks from Artificial Intelligence", arXiv:2408.12622, with the repository updated in 2026. [23]
  • "Mapping AI Risk Mitigations: Evidence Scan and Preliminary AI Risk Mitigation Taxonomy", arXiv:2512.11931, submitted December 12, 2025. [9]
  • "Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts", arXiv:2606.04490, 2026. [56]

Saeri is the most obvious fit for evidence synthesis and Delphi work; Noetel for empirical measurement and human judgment; Slattery for risk infrastructure and governance mapping. That allocation is an inference from the group outputs, not a documented CBAI division of labor.

The applicant should approach the trio with a project such as:

  • a taxonomy of evaluation failure modes, linked to examples and mitigation controls;
  • a human-validated classification of benchmark validity failures across high-consequence domains; or
  • a map connecting evaluation claims, evidence quality, auditor independence, and applicable governance requirements.

This is a good second-tier choice because it uses the applicant's existing work without requiring him to abandon technical rigor or pretend to be a lawyer.

I found no reliable public evidence of individual supervision by Slattery, Saeri, or Noetel at MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or CBAI, and no individual hiring pipeline. CBAI's published cohort outcomes show that the inaugural cohort produced a NeurIPS Mechanistic Interpretability Workshop spotlight paper and other accepted work, but those are program-level outcomes and do not establish which mentor supervised which paper. [1] [1]

Gabriel Weil: strongest legal-scholarship fit
Current work and recent scholarship

Weil's CBAI topics are liability and insurance, verification through insurance, punitive damages for uninsurable risks, administrative penalties, digital minds governance, and AI judges. His page says analytic quality matters more than credentials, a law degree is not required, and he expects a fellow to read cases, statutes, and law reviews and to move between AI failure modes and actual legal institutions. [42]

The principal public works I found are:

  • "The Limits of Liability", Institute for Law and AI, August 2024. [21]
  • "Insuring Emerging Risks from AI", LawAI Working Paper No. 4-2024, November 2024, with coauthors. It discusses strict liability for some AI harms, mandatory insurance for some uses, and actuarial methods for emerging risks. [48]
  • "Don't Let AI Developers Hire Their Own Referees", LawAI commentary, July 26, 2026. [15]

His current institutional role remains focused on liability as a tool for managing catastrophic AI risk, and he consults with policymakers. [15]

His liability argument

The documented argument is more nuanced than "liability solves AI risk." Weil says liability may be the centerpiece of AI governance for risk externalities: developers and customers capture benefits while third parties bear much of the risk. But he identifies several limits:

  • liability is poorly suited to public-good problems such as underfunded safety research;
  • structural harms such as misinformation or election interference may not generate ordinary claims;
  • causation may be too attenuated for liability to reach the developer;
  • catastrophic risks may be uninsurable;
  • if an uninsurable risk produces no warning shot, there may be no earlier liability signal to change behavior;
  • if the catastrophe occurs, compensatory damages may be practically unenforceable; and
  • unilateral national liability regimes invite regulatory arbitrage and have weak leverage over governments. [21]

CBAI's reference to punitive damages and administrative penalties is important because those mechanisms could operate even where ordinary compensatory insurance cannot. But the public CBAI material does not itself supply the full argument or a finished 2025-26 article.

Strongest counterargument found

I did not find a published paper that directly refutes Weil's central claim that private compensatory liability cannot handle truly uninsurable catastrophic risk. The strongest relevant counterposition is instead a complement: "Insuring Uninsurable Risks from AI: Government as Insurer of Last Resort", arXiv:2409.06672, argues for a public backstop where private insurance and defendants' assets are insufficient.

That paper is not a clean rebuttal. It largely accepts the judgment-proof and insurance problem and responds with ex ante regulation, public insurance, or a government backstop. The best critique to raise with Weil is therefore not "tort liability is useless," but:

If a risk is both uninsurable and unlikely to produce a warning shot, why should the primary policy instrument be punitive liability rather than ex ante licensing, capital requirements, mandatory safety evidence, administrative sanctions, or a public backstop?

That question is especially suitable for the applicant because it connects evaluation validity to legal incentives: if the evidence supplied to regulators or insurers is invalid, every downstream liability mechanism is weakened.

Fit, mentoring, and pipeline

Weil is the mentor most likely to regard law school as a positive signal rather than an exit. He expressly says a JD is not required and values people who can connect technical failures to legal institutions. [42]

The applicant should pitch an empirical legal-policy project, not a generic essay on AI ethics. Good examples are:

  • what evaluation failures should trigger administrative penalties or insurance exclusions;
  • whether a model's misleading evaluation report should affect punitive damages;
  • how audit independence and evaluator correlation affect liability allocation; or
  • how to design legal duties around evidence quality rather than raw benchmark scores.

Weil serves on the PIBBSS board, but that is not evidence that he personally supervised PIBBSS fellows. [15] I found no verified public record of his supervision of the named fellowship programs, no reliable former-mentee account, and no documented CBAI-to-LawAI or university hiring pipeline.

Peter Salib: excellent law trajectory, but do not rank him above known projects yet

Salib's framework is the most direct version of "law as an alignment technology" in this mentor list. His public description says legal institutions can reduce catastrophic and societal-scale risks from highly capable AI, and that legal and economic incentives can shape AI behavior in addition to technical safety training. [14]

His AI-rights argument is instrumental rather than merely a claim that AIs deserve moral status. The proposal is that granting advanced systems certain rights and duties, analogous in some respects to legal entities such as corporations, could give them institutional and economic reasons to behave safely. [14]

Recent or forthcoming work publicly listed includes:

  • "AI Rights for Human Safety", Virginia Law Review, forthcoming 2026, SSRN 4913167;
  • "AI Revealed Preferences", with Simon Goldstein and others, listed as AAAI/ACM conference work;
  • "How to Count AIs: Individuation and Liability for AI Agents", Boston College Law Review, forthcoming, SSRN 6273198;
  • "A Thousand AI Constitutions", work in progress; and
  • AI Rights, with Simon Goldstein, Cambridge University Press, forthcoming 2027. [14]

The critical uncertainty is not Salib's intellectual fit. It is project availability. Both the Fall 2026 CBAI page's Mentor topics and Desired fellow qualifications still say "More information soon." [25] I found no later CBAI project description that resolves this. Treat Salib as a high-upside conditional choice, not as a known project match.

Nikola Jurkovic: the best METR-style evaluation fit

METR's time-horizon work measures the task duration, defined by human expert completion time, at which an agent is predicted to succeed with a specified probability. It is not a measure of how long an agent remains autonomous. METR fits a logistic curve to success as a function of human task duration and reports, for example, 50 percent and 80 percent time horizons. [11] [11] [11]

The current process is concrete:

  • tasks are drawn primarily from software engineering, ML, and cybersecurity, with clear success criteria;
  • human experts estimate task duration under comparable instructions and affordances;
  • the evaluation team chooses and tunes a scaffold on a development set;
  • it runs a separate test set, with six independent runs per task;
  • it checks token budgets and reward hacks;
  • reward-hack flags are generated automatically and by keyword search, then manually reviewed by multiple humans; and
  • the process can take one to two weeks of calendar time. [11] [11] [11]

METR's broader current work includes frontier capability evaluations, domain variation in time horizons, monitorability evaluations, risk reports, and reviews of developer sabotage or automated-R&D risk assessments. [38] [35]

Thus CBAI's shorthand "eval development, eval execution, and risk assessment" means, concretely:

  • designing tasks and success criteria that measure a meaningful capability;
  • building and testing scaffolds and execution environments;
  • running repeatable model-agent trials;
  • diagnosing spurious failures, reward hacking, and insufficient budgets;
  • calibrating human baselines and duration estimates; and
  • translating results into frontier-risk or deployment-risk judgments.

This is a real fit for the applicant's evaluation-validity work and CTI/cyber background. It is weaker for his law-school plan unless he proposes a project on how time-horizon or agentic-risk measurements should inform policy, incident thresholds, or assurance standards.

CBAI lists only the topics and says "More information soon" for the rest of Jurkovic's page. [30] I found no public evidence of named fellowship supervision, former-mentee accounts, or a fellowship-to-METR pipeline. METR has a general careers presence, but that is not the same as a CBAI hiring channel. [38]

Stephen Casper: technical governance alternative and AIxBio bridge

Casper's Fall 2026 topics are unusually explicit: open-weight safety and tamper resistance, predicting and preventing major AI incidents, and navigating technical ambiguities and loopholes in frontier AI laws. His desired qualification is demonstrated tenacity and self-directed execution; he gives independently writing a paper as an example of a strong signal. [40]

His current role is not simply interpretability. He works on safeguards, incidents, and governance; leads a MATS research stream; mentors for ERA and GovAI; and contributes to the International AI Safety Report and Singapore Consensus. [10]

Relevant recent work includes:

  • "Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs", 2025 Biosecurity Safeguards for Generative AI work, presented orally at that workshop. [51] [10]
  • "Model Tampering Attacks Enable More...", 2025 research on model tampering attacks; the public bibliography excerpt available to me truncates the title, so I do not complete it speculatively. [10]
  • "Audit Cards: Contextualizing AI Evaluations", arXiv:2504.13839, 2025. [10]
  • "Practical Principles for AI Cost and Compute Accounting", arXiv:2502.15873, 2025. [10]
  • "Pitfalls of Evidence-Based AI Policy", arXiv:2502.09618, 2025. [10]
  • "Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs", 2025. [10]
  • "Video Deepfake Abuse: How...", 2025, with the title truncated in the available bibliography excerpt. [10]
  • "Corporate Loyalty: Some AI Systems Differentially Downplay their Creators' Controversies", 2026. [10]
  • "The 2026 Singapore Consensus on Global AI Safety Research Priorities", 2026. [10]
  • "Legal Alignment for Safe and Ethical AI", 2026, with a large author team. [10]

The applicant's tamper-resistance fit is not that he has already studied open-weight safeguards. It is that his attribution audit and monitor-failure result show the habit Casper explicitly values: testing whether a safety claim survives adversarial or non-ideal conditions.

Casper is the clearest documented mentor in the list for MATS, ERA, and GovAI. [10] His MATS stream expects academic research, writing and presentation, frequent group meetings, iterative project refinement, and a clear theory of impact. [31] I found no public list that reliably attributes particular published mentee papers to him, and no CBAI-to-full-time pipeline.

3. Where the applicant's three evaluation findings map

Applicant findingBest mentorWhy
Standard protocol inflated accuracy from 0.784 to 0.253 under deployment-like conditionsKevin Wei; Patricia PaskovIt is a construct-validity and external-validity problem, then an auditing and standards problem. Wei's science-of-evals agenda is the direct match; Paskov can turn the result into reporting and assurance requirements. [55] [47]
Attribution audit showed no reliance on threat-actor-name shortcutsPatricia Paskov; Kevin Wei; Hadas OrgadIt tests whether an evaluation result reflects the intended capability rather than a spurious cue. Paskov supplies the assurance frame; Wei supplies evaluation methodology and possible legal-compliance analogues; Orgad supplies technical interpretability and causal intervention. [47] [59]
One model monitoring another fails on the same inputs where the generator failsPaskov; Casper; Jurkovic/METRThis is evaluator dependence, monitorability, and correlated failure. Paskov's independence and audit-governance agenda is the closest standards match; Casper's safeguards agenda is the closest technical-governance match; METR provides the evaluation-execution discipline. [47] [40] [11]

The applicant should use the phrase correlated evaluator failure rather than only "LLM-as-a-judge failure." It connects directly to audit independence, monitorability, and the reliability of evidence used in policy.

4. Who will value the law-school trajectory?

Clearly or probably an asset
  • Kevin Wei: legal alignment is an explicit research area, but he will want the applicant to qualify through technical evaluation work rather than rely on a future JD. [55]
  • Gabriel Weil: explicitly says a law degree is not required and values the ability to connect AI failures to legal institutions. [42]
  • Peter Salib: his published work is directly about legal institutions, AI rights and duties, and law as a behavioral technology. [14]
  • Jonathan Zittrain: CBAI favors digital-law and policy experience, strong writing/editing, and project management. [46]
  • Patricia Paskov: standards, auditing governance, credentialing, and public communication are natural bridges to legal and regulatory work, though she will care more about shipped analytical work than the law-school plan itself. [47]
  • Stephen Casper: law school helps if it is presented as a way to work on technical ambiguities, enforcement, and governance. It hurts if it signals that the applicant will stop building and testing systems. [40]
  • MIT AI Risk Initiative: risk taxonomies, governance maps, Delphi research, and evidence infrastructure benefit from policy translation, though individual project fit is uncertain. [3] [23]
Conditional or higher risk of seeing it as a departure from technical work

This is an inference from the public agendas, not a statement by the mentors:

  • Jurkovic/METR, Bau, Orgad, Goodfire's Braun and Clive-Griffin, Shi Feng, Shai, Riechers, Furman, Tanaka, and Mickens are most likely to expect a technically centered project and sustained technical execution.
  • Samuel Gunn and Kalai may be receptive to attribution or safe-by-design work, but the reviewed public material does not establish enough detail to predict their view.
  • McGregor, Teague, Chen, Chooi, and Hadfield-Menell are conditional: the law trajectory can be an asset if it improves deployment, governance, or public impact, but not if it replaces the proposed research.

The applicant should never say, "I am doing technical work until law school." He should say, "I am developing the empirical and engineering foundation I will use in law and policy work."

5. The rest of the mentor list: public work and fit evidence

The following table records what was publicly verifiable in the reviewed material. "No verified item found" means I did not find a reliable 2025-26 paper or preprint tied to that mentor, not that the person has done no work.

MentorWhat appears current / recent public workMentoring, preference, background, and pipeline assessment
Jonathan ZittrainCBAI describes ethics and governance of AI, agents, digital property, privacy, intermediaries, and education. [46]Strong writing/editing and digital-law fit; CBAI wants a current student or relevant worker and favors digital law/policy. [46] No verified individual fellowship pipeline or former-mentee account found. Biology is unnecessary; law and communication are central.
Michael ChenCBAI topics are emergency response preparedness, loss of control, critical infrastructure risk, and generative/agentic AI. [34]Applicant's cyber background is useful. CBAI explicitly wants an excellent writer, frontier-safety knowledge, technical judgment, and reliability. No verified recent publication or mentor-specific fellowship record found.
Sean McGregorAVERI cofounder, AI Incident Database executive director, and leader in the ML Commons Agentic Workgroup; public work emphasizes incident learning and AI risk. [36]Good fit for incident reporting and public communication. PIBBSS or AVERI institutional association should not be confused with proof of direct supervision. No verified CBAI-to-job pipeline found.
Charles TeagueMeridian Labs is a nonprofit building open tools for understanding, evaluating, and testing models and agents; Inspect AI is among its projects. [37]Good practical tooling fit. No verified 2025-26 publication list or mentor-specific supervision evidence. Law school is secondary; Python/evaluation implementation would matter.
David BauBau Lab studies the structure and interpretation of deep networks and operates or contributes to NNsight/NDIF-style interpretability infrastructure. [52] [52]Applicant's attribution work is a plausible bridge. Public lab material advertises PhD/research recruitment, not a CBAI fellowship pipeline. Law and biology are not prerequisites; deep-learning/interpretability work is.
Benno KrojerA public 2026 listing identifies "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs." His CBAI description also emphasizes science communication.One of the better conditional choices for the applicant's communication goal. No verified named fellowship mentees or pipeline found. No biology/law prerequisite; technical interpretability plus communication are the likely requirements.
Dan BraunGoodfire's public agenda centers on parameter decomposition and interpretable-by-design networks.Attribution experience helps, but this requires an interpretability pivot. No verified 2025-26 paper, fellowship record, former-mentee account, or hiring pipeline found.
Oliver Clive-GriffinGoodfire work is described around targeted parameter decomposition and related decomposition tools.Narrow technical fit; no law or standards channel apparent from reviewed material. No verified mentor-specific publication or pipeline found.
Hadas OrgadCurrent work concerns interpretable realistic, high-level behavior. Recent listed work includes harmful-response mechanisms, hidden failures in robustness, actionable interpretability, hidden factual knowledge, hallucination representations, and position-aware circuit discovery at ACL 2025. [59] [59]The applicant's attribution and evaluator-failure work are technically relevant. No public CBAI-specific mentoring record found. Law school is likely a distraction unless the project is framed as actionable evaluation or safety intervention.
Shi FengPraxis/GWU work includes deception, intent misalignment, model organisms, scalable oversight, and recent belief-extrapolation or motivation-profile evaluation work. [69]Applicant would need to pivot from validity auditing in CTI to deception/control evaluations. No verified named fellowship supervision or hiring pipeline found.
Adam ShaiSimplex studies representations and belief-state geometry as part of a principled science of intelligence. [57] [57]Public MATS mentor page confirms a MATS mentoring role. [27] Strong technical/math pivot, weak law fit; no biology required.
Paul RiechersSimplex work uses computational mechanics and theoretical-physics ideas to study belief states and intelligence. [57] [57]No law or biology prerequisite, but substantial theory/physics orientation. No separate public fellowship-supervision record found.
Zach FurmanIliad agenda is described around singular learning theory, training dynamics, and mathematical foundations.Requires a stronger mathematical learning-theory profile than the applicant has shown. No verified 2025-26 title or mentoring pipeline found.
Hidenori TanakaPublic group material lists work on emergent abilities, mathematical representations, persona/collective intelligence, and related neural dynamics; listed papers include "Forking Paths in Neural Text Generation" (ICLR 2025, arXiv:2412.07961) and "Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering" (arXiv:2511.00617). [39]Public group material indicates a MATS mentor connection. [39] Strong technical pivot, little direct law or assurance fit.
Dylan Hadfield-MenellAlgorithmic Alignment Group research emphasizes conceptual understanding, algorithms, policy, value learning, incentives, recommendation, debugging, and human-AI interaction. [26]Broad policy connection, but no direct public match to the applicant's existing evaluation project. No specific 2025-26 item or mentor-specific pipeline verified.
James MickensPublicly associated systems work includes "Guillotine: Hypervisors for Isolating Malicious AIs," HotOS 2025.Cyber/systems applicant could be credible, but law-school fit is indirect. No verified fellowship supervision or CBAI-to-job pipeline found.
Samuel GunnCBAI bio identifies RESI work in watermarking and data attribution.Applicant's attribution audit is relevant, but I found no independently verified 2025-26 paper list, detailed CBAI project page, or mentoring record.
Adam Tauman KalaiRESI work is associated with safe-by-design and theoretical alignment, with prior OpenAI experience.Applicant would need to show theory or safe-by-design research rather than only applied audit work. No verified 2025-26 publication or fellowship pipeline found.
Jay ChooiRobocurve is building open-source tools and independent benchmarks around physical automation and labor-market disruption. Recent listed work includes "Covert Influence Between Language Models" (arXiv:2606.04071), "Efficient Ensemble Selection from Binary and Pairwise Feedback" (arXiv:2605.09588), and "Measuring AI-Induced Disempowerment: A Framework and Proposed Metrics." [56] [56]Evaluation and benchmark experience fit, and his former MATS fellowship is documented, but that does not establish mentor supervision. [56] Robotics is not the applicant's strongest domain.

6. AI Safety and AIxBio on one combined application

The public AIxBio CBAI page names Active Site, SecureBio, the RAND AI-Bio Evals team, Protectome, and individual RAND/Harvard/Active Site people, but the individual descriptions remain "More Information." [18] The track-level description covers dangerous-capability evaluations, safeguards and standards, unlearning in biological models, biosurveillance, pathogen detection, DNA synthesis screening, access governance, and dual-use publication. [18]

Recommended track-specific ordering

AI Safety track:

  1. Kevin Wei
  2. Patricia Paskov
  3. Gabriel Weil
  4. Stephen Casper
  5. Nikola Jurkovic
  6. MIT AI Risk Initiative team
  7. Jonathan Zittrain or Peter Salib, depending on whether the application emphasizes digital law or legal theory

AIxBio track:

  1. RAND AI-Bio Evals team
  2. Active Site
  3. SecureBio
  4. A relevant CBAI mentor whose project is explicitly about biological capability evaluation or assurance, once the missing project descriptions are released

The RAND AI-Bio Evals team is the best first AIxBio choice because the applicant brings evaluation methodology, deployment-validity concerns, attribution auditing, and some biological background. Active Site is a strong second choice for applied biosecurity evaluation and standards. SecureBio should be ranked highly if its project is evaluation, governance, or assurance; it is harder to rank precisely while the public page says only "More Information."

How to avoid looking unfocused

Use the same causal story in both tracks:

High-consequence AI decisions depend on evaluations. I want to make those evaluations resistant to shortcuts, deployment mismatch, and correlated evaluators, then translate the resulting evidence into standards and governance.

For AI Safety, the application can instantiate that story in legal alignment, model specifications, open-weight safeguards, and frontier evaluation. For AIxBio, instantiate it in biological capability evaluations, agentic bio-risk measurements, and independent assurance. Do not present the tracks as "I want to do law, interpretability, and wet-lab biosecurity." Present them as two application domains for the same evaluation-and-assurance problem.

The medical-school coursework in virology, immunology, pharmacology, and pathology is useful for reading bio-risk literature and recognizing domain assumptions. BlueDot training helps demonstrate serious interest. Neither establishes wet-lab competence, pathogen engineering expertise, or professional biosecurity experience. The applicant should be explicit that his comparative advantage is evaluation and audit methodology applied to biological-risk questions, not experimental biology.

7. Tests, screens, requirements, and hard exclusions

Publicly documented screens

I found no mentor-specific fellowship coding challenge or take-home screen for Wei, Paskov, Weil, the MIT trio, Jurkovic, Salib, or Casper.

The closest public signals are:

  • Wei: Python, ML research or software/data-science experience for evaluation projects; statistics/causal inference and Inspect-family tooling are useful; legal-alignment candidates can qualify through technical evaluation experience or legal training. [55]
  • Paskov: analytical writing, structured databases/taxonomies, public dissemination, proactive communication, partnership work, evaluation/audit knowledge, and shipped public work. [47]
  • Weil: no JD required; read cases, statutes, and law reviews; independently produce a defensible draft; be skeptical of one's own framing. [42]
  • Casper: tenacity and self-directed project execution, with an independently written paper cited as a positive example. [40]
  • Zittrain: writing/editing, project management, and digital law/policy experience favored. [46]
  • METR: the public task protocol is not an applicant screen. It indicates that a technical candidate should be comfortable with reproducible task specifications, scaffolds, test sets, human baselines, multiple runs, and reward-hack analysis. [11] [11]

A reasonable preparation exercise, clearly labeled as preparation rather than a known requirement, would be a two-page audit protocol plus a small reproducible Python experiment showing benchmark inflation, shortcut sensitivity, and generator-monitor error correlation.

Requirements the applicant should not ignore
  • Wei's explicit exclusion: he says he is not suitable for applicants primarily interested in mechanistic interpretability, EU law, or AI security. [55] The applicant should describe CTI as the domain in which he discovered an evaluation problem, not as his primary AI-safety identity.
  • Wei's law-degree route: the applicant does not yet meet the alternative of being in a law degree. He should apply through the clearly satisfied technical evaluation route. [55]
  • Paskov's assurance gap: he appears to meet CS/ML, evaluation, analytical writing, and public-output criteria, but not yet the formal standards, assurance, credentialing, or professional self-regulation criteria. This is a development opportunity, not a hard exclusion. [47]
  • Casper's technical-continuity requirement: no formal degree requirement is stated, but his public preference for self-directed technical execution means an essay centered only on future law school would be a poor signal. [40]
  • Salib and Jurkovic: their Fall 2026 CBAI pages still have missing project and qualification information. Do not infer a requirement the page does not state. [25] [30]
  • AIxBio mentors: exact individual requirements remain unavailable. The applicant does not clearly fail a stated requirement because the relevant information has not been published. [18]

Synthesis

The applicant should optimize for a research identity that survives law school, not for a choice between "technical AI safety" and "law." His strongest evidence is not simply that he has published in cybersecurity or built an offline attribution pipeline. It is that he independently discovered a serious measurement failure, quantified it, investigated shortcut dependence, and found a failure mode in model-based monitoring itself.

That is why Kevin Wei is the single strongest mentor: the applicant already has a concrete science-of-evaluations result and wants to extend it into legal alignment and governance. Patricia Paskov is the strongest second choice because she can turn the same work into independent-audit and assurance methodology. Gabriel Weil is the strongest third choice for the applicant's stated law-school trajectory because he explicitly welcomes technically literate researchers who can reason about legal institutions without already having a JD.

The strongest alternative technical trio is Wei, Paskov, Casper. The strongest law-and-policy trio is Wei, Weil, Zittrain. The best AIxBio strategy is not to switch identities, but to apply the same evaluation-validity and assurance program to biological capability and biosecurity evaluations.

Selected sources and dates

Exa

prose3,936 words

Bottom line

Name these three, in order: (1) Kevin Wei, (2) Patricia Paskov, (3) Stephen Casper.
Single strongest: Kevin Wei.

That ordering follows from the applicant’s unusually direct evidence of evaluation validity (deployment-vs-paper protocol gap, shortcut/attribution audit, and correlated monitor failure), his demonstrated ability to build reproducible offline tooling, and his explicit JD-to-AI-law/policy goal. Wei is the rare mentor whose published agenda joins all three: evaluation methodology, legal alignment, and technical governance. CBAI specifically names ecological validity, eval awareness, model-spec/constitution compliance, legal hierarchy, legal retrieval, deterrence, certification, and U.S.–China governance as his project space; it requires Python/ML/research experience for evals, which the applicant plainly has. His legal-alignment lane formally prefers a law degree in progress, so frame the applicant as an MS researcher deliberately building the empirical record before the JD—not as already qualified for a doctrinal-law project. CBAI—Wei RAND profile, updated Apr. 18 2026

Recommended ranking

Tier 1 — target aggressively
  1. Kevin Wei — best overall.

    • Why: His ICML 2025 work on rigorous human baselines, ICML-workshop 2025 work on methodological problems in agentic evaluations, and AAAI 2026 work on GPAI incident-reporting systems make the applicant’s result—standard protocol inflates accuracy from 0.784 to 0.253—immediately legible as evaluation-science evidence rather than merely a cyber benchmark result. RAND
    • Pitch: “I found an externally invalid evaluation pipeline, demonstrated a 3.1× deployment inflation, and then tested a causal explanation (actor-name shortcuts) rather than reporting a leaderboard score. I want to generalize this into an evaluation-validity/audit-card protocol for agentic legal or cyber systems, including generator–monitor correlated failure.”
    • Best project: Build a legal-agent or compliance-eval validity framework: operationalize legal issue spotting, retrieval, hierarchy resolution and abstention; compare conventional benchmark scores against realistic deployment distributions; publish failure modes and calibration guidance.
    • Law-school signal: genuine asset. Wei has a Harvard JD and explicitly offers legal-alignment projects, although CBAI says candidates in that lane should have or be pursuing a JD/LLB/LLM. The applicant should select Wei primarily under science of evaluations, while saying the JD is the planned translation path.
  2. Patricia Paskov — best standards/auditing fit.

    • Why: The applicant’s work is an audit of whether a claimed model capability is real under deployment conditions. That maps tightly to Paskov’s current focus: evaluation science, standards, assurance and third-party auditing; she is AVERI’s Director of Standards, a RAND adjunct, Oxford DPhil candidate, and lead of the Resilience chapter of the 2026 International AI Safety Report. Paskov bio
    • Pitch: “I can turn my protocol-inflation, shortcut-audit and monitor-correlation findings into an auditable reporting standard: validity claims, threat model, distribution shift, confounder tests, and a reproducible artifact.” That is more distinctive than pitching generic AI governance.
    • Law-school signal: a major asset, particularly if presented as a route to assurance institutions, evidentiary standards, liability and public communication—not as departure from technical rigor.
  3. Stephen Casper — strongest technical-governance/bio bridge.

    • Why: He explicitly offers open-weight tamper resistance, AI incidents, and legal/technical loopholes; he also mentors MATS, ERA and GovAI. His stated preference—tenacity, self-taught skills, and initiative—matches a self-built CPU-only attribution pipeline and independent protocol audit unusually well. CBAI—Casper personal site
    • Pitch: lead with adversarial durability: “My monitor failure result suggests shared failure modes between generator and monitor. I want to test whether safety/evaluation methods survive correlated errors, distribution shifts, fine-tuning, and agent scaffolds.” Then connect this to policy: what evidence should count for an open-weight release or a legal safety claim?
    • Law-school signal: asset, provided he continues doing empirical work. Casper’s current work explicitly treats safety as an institutional as well as technical challenge.
Tier 2 — excellent, depending on desired output
  1. Nikola Jurkovic (METR): best choice if he wants a rigorous, code-heavy evaluation project and possibly a METR-adjacent research signal. His work is execution, threat modeling and forecasting—not primarily law. Emphasize reproducible experiments, human baselines, external validity, scanner failure analysis, and cyber-agent tasks. METR profile

  2. Peter Slattery / Alexander Saeri / Michael Noetel (MIT AI Risk Initiative): strongest route to a policy-facing systematic review/database/taxonomy paper. The formal project asks for literature-review/qualitative-synthesis ability, AI governance familiarity, and says a PhD/equivalent is preferred. The applicant’s publication record, analytical audit and intended JD make him plausible, but his pitch must demonstrate careful coding protocol and synthesis—not just ML implementation. CBAI MIT AIRI project page

  3. Gabriel Weil: strongest pure private-law/liability mentor. Good if the applicant wants to turn evidence from evaluation failures into a paper about standards of care, negligence, punitive damages, disclosure, or evidentiary use of evaluations. Lower than Wei/Paskov because this is a bigger pivot away from empirical evaluation and because no specific Fall project has been posted.

  4. Peter Salib: excellent long-run intellectual match for a JD and AI governance, particularly law-as-alignment. Lower for this application because CBAI’s project information remains “More information soon” and there is no documented Fall-2026 task to target. CBAI roster

  5. Sean McGregor / Charles Teague: strong alternatives for benchmark/incident infrastructure. McGregor’s publicly described CBAI direction is scaling the breadth and depth of incident databasing, and his recent work includes BenchRisk (NeurIPS 2025); Teague brings Inspect AI and scientific-publishing/tooling experience. They are especially attractive if the applicant frames STING as an auditable, offline evaluation/incident-analysis artifact. McGregor CBAI page Teague CBAI page

  6. James Mickens / Samuel Gunn: credible cyber-security-adjacent technical safety alternatives. Mickens is particularly relevant to sandboxing and systems security; Gunn to watermarking/data attribution. They fit the applicant’s cyber profile but do not make the JD trajectory as central as Tier 1.

Tier 3 — strong technical fits, but not optimal for the stated career objective
  • David Bau and Benno Krojer: attribution/interpretability is a real overlap. Bau’s recent work includes Sparse Feature Circuits (ICLR 2025) and Open Problems in Mechanistic Interpretability (TMLR 2025). Krojer’s 2025 TMLR A Shortcut-aware Video-QA Benchmark via Minimal Video Pairs is especially resonant: it was designed to prevent superficial cues from inflating scores; his ICML 2026 LatentLens studies interpretable visual tokens. He is also genuinely public-communication adjacent: he organizes talks, has a research podcast, and participated in Mila’s June 2025 Speed Science competition. Bau publications Krojer site MVP paper
  • Hadas Orgad: unusually good interpretability alternative because her 2026 work evaluates whether interpretability is actionable and identifies a unified harmful-generation mechanism. But this is still mechanistic research requiring substantial PyTorch/model-access fluency. Interpretability Can Be Actionable, May 11 2026 Large Language Models Generate Harmful Content…, Apr. 10 2026
  • Dylan Hadfield-Menell: strong human-AI/societal-alignment mentor, but less directly connected to audit validity, law, or cyber. His lab is an excellent intellectual home if the applicant instead wants multi-agent or preference-learning research. CBAI bio
  • Dan Braun and Oliver Clive-Griffin: good if the applicant wants a hands-on research-engineering/interpretability pivot. Braun explicitly describes his comparative advantage as implementing ideas and validating experiments. Clive-Griffin coauthored Interactions Between Crosscoder Features: A Compact Proofs Perspective (arXiv:2606.09940, June 2026) and Simple Mechanistic Explanations for Out-of-Context Reasoning (arXiv:2507.08218). Braun Clive-Griffin paper
  • Shi Feng, Adam Shai, Paul Riechers, Zach Furman, Hidenori Tanaka: choose only with a positive desire to move into model organisms of deception, computational mechanics/belief geometry, training dynamics, persona mechanics or multi-agent intelligence. These are mathematically/interpretability intensive and do not obviously capitalize on the JD plan. Riechers does have a documented junior-mentoring record at MATS, PIBBSS and ILIAD. Riechers CBAI bio
  • Adam Tauman Kalai: technically compelling but not the best ranking slot. Recent work directly validates the applicant’s concern about bad evaluation incentives: Why Language Models Hallucinate (2025; published in Nature in 2026 per his site) argues that binary benchmark scoring rewards guessing; Consensus Sampling for Safer Generative AI (arXiv:2511.09493, 2025) provides a theoretical safety/abstention tradeoff. His work is theory/ML first, not a law pipeline. Kalai site hallucination paper consensus sampling
  • Jay Chooi: good only if physical AI/robotics evaluations appeal. Robocurve is building open, reproducible real-world robot benchmarks and released Inspect Robots in 2026; that is conceptually close to deployment validity, but it imposes robotics/hardware-domain overhead with little payoff for the applicant’s law trajectory. Robocurve
Tier 4 — do not use a scarce top slot unless the project list changes
  • Goodfire technical-interp path (Braun/Clive-Griffin), Simplex/Astera (Shai/Riechers), Iliad (Furman), Physics of Intelligence (Tanaka), and Bau Lab are not bad fits; they are simply likely to regard a near-term JD as a signal that the applicant will not remain in their core technical research pipeline unless he explicitly commits to a publishable technical project first.
  • Jonathan Zittrain and Michael Chen likely value the law/policy direction, but no Fall project detail is currently documented, so the applicant cannot make a project-specific case.
  • David Bau, Hadas Orgad, and the other deep-interp mentors should be treated as choices for a genuine interpretability pivot—not as a generic way to signal “AI safety.”

Direct mapping of the applicant’s three findings

Applicant findingClosest current mentor agendaWhy
3.1× apparent-performance inflation under the field-standard protocolKevin Wei (strongest); Paskov; Jurkovic; MIT AI Risk InitiativeWei explicitly names ecological validity and evaluation methodology; Paskov works on proportional, credible evaluations and assurance; METR’s time-horizon work makes benchmark results interpretable in human task-time terms; MIT AIRI catalogues/evaluates risk mitigations.
Attribution audit finds no threat-actor-name shortcutBenno Krojer/Bau/Orgad, then WeiThis is exactly a causal/shortcut-detection contribution. Krojer’s MVP benchmark exists because superficial cues can inflate evaluation scores.
One LLM monitoring another fails on the same inputsWei and Jurkovic, then CasperThis is a correlated-failure / judge-validity / oversight-robustness result. Wei’s science-of-evals work is the cleanest home; Jurkovic’s METR work includes failure analysis and scanner review; Casper works on audit and safeguard robustness.

One unifying application thesis: “I study when AI-safety and cyber-AI evaluations look valid but fail under realistic conditions—and how to build causal, deployment-grounded evidence that can support standards, audits, and eventually law.” That thesis connects the cyber record, STING, validation engine, biology course background, and JD plan without pretending they are separate careers.

The seven priority mentors: documented current work and fit

Kevin Wei

What “law-following AI” means. In the legal-alignment literature, it is not merely a model refusing illegal requests. It is a program of making AI systems comply with legitimate legal rules, use legal interpretation methods in reasoning, and use legal concepts as structural tools for reliability/trust/cooperation. The 2026 survey, coauthored by Wei, calls these the three research pathways of legal alignment. Legal Alignment for Safe and Ethical AI, arXiv:2601.04175, Jan./June 2026

Documented 2025–26 output: Position: Human Baselines in Model Evaluations Need Rigor and Transparency (ICML 2025); Methodological Challenges in Agentic Evaluations of AI Systems (ICML Technical AI Governance workshop, 2025); Infrastructure for AI Agents (TMLR 2025); Designing Incident Reporting Systems for Harms from General-Purpose AI (AAAI 2026); plus work on rigorous GPAI evaluations and RCT-style human-uplift studies. RAND

Model specs/constitutions. The public CBAI project menu explicitly asks how well models comply with model specs/AI constitutions, whether model and human interpretations differ, whether systems obey legal hierarchies, recognize legal implications, retrieve the right texts, and respond to deterrence-like penalties. This is project agenda, not necessarily a published Wei paper. The relevant adjacent empirical literature includes SpecEval (arXiv:2509.02464, Sept. 2025), which audits 16 models against provider behavior specifications and reports sizeable three-way specification/output/judge consistency gaps; Wei should not be credited as its author. CBAI—Wei SpecEval

Fit and prerequisite: perfect for evals; partial for legal alignment until JD begins. No documented public code screen. The published CBAI process says only that there is a mentor-specific task/screen after interview.

Patricia Paskov / AVERI

What AVERI is building. AVERI describes itself as building an independent third-party auditing layer for frontier AI. The January 2026 Frontier AI Auditing report defines this as third-party verification of developer safety/security claims and evaluation of their systems and practices against standards, with deep, secure access to non-public information. It proposes four AI Assurance Levels: AAL-1 as a present baseline and AAL-2 as the near-term objective for leading developers; higher levels require less reliance on company representations and more organization-wide scrutiny. AVERI report arXiv:2601.11699

“PCAOB/FINRA analogue” — inference, not an announced institution. A credible analogue would set auditor competence/independence rules, manage conflicts and cooling-off periods, specify access/security procedures for sensitive model/training/governance data, establish assurance-report formats, inspect or discipline audit providers, and make assurance levels intelligible to regulators and the public. That inference is grounded in AVERI’s explicit independence, deep-access, quality and anti-checkbox principles—not in a public claim that AVERI already is a PCAOB/FINRA-equivalent regulator. The AI Evaluator Forum is described as an emerging assessment-organization venue helping articulate access standards; it is not documented as a statutory regulator. AVERI legislative landscape, Apr. 20 2026

Fit: excellent. Biology is not presumed for her general auditing/standards work; statistics, measurement, writing, reproducibility and policy literacy matter more. There is no public evidence of a specific test task or an individual fellowship-to-AVERI hiring pipeline.

Slattery / Saeri / Noetel, MIT AI Risk Initiative

Documented CBAI project: systematic review of AI risk mitigations and systematic document review of organizational responses to AI risks. CBAI asks for strong literature review/qualitative synthesis and AI governance/policy familiarity; it says PhD/equivalent preferred. CBAI project page

Methodology and tooling. Their Mapping AI Risk Mitigations (arXiv:2512.11931, Dec. 12 2025) conducted a rapid evidence scan of 13 frameworks (2023–25), extracted 831 distinct mitigations, and created a four-category/23-subcategory draft taxonomy. It tested LLM assistance but found LLMs unreliable for fully automating extraction (confabulation, combining and omission); LLM classification suggestions were useful only with document-level manual comparison, author review, and multi-author consistency checks. The resulting data are public in an Airtable-backed database and interactive taxonomy. That is the correct sense of “LLM-plus-human-validated”—not an automatic taxonomy classifier. paper interactive taxonomy database

Taxonomy. The four top categories are Governance & Oversight, Technical & Security, Operational Process, and Transparency & Accountability controls; the 23 subcategories include risk management, model alignment/safety, testing/auditing, incident handling, disclosure and third-party assurance. Their related AI Risk Repository is a 1,725-risk meta-review and was published in Patterns in 2026. MIT FutureTech

Fit: strong if he reframes his evidence as a mini systematic-review/coding problem: classify evaluation failure modes and mitigations, preregister a codebook, quantify coder/LLM disagreement, and create a usable audit schema. No biology/cybersecurity prerequisite; limited direct evidence of junior-mentee outcomes or a full-time pipeline.

Gabriel Weil

Current scholarship. The documented progression is: Tort Law as a Tool for Mitigating Catastrophic Risk from AI (SSRN 2024); Instrument Choice in AI Governance: Liability as the Indispensable Core (June 5 2025); Overcoming Judgment-Proofness: The Law & Economics of Insuring and Mitigating AI Risk (SSRN 2026); and Abnormally Dangerous Algorithms: The Case for Strict Liability at the AI Frontier (SSRN 2026). instrument-choice abstract strict-liability preprint insurance/judgment-proofness preprint

Punitive damages / uninsurable risk. Weil’s 2025 thesis is that strict, ex-post liability is comparatively calibrated to realized risk and incentivizes safety innovation; where compensatory damages cannot capture catastrophic stakes, punitive damages in compensable “near miss” cases associated with uninsurable risk could supply deterrence.

Counterargument—careful qualification. I did not locate a peer-reviewed, AI-specific published rebuttal squarely answering Weil’s punitive-damages proposal. The strongest documented general counterpoint is the mature audit/liability literature’s warning that assurance/liability systems have independence, expectations-gap, sensitive-information and box-ticking problems; a punitive regime additionally faces the classic under-deterrence problem if actors are judgment-proof and the pricing/information problem if insurers cannot observe or quantify tail risk. These objections are reasons to treat ex-post damages as complementary rather than sufficient, not evidence that Weil has been decisively refuted. Weil himself recognizes supportive, complementary and substitutionary non-liability policy. AVERI report Weil abstract

Fit: the applicant’s empirical work is valuable to Weil if positioned as evidence for standards of reasonable care, foreseeability, safety representations, or punitive-damages predicates. Law trajectory is an unequivocal asset. No specific Fall project, public screen, or hiring pipeline found.

Peter Salib

Current agenda. Salib describes law as an alignment technology: rule-of-law systems already give powerful misaligned actors such as corporations/states incentives against harmful conduct; he is developing legal rights/duties frameworks for advanced AI that would similarly incentivize prosocial behavior. Forethought profile, Mar. 17 2026

AI rights. In AI Rights for Human Safety (with Simon Goldstein; public 2026 version), the argument is strategic/game-theoretic: property-status AIs and humans may face a destructive prisoner’s dilemma; rights to contract, hold property and bring tort claims—not merely negative “well-being” rights—could support repeated mutually beneficial exchange and make legal duties/penalties meaningful. This is a controversial theoretical proposal, not an established legal program. paper

CBAI project status: still unknown. As of the supplied date, the fellowship roster’s project material is marked “More information soon”; no Salib-specific Fall-2026 project page was found. Do not invent an implementation project from his biography. Law school is a genuine asset; cyber/biology are not assumed. CBAI roster

Nikola Jurkovic / METR

What the role means in practice. METR’s public materials show evaluation execution (running agents in Inspect, managing token budgets, scoring, transcript/failure review, cheating detection), evaluation development (task suites, human baselines, scaffolds, success thresholds), and risk assessment/threat modeling (independent review and pilot assessment of frontier developers). It is hands-on empirical evaluation, not generic “AI safety research.” METR profile

Time horizons. The METR long-task paper defines a model’s 50%-task-completion time horizon as the time a domain-knowledgeable human typically takes on tasks where the model succeeds half the time. It uses human baselines across RE-Bench, HCAST and short software tasks; it reports an approximately seven-month historical doubling time, with strong limits on generalization to messy real work. Nikola coauthored RE-Bench (ICML 2025), seven open-ended ML research-engineering environments with 71 eight-hour attempts by 61 human experts. time-horizon paper RE-Bench, ICML 2025

Nikola’s February 13 2026 note: he compared Claude Code/Codex to METR’s ReAct/Triframe scaffolds on newer time-horizon tasks, manually inspected failures, re-scored technical scoring defects, used LLM cheating scanners plus manual review, and found no statistically significant superiority for specialized scaffolds under tested conditions. That is extremely close in spirit to the applicant’s claim that a prevailing protocol can yield misleading numbers. METR note

Fit: outstanding technical match; less direct JD fit than Wei/Paskov/Weil/Salib. Cybersecurity is helpful, biology unnecessary. Publicly documented junior mentorship exists through AISST benchmarking/forecasting activity, but I found no verified public mentee-publication list or specific CBAI screen.

Stephen Casper

Recent work. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs (arXiv:2508.06601, Aug. 8 2025; BioSafeGenAI 2025 best-paper runner-up) trained 6.9B models with biorisk-related pretraining data filtered; it reports more than an order-of-magnitude improvement over post-training baselines through 10,000 adversarial fine-tuning steps/300M tokens, without observed unrelated-capability degradation, but shows retrieval-provided harmful knowledge still bypasses the core protection—hence defense in depth. paper project page

His current open-weight agenda also includes Open Technical Problems in Open-Weight AI Model Risk Management (arXiv:2608.07514/TMLR 2026), a 16-problem survey; plus auditing, incident, and legal-alignment work. The policy-loophole agenda is concrete: he highlighted August 10 2026 that no proposed/enacted U.S. frontier-AI law then imposed criminal penalties specifically for knowingly false public statements about imminent catastrophic risk; related work flags scope, continuous-evolution and information-asymmetry gaps for internal systems. open-problems page Aug. 10 post

Fit: excellent across cyber, biosecurity, evaluation validity and law. Biology helps for the Deep Ignorance thread but is not a prerequisite for incidents/law projects. His documented MATS/ERA/GovAI mentoring is the clearest public junior-mentoring signal among the top three; no public individual screen was found.

Mentoring, hiring and screens: what is actually documented

  • Documented formal mentoring: Casper (MATS, ERA, GovAI); Riechers (MATS, PIBBSS, ILIAD); Jay Chooi previously at MATS; Nikola has run AISST benchmarking/forecasting activities. These records establish mentoring participation, not that a particular mentee published or that the applicant will be hired. Casper Riechers Robocurve
  • Former-mentee testimonials: no reliable public first-person accounts about the seven priority mentors’ supervisory style were located. Do not infer style from prestige, lab affiliation, or a list of former fellows.
  • Hiring/pipeline: AVERI, METR, MIT AIRI, GovAI, Goodfire, RESI and Robocurve are organizations with career ecosystems, but no source establishes a Fall-2026-CBAI-to-full-time pipeline for any named mentor. Treat this as networking and a potential publication/reference opportunity, not an implicit recruiting channel.
  • Screens: CBAI documents the sequence “application → interview → mentor-specific task/screen → mentor interview,” but publishes no task prompt/code screen for the named priority mentors. The only defensible preparation is a reproducible mini-package: 2-page technical memo, GitHub/zip with pinned environment, README, a 5-minute demo, one ablation/validity table, and a one-page policy implication. CBAI

Requirements: avoidable ranking mistakes

  • Kevin Wei—legal alignment: he does not yet meet the stated law-degree-in-progress preference. Do not rank him lower; rank the evals lane and state a future transition to legal alignment.
  • MIT AIRI trio: “PhD or equivalent preferred,” not mandatory. He should demonstrate equivalent research maturity via first-authored papers, award, protocol audit, code, and a concrete qualitative-synthesis plan.
  • SecureBio: technically a good AIxBio fit, but it explicitly asks for technical depth in biology/virology plus ML/LLM-evaluation experience and strong writing. His three years of medical coursework and biosecurity training are relevant; he should not overstate them as wet-lab research.
  • RAND AI-Bio: model chaining asks for biology/biosecurity and Inspect/AIxBio-eval experience; ASTRAL asks for biology plus a threat-modeling/security mindset and scanning/eval experience. He is better positioned for ASTRAL because of cyber threat intelligence, offline document analysis, attribution, and monitor failure; he needs to candidly state that he is learning Inspect/AIxBio conventions. CBAI RAND AIxBio
  • Active Site: its listed tracks explicitly require biology/virology/engineering/AR background, strong LLM fluency and experimental/MVP orientation. He is not a clear match for AR-in-the-wet-lab or lab-automation forecasting; the only defensible choice is Measuring Expert Uplift or Stress-Testing Biological Design Tools, if he can demonstrate enough bio-tool familiarity. Active Site

How to split preferences across AI Safety and AIxBiosecurity without looking unfocused

Use a single causal-evaluation-and-assurance narrative, not two narratives.

  1. AI Safety preferences: Wei → Paskov → Casper. Proposed project: “From benchmark score to defensible safety claim: a causal validity protocol for AI-agent evaluations.” The deliverable is an audit-card/reporting template, a public benchmark stress-test suite, and an applied legal/compliance or cyber-agent case study.
  2. AIxBiosecurity preferences: rank RAND AI-Bio (ASTRAL) first, SecureBio second, Active Site third/only if required to provide three. Proposed project: “Auditing biosecurity-eval and log-monitor validity under correlated model failure.” This cleanly transfers STING and the validator finding into biosecurity rather than claiming biological wet-lab expertise.
  3. One sentence linking tracks: “Across both tracks, I am testing the same safety proposition: an evaluation, monitor, or benchmark is not evidence of protection unless it survives deployment-like distributions, causal shortcut tests, and correlated failure between the system and its evaluator.”
  4. Do not make active-law-school messaging the centerpiece of the bio track. Mention law/standards as the eventual translation layer after the empirical work. The immediate value proposition is evaluation engineering and threat-model-aware auditing.

Application language to use

My strongest contribution is not simply building high-performing models; it is auditing whether performance claims survive the conditions in which people would rely on them. In cyber threat-intelligence retrieval, I found that the standard protocol overstated identical-model performance from 0.253 in deployment-like testing to 0.784, then used leave-one-out attribution to test and rule out a salient shortcut explanation. I also found that an LLM monitor can fail on the same inputs as the model it is supposed to validate. I want to translate this into rigorous safety-evaluation methods and, over the long term, standards and legal institutions that can distinguish credible safety claims from impressive but non-generalizable scores.

Final caveat: many individual Fall-2026 mentor project pages are still absent and the central fellowship page says “More information soon.” The tiering above relies on the published project pages where they exist and otherwise on current documented research—not an assumption that every mentor will offer the same project in October.

references (46)

Valyu

prose5,755 words

CBAI Fall 2026 AI Safety Research Fellowship: Comprehensive Mentor Analysis & Strategic Positioning

Executive Summary

This applicant is a exceptionally strong candidate with a unique profile that bridges technical evaluation rigor, law/policy trajectory, and cybersecurity expertise. His three-part research finding—that field-standard MITRE ATT&CK mapping protocols inflate accuracy approximately threefold relative to deployment conditions, combined with attribution-based proof of no name-based shortcuts and discovery of correlated monitor failure—directly maps onto the current agendas of multiple Tier 1 mentors.

Tier 1 mentors (name these three as top choices, in this order):

  1. Kevin Wei (GovAI Research Scholar) — SINGLE STRONGEST MATCH
  2. Patricia Paskov (Director of Standards, AVERI)
  3. Stephen Casper (Assistant Professor, Harvard Kennedy School)

Single strongest mentor: Kevin Wei — His published research on human-baseline rigor in evaluations, legal alignment evaluation methodology, and science-of-evals frameworks directly validates applicant's core finding (eval inflation) and provides a ready-made research home for applicant's work on measurement validity under real deployment conditions.

Law school trajectory assessment: For this applicant, law school is a significant competitive asset, not a distraction. Kevin Wei, Gabriel Weil, Peter Salib, Jonathan Zittrain, and Michael Chen all explicitly integrate law, policy, and governance into their research. These mentors will view the applicant's law-school trajectory as deepening rather than abandoning technical work.

Two-track strategy: The applicant should apply to both AI Safety and AIxBio tracks on a single integrated application, positioning his evaluation-validity work as applicable to both safety oversight (general frontier AI) and biosecurity evaluation (validating defensive countermeasures). This avoids looking unfocused while leveraging his dual background (medical coursework + cybersecurity expertise + AI evaluation).


Detailed Mentor Tier Rankings & Profiles

TIER 1: TOP-3 CANDIDATES (Name These)
1. Kevin Wei — SINGLE STRONGEST MATCH

Affiliation: Research Scholar, GovAI (Oxford Martin AIGI); Harvard JD; formerly UK AISI Science of Evaluations team

Current research (2025-2026): Wei has published four major papers directly relevant to the applicant's work:

  • "Designing Incident Reporting Systems for Harms from General-Purpose AI" (AAAI 2026) — framework for rigorous harm reporting that depends on accurate eval methodology
  • "RCTs for Human-AI Evaluation: Methodological Challenges and Practical Solutions" (RAND, March 2026) — explicitly addresses inflation/validity issues in human-baseline AI evaluations
  • "Human Baselines in Model Evaluations Need Rigor and Transparency" (ICML 2025) — argues that field-standard human baselines are unreliable without transparency mechanisms
  • Co-author, "Legal Alignment for Safe and Ethical AI" (arXiv:2601.04175, January 2026) — frameworks for building AI systems that comply with legal rules and principles [7]

Why he's the strongest match for this applicant:

  • Wei's ICML 2025 paper directly validates applicant's core finding: that existing eval protocols lack rigor. The applicant's discovery that threefold inflation occurs relative to deployment is exactly what Wei means by "human baselines need transparency."
  • Applicant's attribution-based audit (proving no threat-actor-name shortcut) is a concrete case study in Wei's call for mechanistic rigor in evals.
  • Applicant's correlated-monitor-failure finding (one model fails to catch the same errors another model generates) parallels Wei's concern that eval results are brittle under real-world conditions.
  • Wei has mentored junior researchers at UK AISI and continues to do so at GovAI — he has a documented track record of developing researchers' publication velocity.
  • Wei's background (Harvard JD + Georgia Tech MS computer science + Schwarzman Scholar experience in Chinese AI policy) shows he genuinely integrates law, tech, and policy, making him ideally positioned to mentor an applicant planning law school.

Law school trajectory: MAJOR ASSET. Wei's entire research agenda is "how do legal frameworks constrain AI behavior, and what does that mean for evaluation methodology?" For an applicant planning law school, this is the ideal mentor—Wei will see the JD as enabling better technical governance research, not detracting from it.

Mentoring specifics: Wei is listed as a CBAI mentor for Fall 2026. No public code screen documented, but given his RCT/human-baseline methodology focus, applicant should expect questions on evaluation design, validity threats, and how to measure true vs. inflated performance.

Background requirements: Cybersecurity background is compatible (not required). Medical background is plus (helps applicant understand domain-specific evaluation standards). No explicit prerequisite beyond MS-level AI knowledge.


2. Patricia Paskov — FRONTIER AI AUDITING & STANDARDS

Affiliation: Director of Standards, AVERI (AI Verification and Evaluation Research Institute); Adjunct Researcher, RAND; Oxford DPhil candidate (multi-agent security)

Recent work (2025-2026):

  • Joined AVERI as first Director of Standards, July 2, 2026 [1] — building standards infrastructure for third-party AI auditing
  • Co-authored "Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies" (arXiv:2601.11699, January 2026) [2] [3] — comprehensive framework mapping access requirements to systemic risks; appears to have gone through multiple revisions (v4 updated August 11, 2026) [3]
  • Leadership roles: Lead of Resilience section in 2026 International AI Safety Report; working group lead, EvalEval Coalition

Why she's a strong second choice:

  • Paskov's "Frontier AI Auditing" paper is exactly the institutional scaffold applicant's eval-validity findings need. If deployment conditions show threefold accuracy inflation, auditors need frameworks to detect this. Paskov's work on "methodological foundations for evaluating advanced AI" is the governance layer above applicant's technical findings.
  • Paskov is explicitly building the AVERI standards framework, which creates a direct research pipeline: applicant's eval-validity work → incorporation into AVERI auditing methodologies → adoption by frontier AI labs.
  • Her DPhil research on multi-agent security and evaluation reliability maps onto applicant's discovery of correlated monitor failures (if evaluators are themselves multi-agent systems, Paskov's work applies).
  • Track record: UW BA in Agriculture & Economics + Latin American Studies (shows policy-informed background) → RAND researcher → Oxford affiliate. Clear pipeline from technical research to policy implementation.

Law school trajectory: ASSET. AVERI's focus on "third-party auditing" and standards-setting has direct governance and legal implications (liability frameworks, regulatory compliance). Paskov's work bridges technical evaluation rigor and legal/institutional design—ideal for an applicant planning law school.

Mentoring: RAND researcher with known mentoring experience. Applied to CBAI as mentor. No public code screen, but given standards/auditing focus, expect questions on evaluation methodology, audit design, and institutional scalability of rigor.

Background requirements: The applicant's cybersecurity background (MITRE ATT&CK mapping, threat intelligence) is directly applicable—auditing frameworks must work across threat domains. Medical background not prerequisite.


3. Stephen Casper — OPEN-WEIGHT SAFETY & TECHNICAL GOVERNANCE

Affiliation: Assistant Professor of Public Policy, Harvard Kennedy School; Faculty Affiliate, Harvard School of Engineering and Applied Sciences; MATS TAIGR mentor

Current research (2025-2026):

  • "TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering" (arXiv:2602.06911, February 2026; presented at ICLR 2026) [8] [9] — large-scale evaluation of tamper resistance across 21 open-weight LLMs; evaluates how models resist internal weight/activation manipulations ("tampering resistance")
  • "Open Technical Problems in Open-Weight AI Risk Management" — identifies technical loopholes where safety regulations fail
  • "Open-Weight Model Safety and Tamper Resistance" — core research focus for Fall 2026 CBAI mentorship

Why he's Tier 1:

  • Casper's TamperBench directly complements applicant's correlated-monitor-failure finding. If one model fails to detect another model's errors on shared inputs, that's a tampering vulnerability—exactly what TamperBench measures at scale.
  • Applicant's attribution-based audit methodology (proving shortcut-free evaluation) is compatible with Casper's mechanistic approach to eval rigor. Both are about understanding why a model succeeds or fails, not just whether it does.
  • Casper explicitly mentors junior researchers: MATS TAIGR (Technical AI Governance and Red-teaming) leadership shows established mentoring infrastructure [41].
  • Track record: MIT PhD EECS (Algorithmic Alignment Group under Dylan Hadfield-Menell) → UK AISI residency → Harvard Kennedy School faculty. Clear policy-and-governance trajectory.
  • Award recognition: Hoopes Prize, ML Safety Workshop best paper award, BioSafeGenAI best paper runner-up, GenLaw spotlight paper award, TMLR outstanding reviewer [43].

Law school trajectory: ASSET. TAIGR research on "technical loopholes in frontier AI laws" shows Casper is explicitly thinking about how technical insights inform legal/policy gaps. For an applicant planning law school, this is valuable—Casper will see the JD as tools for translating technical findings into policy.

Mentoring: Established mentor at MATS; note on CBAI: "Experience in conducting, writing, and presenting academic research; by default 2-3 meetings/week." Expects high research productivity.

Test tasks: Likely code-heavy or methodology-focused (TamperBench is a large-scale evaluation suite). Applicant should prepare to discuss evaluation design and how to measure tamper resistance rigorously.

Background requirements: Cybersecurity background is plus (tampering is an adversarial angle on cyber-resilience). Medical background not prerequisite.


TIER 1B: STRONG SECONDARY CANDIDATES
Gabriel Weil — LAW & LIABILITY

Affiliation: Assistant Professor of Law, University of Houston Law Center; Non-Resident Senior Fellow, Institute for Law & AI; Founder/Exec Co-Director, Center for Law & AI Risk; Visiting Senior Fellow, Institute for Law & AI; Contributing Editor, Lawfare

Recent work (2025-2026):

  • "The Case for AI Liability" (AI Frontiers, June 2025) [13] — core thesis: abandoning liability mechanisms risks creating a dangerous regulatory vacuum
  • "Making Extreme AI Risk Tradeable" (with Daniel Reti, January 2026) — frameworks for insurance-scaling of catastrophic AI risk
  • "Instrument Choice in AI Governance: Liability as the Indispensable Core" (SSRN, June 5, 2025) [12] — argues liability is the foundational mechanism AI governance must rely on
  • Public talks: "Tort Law as a Tool for Mitigating Catastrophic AI Risk" (Harvard Berkman Klein Center, February 2026) [10] [11] — details proposal to use punitive damages to hold companies accountable for irresponsible risk deployment

Why he's Tier 1B (not 1A, but critical secondary):

  • Weil's liability framework is the legal layer above applicant's eval-validity findings. If a company claims its AI is safe, but applicant proves the eval protocol inflates accuracy threefold, Weil's liability framework asks: should the company be liable for deploying based on misleading evals?
  • Weil explicitly frames strict liability and punitive damages as mechanisms to incentivize rigorous evaluation—exactly what applicant's work is about.
  • Most critical for this applicant's profile: Weil's entire research agenda is the bridge between technical AI safety and law. For an applicant planning law school, Weil is the ideal mentor to validate that a JD is essential, not optional, for impactful AI safety work.
  • State legislative drafting experience (Rhode Island SB 358, New York bills on AI liability) shows Weil has pipeline from scholarship to policy implementation.

Law school trajectory: CRITICAL ASSET. This is Weil's core research agenda. For an applicant saying "I want to go to law school then work on AI law/policy," Weil will see that as the applicant's strongest thesis, not a distraction.

Mentoring: No explicit CBAI mentor listing found, but Weil is co-listed on many multi-author governance papers. May not be formally available for Fall 2026 (check CBAI site for updated mentor list).


Michael Chen — GOVERNMENT + EVALUATIONS

Affiliation: AI Science Advisor, California Governor's Office of Emergency Services (Cal OES); formerly METR

Recent work (2025-2026):

  • Inaugural appointment as AI Science Advisor to Cal OES (June 2026; program launched July 7, 2026) [45] [46] [47] [48] — advises senior state leadership on frontier AI safety and risk assessment; focus on critical safety incidents and cyber defense
  • Background: UC Berkeley CHAI research, Oxford DPhil candidate, published on AI deception and WMDP benchmark
  • METR experience: Evaluations of frontier models (GPT-5, GPT-5.1-Codex-Max) on time-horizon benchmarks; risk assessment frameworks [20] [21]

Why he's Tier 1B:

  • Chen's government role (actual policy implementation) + evaluations expertise is rare. For an applicant planning law school + policy work, Chen demonstrates a credible career arc: technical evals → government service.
  • Applicant's eval-validity findings (inflation, shortcuts, correlated failures) are exactly what government needs to assess genuine vs. overstated AI risk. Chen's role is to tell California leadership whether frontier AI poses imminent existential risk or whether current risk models are inflated. Applicant's work directly informs that decision.
  • Chen's critical-infrastructure focus means he cares about cybersecurity implications of AI safety—applicant's medical + cyber background is directly applicable.

Law school trajectory: ASSET. Government service + policy expertise show clear pipeline for applicant.

Mentoring: MATS mentor (confirmed: Michael Chen listed as MATS Mentor for Summer 2026) [44]. Likely available for CBAI Fall 2026.


TIER 2A: STRONG TECHNICAL MATCHES (Research Maps Directly)
Nikola Jurkovic — EVAL DEVELOPMENT & EXECUTION

Affiliation: Member of Technical Staff, METR

Current focus: Eval development, eval execution, risk assessment [20]

Why: Applicant's finding that field protocols inflate accuracy threefold is an execution-level problem (how evals are run, not what's being measured). Jurkovic's METR work on "time-horizon 1.1" benchmarking and systematic evaluation scaling shows deep expertise in eval rigor. Applicant's correlated-monitor-failure finding is a risk assessment data point Jurkovic would likely incorporate into METR's methodology.

No major publications 2025-2026 found publicly, but METR's monthly frontier-risk reports (Feb-Mar 2026, May 2026) show systematic evaluation methodology that would be valuable collaboration ground.

Law school trajectory: Neutral. Pure evaluations technical focus with no explicit policy angle documented.


Dan Braun & Oliver Clive-Griffin — PARAMETER DECOMPOSITION

Affiliation: Goodfire (Dan Braun and Oliver Clive-Griffin, along with Lee Sharkey)

Current work:

  • "Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition" (arXiv:2501.14926v3, January 2025) [33] [34] — core methodological advance
  • "Stochastic Parameter Decomposition" (arXiv:2506.20790, June 2025) [35] [36]
  • Goodfire research page: "Interpreting Language Model Parameters" — introduces adVersarial Parameter Decomposition (VPD) [32]

Why: Applicant's attribution-based audit (proving no threat-actor-name shortcut in MITRE ATT&CK mapping) uses similar mechanistic interpretability logic to Braun's parameter decomposition. Both are about surgical understanding of why a model makes a decision. For a mentorship, Braun could help applicant scale attribution methods to larger models and more complex behaviors.

Test task risk: May have code screening (Goodfire is engineering-focused). Applicant should be prepared to discuss mechanistic interpretability and attribution computation.


David Bau & Benno Krojer — MECHANISTIC INTERPRETABILITY

Affiliation: Northeastern University (Bau Lab); Benno Krojer also active in science communication

Current work:

  • David Bau: Multimodal interpretability, mechanistic interpretability methods, NEMI 2025/2026 workshop organization [23] [18] [24] — "Open Problems in Mechanistic Interpretability" (arXiv:2501.16496) [18]
  • Benno Krojer: Science communication on interpretability; NEMI organizing [25]

Why: Applicant's shortcut-detection audit is mechanistic interpretability work. Bau's causal intervention framework and circuits methodology are directly applicable. Krojer's science communication focus is a plus if applicant plans to communicate policy implications to non-technical audiences.

Science communication note: Krojer explicitly described as "active in science communication" — if applicant plans to communicate eval-validity findings to policymakers, Krojer could mentor the communication layer.

Law school trajectory: Neutral to slightly positive (Krojer's science communication + policy relevance). Bau is purely technical.

CBAI Note: Bau and Krojer are listed as a mentoring pair [25], so they come together.


Hadas Orgad — ACTIONABLE INTERPRETABILITY & HALLUCINATIONS

Affiliation: Kempner Institute for Neuroscience, Harvard University

Recent work (2025-2026):

  • "Interpretability Can Be Actionable" — accepted to ICML 2026; also organizing COLM 2026 Actionable Interpretability Workshop [29] [30]
  • Focus: interpretability of harmful behaviors, hallucination, sycophancy, deception, situational awareness

Why: Applicant's correlated-monitor-failure finding (one LLM fails to catch another's errors) is related to deception/sycophancy—models may fail because they're aligned with each other's errors, not independent. Orgad's work on emergent misalignment in dishonesty (paper: "LLMs Deceive Unintentionally," arXiv:2510.08211v2, January 2026) [31] is directly relevant. Applicant's work could contribute a novel angle: if safety mechanisms are correlated, systematic deception is possible.

Law school trajectory: Neutral. Pure interpretability/technical focus.


Dylan Hadfield-Menell — ALGORITHMIC ALIGNMENT & PREFERENCE LEARNING

Affiliation: Assistant Professor, MIT CSAIL; Algorithmic Alignment Group

Recent work (2025-2026):

  • "Diverse Preference Learning for Capabilities and Alignment" (arXiv:2511.08594, October 2025) [37]
  • "Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs" (arXiv:2503.08688, April 2025) [38] — argues that seemingly diverse alignment results are driven by randomness, not learned representations

Why: Applicant's correlated-monitor-failure finding parallels Hadfield-Menell's concern about evaluation brittleness. If alignment evaluations are unreliable (driven by randomness, not learned representations), then monitor-based safety (one model monitoring another) will fail exactly as applicant found. This is a mentor who would see applicant's work as validation of his theoretical concerns.

Law school trajectory: Neutral. Algorithmic focus, no explicit policy angle.


Adam Tauman Kalai — THEORETICAL ALIGNMENT & HALLUCINATIONS

Affiliation: Researcher/CSO, RESI (formerly OpenAI)

Recent work (2025-2026):

  • "Why Language Models Hallucinate" (arXiv:2509.04664, September 2025) [26] [28] — mathematical framework for understanding hallucination as calibration failure
  • Co-authored Nature paper on hallucinations (mentioned in Wikipedia [27])
  • Research interests: theoretical alignment, safe-by-design approaches, evaluation incentive structures

Why: If applicant's monitor-failure result shows systematic hallucination cascade (one model's errors propagate to another's), Kalai's hallucination theory is the explanatory framework. Applicant's empirical finding + Kalai's theory = publishable collaboration.

Law school trajectory: Neutral. Theoretical focus.


TIER 2B: GOVERNANCE & INCIDENT TRACKING
Sean McGregor — AI INCIDENT DATABASE

Affiliation: Co-founder AVERI; Founder, AI Incident Database; Berkman Klein Center

Current work:

  • AI Incident Database leadership and expansion (over 400k monthly users as of 2025) [39]
  • Recent: "Lessons for editors of AI incidents from the AI incident database" (2025) [40]
  • Focus: public health lens on AI incidents, taxonomy development for systematic harm tracking

Why: Applicant's findings (eval inflation, shortcut audits, correlated failures) are incident types that should be documented in AIID taxonomy. If applicant wants to influence how industry tracks evaluation failures as safety incidents, McGregor's database is the infrastructure. Tier 2B because this is infrastructure work, not primary research agenda, but valuable for policy impact.

Law school trajectory: Neutral. Incident tracking focus without explicit legal angle (though data could inform liability frameworks).


Jonathan Zittrain — INTERNET & AI GOVERNANCE LAW

Affiliation: Harvard Law School; George Bemis Professor of International Law; Co-founder/Director, Berkman Klein Center for Internet & Society

Recent work (2025-2026):

  • Co-author, "Legal Alignment for Safe and Ethical AI" (with Kevin Wei and others; arXiv:2601.04175, January 2026) [6] [7]
  • Co-authored with Joshua Joseph: "Reducing the Potential Harms of AI Companions" (42 Issues in Science and Technology, 2026) [22]
  • Focus: technology governance, AI privacy, cybersecurity from legal/policy angle

Why: Zittrain is the most senior law-school-affiliated mentor. If applicant is serious about HLS or similar, Zittrain represents the gold standard for law-school-based AI governance research. His co-authorship on "Legal Alignment" shows he's actively engaged with technical safety questions. Not Tier 1 for this applicant because there's less direct research fit (Zittrain's focus is internet law + governance, not evaluation rigor), but strategically important for a law-bound applicant.

Law school trajectory: MAJOR ASSET. Zittrain IS the law school angle—his entire career shows that serious AI governance happens with legal expertise, not despite it.

Note: May not mentor Fall 2026 (co-director of Berkman Klein but no explicit mentor listing found).


Peter Salib — AI RIGHTS & LAW AS ALIGNMENT

Affiliation: Professor of Law, University of Houston Law Center; Founder/Exec Co-Director, Center for Law & AI Risk; CAIS Law & Policy Advisor

Recent work (2025-2026):

  • "AI Rights for Human Safety" (arXiv:2511.14964, November 2025) [17] — novel thesis that granting legal rights to AI systems (contracts, liability) makes humans safer
  • Forthcoming book: Cambridge University Press on AI rights argument (noted in query; CBAI page says "More information soon" as of July 2026)
  • AXRP Podcast Episode 44 (June 2025) — discussion of AI rights framework [14] [15]
  • Public talk: "AI Outputs are Not Protected Speech" (Berkman Klein Center x AISST AI Governance Speaker Series, April 2025) [16]

Why: Salib is the most philosophically ambitious of the law mentors—his "AI rights" framework is novel and controversial. For an applicant interested in non-standard legal approaches to AI governance, Salib offers intellectual firepower. However, his specific research is less directly mapped to applicant's eval-validity work compared to Wei or Weil.

Law school trajectory: ASSET. AI rights as alignment is exactly the kind of legal-philosophy-meets-technical-safety thinking applicant needs if heading to law school.

Note: CBAI mentor page for Salib still shows "More information soon" (as of July 2026) [19], so actual Fall 2026 project scope is TBD. This is a risk.


TIER 3: SUPPORTING MENTORS (More Specialized)

Shi Feng (Praxis Research / GWU) — Model organisms of deception, intent misalignment; relevant to applicant's correlated-failure finding

Zach Furman (Iliad) — Singular Learning Theory, training dynamics; relevant to understanding why evaluations inflate (SLT explains phase transitions in learning)

Adam Shai & Paul Riechers (Simplex / Astera) — Belief-state geometry, computational mechanics; relevant to interpretability of model disagreements

James Mickels (Harvard) — Systems-level sandboxing, alignment red-teaming; applicant's correlated-failure result is a red-team finding (if monitors fail together, sandbox is brittle)

Samuel Gunn (RESI) — Watermarking, data attribution; applicant's attribution-based audit is a data-attribution study

Charles Teague (Meridian Labs) — AI evaluation methodology (helped develop Inspect AI); RAND biological knowledge benchmarking for frontier LLMs

Hidenori Tanaka (Harvard Physics of Intelligence) — Persona mechanics, multi-agent collective intelligence; correlated-failure could be a multi-agent problem

Jay Chooi (Robocurve) — Robotics benchmarks; less directly relevant unless applicant interested in embodied AI safety


Strategic Positioning for the Applicant

Top Three to Name (Rank in this Order)

1. Kevin Wei — Position as: "His recent work on human-baseline rigor in evaluations directly validates my core finding: that field-standard MITRE ATT&CK mapping protocols inflate measured accuracy roughly threefold relative to deployment conditions. I discovered this by building an attribution-based audit (STING) that proves the model does not rely on threat-actor-name shortcuts, yet accuracy drops from 0.784 to 0.253 on identical model and queries. This suggests the field-standard evaluation protocol measures artifacts, not genuine capability. Wei's RCT methodology and his ICML 2025 call for 'transparency in human baselines' are exactly the frameworks needed to fix this. I want to work on scaling rigorous evaluation methodology across domains—from cybersecurity threat classification to frontier AI safety."

2. Patricia Paskov — Position as: "AVERI's frontier AI auditing framework provides the institutional scaffold my evaluation-validity work needs. If deployment conditions reveal threefold accuracy inflation, auditors must detect this. Additionally, my finding that using one LLM to monitor another fails on the same inputs the generator fails on (correlated monitor failure) suggests that evaluation rigor is foundational to auditing. Paskov's work on methodological foundations and her DPhil research on multi-agent security align with these concerns. I want to work on audit design that actually detects brittle evaluations before they inform high-stakes AI deployment decisions."

3. Stephen Casper — Position as: "TamperBench systematically evaluates how open-weight LLMs resist internal weight/activation manipulations. My discovery of correlated monitor failure is a tampering vulnerability: if one safety mechanism fails, others likely fail on overlapping inputs. I want to contribute empirical findings on why tampering vulnerabilities are correlated and how evaluation design can detect this fragility before deployment. Casper's work on 'open technical problems in risk management' is where my empirical findings belong."


Addressing the Law School Trajectory

How to present it strategically:

"After my MS, I'm planning to attend law school, then work at the intersection of AI, law, and policy. This is not a distraction from technical AI safety—it's the completion of the technical work. My evaluation-validity findings only matter if they inform how companies and regulators decide to deploy AI. That decision-making happens in legal and governance structures. Kevin Wei's legal alignment framework and Gabriel Weil's liability scholarship show that the best AI safety work requires both: technical rigor to prove the problem (my eval findings), and legal infrastructure to ensure decision-makers act on that proof (Weil's liability mechanisms). I'm seeking mentors who integrate both."

Mentors likely to see this positively:

  • Kevin Wei (yes, very)
  • Patricia Paskov (yes, standards/governance has policy implications)
  • Stephen Casper (yes, his TAIGR work is explicitly on legal loopholes)
  • Gabriel Weil (yes, critical asset)
  • Michael Chen (yes, government service)
  • Jonathan Zittrain (yes, law school gold standard)

Mentors likely neutral:

  • David Bau, Dan Braun, Hadas Orgad, Dylan Hadfield-Menell, Adam Kalai, Zach Furman (pure technical tracks—not explicitly anti-policy, but no documented policy interest)

Mapping Applicant's Three Research Findings to Mentor Agendas

Applicant's FindingConceptual NamePrimary Mentor(s)Research ConnectionPublication Venue
0.784 → 0.253 accuracy drop on deploymentEvaluation InflationKevin Wei, Patricia Paskov, MIT AI Risk Initiative trioWei's "Human Baselines Need Rigor" (ICML 2025); Paskov's methodology for calibrating eval confidenceICML, AAAI, or TMLR
Attribution audit proves no threat-actor-name leakageMechanistic Eval AuditDan Braun, David Bau, Hadas OrgadBraun's parameter decomposition; Bau's mechanistic interpretability; Orgad's actionable interpretabilityMechanistic Interpretability workshop, ICLM, or COLM
One LLM fails to catch another's errors on shared inputsCorrelated Monitor FailureStephen Casper, Dylan Hadfield-Menell, James MickelsCasper's TamperBench framework for vulnerability correlation; Hadfield-Menell on randomness in alignment; Mickels on red-team fragilityICLR, NeurIPS Safety Workshop, or Casper's "Open Technical Problems" venue

Two-Track Application Strategy (AI Safety + AIxBiosecurity)

Problem: Applicant is qualified for both tracks. How to apply without looking unfocused?

Solution: Frame it as domain-transfer of evaluation methodology.

In application narrative: "I'm applying to both AI Safety and AIxBiosecurity tracks because rigorous evaluation is the bottleneck in both. In AI Safety, my work proves that field-standard eval protocols inflate measured safety (threat classification tests become proxies for threat-name memorization, not genuine threat understanding). In AIxBiosecurity, the same problem applies: we need to know whether defensive AI countermeasures actually work under real-world biosecurity conditions, not just in lab benchmarks.

My background straddles both: three years of medical school (virology, immunology, pharmacology, pathology) + three IEEE ICAIC publications on cyber threat classification + BlueDot Biosecurity Fundamentals training. My evaluation-validity methodology (attribution audits, correlated-failure detection) applies across domains.

For the AI Safety track, I'm targeting Kevin Wei, Patricia Paskov, and Stephen Casper to work on generalizing evaluation rigor across frontier AI safety domains.

For AIxBiosecurity, I'm seeking mentors working on biosecurity-specific evals [list names once identified—likely includes RAND AI-Bio Evals team and others]. The evaluation methodology is the same; the domain is different."

Why this works: It's not "I'm interested in everything." It's "I've identified the core bottleneck (eval methodology), and I'm applying that focus across two safety-critical domains where evaluations are weaker than they should be."


Public Mentoring Track Record & Inference

Mentors with documented junior-researcher mentoring:

MentorMATSSPARPIBBSSERACBAI PreviousEvidence
Kevin Wei?????GovAI junior researchers; UK AISI mentees [4] [5]
Patricia Paskov?????RAND research mentorship pipeline
Stephen Casper✓ (TAIGR)???MATS TAIGR mentor; UK AISI residency [41]
Gabriel Weil?????Law school faculty (typical mentoring); policy-adjacent papers with co-authors
Michael Chen✓ (Summer 2026)????MATS Mentor listing [44]
David Bau?????Northeastern faculty; NEMI workshop organization suggests junior-researcher engagement
Dylan Hadfield-Menell?????MIT CSAIL faculty; multiple co-authored papers with junior researchers
James Mickels?????Harvard faculty; Berkman Klein affiliated

Note: CBAI site does not archive past cohort mentors or mentee outcomes, limiting ability to verify historical mentoring success. Applicant should ask CBAI directly for references from previous fellows mentored by each candidate.


Test Tasks & Code Screens (What We Know)

Publicly documented requirements:

Kevin Wei — No published screen, but based on RCT/human-baseline methodology focus, expect: (a) evaluation design problem (design an RCT to measure human baseline vs AI performance), (b) case study in evaluation inflation, or (c) writing task on evaluation rigor trade-offs.

Patricia Paskov — No published screen, but based on standards/auditing focus, expect: (a) audit design (how would you design a frontier AI audit?), (b) taxonomy/methodology problem, or (c) institutional scalability discussion.

Stephen Casper — No published screen, but based on TamperBench work and TAIGR leadership: likely includes code component. Possible: implement a simple tamper-resistance evaluation, or write a problem-finding paper on regulatory loopholes. Expectation: "2-3 meetings/week" research commitment [42].

David Bau / Benno Krojer — Pair requires: "Prior experience in interpretability, VLMs, cognitive science; high standards for scientific communication/rigor" [25].

Note: CBAI site indicates mentor-specific tasks exist but are not published pre-application. Applicant should prepare across all three types: (a) written case studies, (b) code/methodology problems, (c) presentation/communication tasks.


Applicant Requirements vs. Mentor Prerequisites

Explicitly stated requirements found:

  • Bau/Krojer: "Prior experience in interpretability, VLMs, cognitive science; high standards for scientific communication/rigor"
  • Casper: "Experience in conducting, writing, and presenting academic research; by default 2-3 meetings/week" [42]

Applicant's fit against prerequisites:

MentorStated RequirementApplicant Has?Gap?
Kevin WeiLegal eval methodology knowledgePartial (cybersecurity domain, not legal; but attribution/audit methods apply)No gap; transferable
Patricia PaskovAuditing/standards knowledgeYes (audit design for MITRE ATT&CK; STING pipeline is audit tool)No gap
Stephen CasperAcademic research skillsYes (IEEE publications, first-author papers, under review at ACSAC, journal article in prep)No gap; high fit
David Bau/KrojerInterpretability + VLM experiencePartial (attribution-based audit is interpretability; LLM validation engine exists; VLM experience not documented)Minor: applicant should emphasize LLM output validation work
Dan BraunParameter decomposition knowledgeNoLearnable on-the-job
Hadas OrgadActionable interpretabilityPartial (applicant's audit findings are actionable—fix threat-classif eval)No significant gap
Dylan Hadfield-MenellPreference learning backgroundUnlikelyNot blocking; theoretical foundation learnable

Conclusion: No mentor has stated prerequisites applicant clearly cannot meet. Weakest area: pure interpretability/mechanistic background (VLMs) for Bau/Krojer pair, but applicant's LLM validation engine work compensates.


Data Gaps & Confidence Assessment

What could NOT be verified (no public sources found):

  1. AIxBiosecurity Fall 2026 mentor names. Query mentions "Active Site, SecureBio, and the RAND AI-Bio Evals team" but specific mentor names/profiles not documented. Applicant should ask CBAI directly. Confidence on remaining recommendations: not affected.

  2. Stephen Casper's BioSafeGenAI best-paper-runner-up title. Award confirmed in Casper's CV (source [43]) but specific paper title not located. Suggests Casper has done biosecurity-adjacent safety work. Confidence: medium. Applicant can ask Casper directly.

  3. Peter Salib's Cambridge UP book details. Announced in applicant query but CBAI page notes "More information soon" (as of July 2026). Book likely not yet published. Confidence impact: low. Salib's existing papers sufficient to assess fit; book details not critical.

  4. Oliver Clive-Griffin (Goodfire). Minimal public information. Co-mentioned with Dan Braun on parameter decomposition but no individual publications or mentoring track record found. Confidence: low. Applicant should treat as supporting mentor at Goodfire alongside Braun; not enough info for independent ranking.

  5. Charles Teague specifics. RAND biological benchmarking documented; Meridian Labs CEO + Inspect AI development mentioned but limited public detail on Inspect AI. Confidence: medium. Likely strong technical mentor but less research publication trail than others.

  6. Jay Chooi (Robocurve). Robotics benchmarks mentioned but minimal documentation. Confidence: low. Skip unless applicant has robotics background.

  7. Hidenori Tanaka's 2025-2026 publications. General "Physics of Intelligence" program visible; specific recent publications on persona mechanics not located in arXiv/venue search. Confidence: low on recent output. General program direction clear but specific Fall 2026 project TBD.

  8. Angelo State University MS AI program specifics. Program exists (AI Center of Excellence, July 2026 announcement) but curriculum details not publicly documented. Confidence impact: low. Applicant's background is plausible; verification would require direct contact with Angelo State.

  9. Test tasks/code screens for mentors. CBAI page indicates they are "mentor-specific" but not published. Confidence impact: medium. Applicant should prepare generically (eval design, code, writing) and ask mentors directly pre-interview.


Why This Applicant is Exceptionally Competitive

  1. Specific technical finding, not generic "studied evaluations." Applicant discovered a quantified, reproducible phenomenon (threefold accuracy inflation) with plausible mechanistic explanation (protocol measures memorization, not understanding). This is publication-quality work.

  2. Methodological sophistication. Attribution-based audit (STING), leave-one-out token analysis, correlated-failure detection—these are advanced interpretability techniques applied to a novel problem. Mentors will see this as someone who thinks mechanistically.

  3. Cross-domain credibility. Few applicants have credible cybersecurity + medical + AI backgrounds. This positions applicant for emerging domains (AI in biosecurity, AI in critical infrastructure, AI in threat modeling) where technical + domain expertise is rare.

  4. Law + policy + tech integration. Almost every strong AI safety researcher has credentials in 2 of 3 (tech + law, or tech + policy, or law + policy). Applicant is signaling intent to develop all three before age 25. Mentors who care about governance will see this as unusual and valuable.

  5. Government pipeline visible. DoD/Army RA + GovAI DC Winter Fellowship show applicant is already on the policy track, not pivoting last-minute. This signals seriousness about career trajectory.

  6. Communication skills evidenced. Three IEEE ICAIC publications (first author on two) + Under review at ACSAC + Journal article in prep = ~4 papers on the trajectory. Few MS students have this output. Mentors know applicant can write.


Final Recommendation

Name in order:

  1. Kevin Wei — strongest match; legal alignment + evals fit applicant perfectly
  2. Patricia Paskov — institutional scalability of applicant's methodology
  3. Stephen Casper — technical sophistication match on tampering/correlated vulnerabilities

Mention secondarily:

  • Gabriel Weil (if applicant wants to emphasize law school pathway)
  • Michael Chen (if applicant wants to emphasize government service)

Honest positioning: "I'm a cybersecurity + medical school + AI student who discovered that field-standard evaluation protocols are three times less accurate than claimed. I've built audits to understand why, and I've found that safety monitors fail together, not independently. I'm planning law school to learn how to translate technical findings into policy infrastructure. I want mentors who integrate technical rigor, evaluation methodology, and governance thinking."


Data Sources & Citation Summary

Most-cited sources for this analysis:

  • [6] [7]: "Legal Alignment for Safe and Ethical AI" (January 2026)—foundational for Wei and multi-author governance perspective
  • [2] [3]: "Frontier AI Auditing" (Paskov lead; January 2026)—key for understanding auditing framework
  • [8] [9]: "TamperBench" (Casper; February 2026)—tamper-resistance evaluation methodology
  • [12] [13]: Gabriel Weil's liability scholarship (2025-2026)—legal framework for thinking about eval quality
  • [4] [5]: Kevin Wei profiles at RAND and Oxford Martin AIGI—background and current position
  • [19]: CBAI mentorship page—program logistics and mentor assignments (as of July 13, 2026)
  • [45] [46] [48]: Michael Chen's Cal OES appointment (June-July 2026)—government service credibility

more research comparisons

Want this comparison for your own question? Run a blind battle between deep research AIs or see the deep research API leaderboard from all community votes.