| metric | Parallel | Exa | Valyu |
|---|---|---|---|
| format | prose | prose | prose |
| word count | 8,350 | 3,936 | 5,755 |
| sources | 118 | 18 | 80 |
| processing time | 1288s | 0s | 0s |
| has images | no | no | no |
| has tables | no | no | no |
| citation style | — | — | — |
Name these three, in order: (1) Kevin Wei, (2) Patricia Paskov, (3) Stephen Casper.
Single strongest: Kevin Wei.
That ordering follows from the applicant’s unusually direct evidence of evaluation validity (deployment-vs-paper protocol gap, shortcut/attribution audit, and correlated monitor failure), his demonstrated ability to build reproducible offline tooling, and his explicit JD-to-AI-law/policy goal. Wei is the rare mentor whose published agenda joins all three: evaluation methodology, legal alignment, and technical governance. CBAI specifically names ecological validity, eval awareness, model-spec/constitution compliance, legal hierarchy, legal retrieval, deterrence, certification, and U.S.–China governance as his project space; it requires Python/ML/research experience for evals, which the applicant plainly has. His legal-alignment lane formally prefers a law degree in progress, so frame the applicant as an MS researcher deliberately building the empirical record before the JD—not as already qualified for a doctrinal-law project. CBAI—Wei RAND profile, updated Apr. 18 2026
Kevin Wei — best overall.
Patricia Paskov — best standards/auditing fit.
Stephen Casper — strongest technical-governance/bio bridge.
Nikola Jurkovic (METR): best choice if he wants a rigorous, code-heavy evaluation project and possibly a METR-adjacent research signal. His work is execution, threat modeling and forecasting—not primarily law. Emphasize reproducible experiments, human baselines, external validity, scanner failure analysis, and cyber-agent tasks. METR profile
Peter Slattery / Alexander Saeri / Michael Noetel (MIT AI Risk Initiative): strongest route to a policy-facing systematic review/database/taxonomy paper. The formal project asks for literature-review/qualitative-synthesis ability, AI governance familiarity, and says a PhD/equivalent is preferred. The applicant’s publication record, analytical audit and intended JD make him plausible, but his pitch must demonstrate careful coding protocol and synthesis—not just ML implementation. CBAI MIT AIRI project page
Gabriel Weil: strongest pure private-law/liability mentor. Good if the applicant wants to turn evidence from evaluation failures into a paper about standards of care, negligence, punitive damages, disclosure, or evidentiary use of evaluations. Lower than Wei/Paskov because this is a bigger pivot away from empirical evaluation and because no specific Fall project has been posted.
Peter Salib: excellent long-run intellectual match for a JD and AI governance, particularly law-as-alignment. Lower for this application because CBAI’s project information remains “More information soon” and there is no documented Fall-2026 task to target. CBAI roster
Sean McGregor / Charles Teague: strong alternatives for benchmark/incident infrastructure. McGregor’s publicly described CBAI direction is scaling the breadth and depth of incident databasing, and his recent work includes BenchRisk (NeurIPS 2025); Teague brings Inspect AI and scientific-publishing/tooling experience. They are especially attractive if the applicant frames STING as an auditable, offline evaluation/incident-analysis artifact. McGregor CBAI page Teague CBAI page
James Mickens / Samuel Gunn: credible cyber-security-adjacent technical safety alternatives. Mickens is particularly relevant to sandboxing and systems security; Gunn to watermarking/data attribution. They fit the applicant’s cyber profile but do not make the JD trajectory as central as Tier 1.
| Applicant finding | Closest current mentor agenda | Why |
|---|---|---|
| 3.1× apparent-performance inflation under the field-standard protocol | Kevin Wei (strongest); Paskov; Jurkovic; MIT AI Risk Initiative | Wei explicitly names ecological validity and evaluation methodology; Paskov works on proportional, credible evaluations and assurance; METR’s time-horizon work makes benchmark results interpretable in human task-time terms; MIT AIRI catalogues/evaluates risk mitigations. |
| Attribution audit finds no threat-actor-name shortcut | Benno Krojer/Bau/Orgad, then Wei | This is exactly a causal/shortcut-detection contribution. Krojer’s MVP benchmark exists because superficial cues can inflate evaluation scores. |
| One LLM monitoring another fails on the same inputs | Wei and Jurkovic, then Casper | This is a correlated-failure / judge-validity / oversight-robustness result. Wei’s science-of-evals work is the cleanest home; Jurkovic’s METR work includes failure analysis and scanner review; Casper works on audit and safeguard robustness. |
One unifying application thesis: “I study when AI-safety and cyber-AI evaluations look valid but fail under realistic conditions—and how to build causal, deployment-grounded evidence that can support standards, audits, and eventually law.” That thesis connects the cyber record, STING, validation engine, biology course background, and JD plan without pretending they are separate careers.
What “law-following AI” means. In the legal-alignment literature, it is not merely a model refusing illegal requests. It is a program of making AI systems comply with legitimate legal rules, use legal interpretation methods in reasoning, and use legal concepts as structural tools for reliability/trust/cooperation. The 2026 survey, coauthored by Wei, calls these the three research pathways of legal alignment. Legal Alignment for Safe and Ethical AI, arXiv:2601.04175, Jan./June 2026
Documented 2025–26 output: Position: Human Baselines in Model Evaluations Need Rigor and Transparency (ICML 2025); Methodological Challenges in Agentic Evaluations of AI Systems (ICML Technical AI Governance workshop, 2025); Infrastructure for AI Agents (TMLR 2025); Designing Incident Reporting Systems for Harms from General-Purpose AI (AAAI 2026); plus work on rigorous GPAI evaluations and RCT-style human-uplift studies. RAND
Model specs/constitutions. The public CBAI project menu explicitly asks how well models comply with model specs/AI constitutions, whether model and human interpretations differ, whether systems obey legal hierarchies, recognize legal implications, retrieve the right texts, and respond to deterrence-like penalties. This is project agenda, not necessarily a published Wei paper. The relevant adjacent empirical literature includes SpecEval (arXiv:2509.02464, Sept. 2025), which audits 16 models against provider behavior specifications and reports sizeable three-way specification/output/judge consistency gaps; Wei should not be credited as its author. CBAI—Wei SpecEval
Fit and prerequisite: perfect for evals; partial for legal alignment until JD begins. No documented public code screen. The published CBAI process says only that there is a mentor-specific task/screen after interview.
What AVERI is building. AVERI describes itself as building an independent third-party auditing layer for frontier AI. The January 2026 Frontier AI Auditing report defines this as third-party verification of developer safety/security claims and evaluation of their systems and practices against standards, with deep, secure access to non-public information. It proposes four AI Assurance Levels: AAL-1 as a present baseline and AAL-2 as the near-term objective for leading developers; higher levels require less reliance on company representations and more organization-wide scrutiny. AVERI report arXiv:2601.11699
“PCAOB/FINRA analogue” — inference, not an announced institution. A credible analogue would set auditor competence/independence rules, manage conflicts and cooling-off periods, specify access/security procedures for sensitive model/training/governance data, establish assurance-report formats, inspect or discipline audit providers, and make assurance levels intelligible to regulators and the public. That inference is grounded in AVERI’s explicit independence, deep-access, quality and anti-checkbox principles—not in a public claim that AVERI already is a PCAOB/FINRA-equivalent regulator. The AI Evaluator Forum is described as an emerging assessment-organization venue helping articulate access standards; it is not documented as a statutory regulator. AVERI legislative landscape, Apr. 20 2026
Fit: excellent. Biology is not presumed for her general auditing/standards work; statistics, measurement, writing, reproducibility and policy literacy matter more. There is no public evidence of a specific test task or an individual fellowship-to-AVERI hiring pipeline.
Documented CBAI project: systematic review of AI risk mitigations and systematic document review of organizational responses to AI risks. CBAI asks for strong literature review/qualitative synthesis and AI governance/policy familiarity; it says PhD/equivalent preferred. CBAI project page
Methodology and tooling. Their Mapping AI Risk Mitigations (arXiv:2512.11931, Dec. 12 2025) conducted a rapid evidence scan of 13 frameworks (2023–25), extracted 831 distinct mitigations, and created a four-category/23-subcategory draft taxonomy. It tested LLM assistance but found LLMs unreliable for fully automating extraction (confabulation, combining and omission); LLM classification suggestions were useful only with document-level manual comparison, author review, and multi-author consistency checks. The resulting data are public in an Airtable-backed database and interactive taxonomy. That is the correct sense of “LLM-plus-human-validated”—not an automatic taxonomy classifier. paper interactive taxonomy database
Taxonomy. The four top categories are Governance & Oversight, Technical & Security, Operational Process, and Transparency & Accountability controls; the 23 subcategories include risk management, model alignment/safety, testing/auditing, incident handling, disclosure and third-party assurance. Their related AI Risk Repository is a 1,725-risk meta-review and was published in Patterns in 2026. MIT FutureTech
Fit: strong if he reframes his evidence as a mini systematic-review/coding problem: classify evaluation failure modes and mitigations, preregister a codebook, quantify coder/LLM disagreement, and create a usable audit schema. No biology/cybersecurity prerequisite; limited direct evidence of junior-mentee outcomes or a full-time pipeline.
Current scholarship. The documented progression is: Tort Law as a Tool for Mitigating Catastrophic Risk from AI (SSRN 2024); Instrument Choice in AI Governance: Liability as the Indispensable Core (June 5 2025); Overcoming Judgment-Proofness: The Law & Economics of Insuring and Mitigating AI Risk (SSRN 2026); and Abnormally Dangerous Algorithms: The Case for Strict Liability at the AI Frontier (SSRN 2026). instrument-choice abstract strict-liability preprint insurance/judgment-proofness preprint
Punitive damages / uninsurable risk. Weil’s 2025 thesis is that strict, ex-post liability is comparatively calibrated to realized risk and incentivizes safety innovation; where compensatory damages cannot capture catastrophic stakes, punitive damages in compensable “near miss” cases associated with uninsurable risk could supply deterrence.
Counterargument—careful qualification. I did not locate a peer-reviewed, AI-specific published rebuttal squarely answering Weil’s punitive-damages proposal. The strongest documented general counterpoint is the mature audit/liability literature’s warning that assurance/liability systems have independence, expectations-gap, sensitive-information and box-ticking problems; a punitive regime additionally faces the classic under-deterrence problem if actors are judgment-proof and the pricing/information problem if insurers cannot observe or quantify tail risk. These objections are reasons to treat ex-post damages as complementary rather than sufficient, not evidence that Weil has been decisively refuted. Weil himself recognizes supportive, complementary and substitutionary non-liability policy. AVERI report Weil abstract
Fit: the applicant’s empirical work is valuable to Weil if positioned as evidence for standards of reasonable care, foreseeability, safety representations, or punitive-damages predicates. Law trajectory is an unequivocal asset. No specific Fall project, public screen, or hiring pipeline found.
Current agenda. Salib describes law as an alignment technology: rule-of-law systems already give powerful misaligned actors such as corporations/states incentives against harmful conduct; he is developing legal rights/duties frameworks for advanced AI that would similarly incentivize prosocial behavior. Forethought profile, Mar. 17 2026
AI rights. In AI Rights for Human Safety (with Simon Goldstein; public 2026 version), the argument is strategic/game-theoretic: property-status AIs and humans may face a destructive prisoner’s dilemma; rights to contract, hold property and bring tort claims—not merely negative “well-being” rights—could support repeated mutually beneficial exchange and make legal duties/penalties meaningful. This is a controversial theoretical proposal, not an established legal program. paper
CBAI project status: still unknown. As of the supplied date, the fellowship roster’s project material is marked “More information soon”; no Salib-specific Fall-2026 project page was found. Do not invent an implementation project from his biography. Law school is a genuine asset; cyber/biology are not assumed. CBAI roster
What the role means in practice. METR’s public materials show evaluation execution (running agents in Inspect, managing token budgets, scoring, transcript/failure review, cheating detection), evaluation development (task suites, human baselines, scaffolds, success thresholds), and risk assessment/threat modeling (independent review and pilot assessment of frontier developers). It is hands-on empirical evaluation, not generic “AI safety research.” METR profile
Time horizons. The METR long-task paper defines a model’s 50%-task-completion time horizon as the time a domain-knowledgeable human typically takes on tasks where the model succeeds half the time. It uses human baselines across RE-Bench, HCAST and short software tasks; it reports an approximately seven-month historical doubling time, with strong limits on generalization to messy real work. Nikola coauthored RE-Bench (ICML 2025), seven open-ended ML research-engineering environments with 71 eight-hour attempts by 61 human experts. time-horizon paper RE-Bench, ICML 2025
Nikola’s February 13 2026 note: he compared Claude Code/Codex to METR’s ReAct/Triframe scaffolds on newer time-horizon tasks, manually inspected failures, re-scored technical scoring defects, used LLM cheating scanners plus manual review, and found no statistically significant superiority for specialized scaffolds under tested conditions. That is extremely close in spirit to the applicant’s claim that a prevailing protocol can yield misleading numbers. METR note
Fit: outstanding technical match; less direct JD fit than Wei/Paskov/Weil/Salib. Cybersecurity is helpful, biology unnecessary. Publicly documented junior mentorship exists through AISST benchmarking/forecasting activity, but I found no verified public mentee-publication list or specific CBAI screen.
Recent work. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs (arXiv:2508.06601, Aug. 8 2025; BioSafeGenAI 2025 best-paper runner-up) trained 6.9B models with biorisk-related pretraining data filtered; it reports more than an order-of-magnitude improvement over post-training baselines through 10,000 adversarial fine-tuning steps/300M tokens, without observed unrelated-capability degradation, but shows retrieval-provided harmful knowledge still bypasses the core protection—hence defense in depth. paper project page
His current open-weight agenda also includes Open Technical Problems in Open-Weight AI Model Risk Management (arXiv:2608.07514/TMLR 2026), a 16-problem survey; plus auditing, incident, and legal-alignment work. The policy-loophole agenda is concrete: he highlighted August 10 2026 that no proposed/enacted U.S. frontier-AI law then imposed criminal penalties specifically for knowingly false public statements about imminent catastrophic risk; related work flags scope, continuous-evolution and information-asymmetry gaps for internal systems. open-problems page Aug. 10 post
Fit: excellent across cyber, biosecurity, evaluation validity and law. Biology helps for the Deep Ignorance thread but is not a prerequisite for incidents/law projects. His documented MATS/ERA/GovAI mentoring is the clearest public junior-mentoring signal among the top three; no public individual screen was found.
Use a single causal-evaluation-and-assurance narrative, not two narratives.
My strongest contribution is not simply building high-performing models; it is auditing whether performance claims survive the conditions in which people would rely on them. In cyber threat-intelligence retrieval, I found that the standard protocol overstated identical-model performance from 0.253 in deployment-like testing to 0.784, then used leave-one-out attribution to test and rule out a salient shortcut explanation. I also found that an LLM monitor can fail on the same inputs as the model it is supposed to validate. I want to translate this into rigorous safety-evaluation methods and, over the long term, standards and legal institutions that can distinguish credible safety claims from impressive but non-generalizable scores.
Final caveat: many individual Fall-2026 mentor project pages are still absent and the central fellowship page says “More information soon.” The tiering above relies on the published project pages where they exist and otherwise on current documented research—not an assumption that every mentor will offer the same project in October.
The ranking below is not a ranking of research quality. It is a ranking of expected fit given three constraints: the applicant's demonstrated evaluation work, his intended move toward law and policy, and his need to preserve a credible technical research identity.
| Rank/tier | Mentor | Why this fit is strong or weak | Main uncertainty or risk |
|---|---|---|---|
| 1, Tier A | Kevin Wei | Exact overlap with evaluation methodology, legal-alignment evaluations, model-spec/constitution compliance, and technically informed US governance. | Must demonstrate statistical evaluation literacy and avoid presenting as primarily a cybersecurity or mechanistic-interpretability applicant. [55] |
| 2, Tier A | Patricia Paskov | Direct match to evaluation validity, independent auditing, assurance standards, structured knowledge systems, and public-facing governance work. | The applicant has limited documented assurance, standards-body, credentialing, and external-partnership experience. [47] |
| 3, Tier A | Gabriel Weil | Best law-school fit; explicitly welcomes non-lawyers who can read law and connect technical failure modes to legal institutions. | Less direct technical supervision fit than Wei or Paskov; recent public output is more legal scholarship and commentary than empirical ML. [42] |
| 4, Tier A/B | Stephen Casper | Strong match to evaluation failure, tamper resistance, incidents, technical AI governance, and loopholes in frontier AI laws. | He will likely expect the applicant to continue doing technically serious work rather than mainly prepare for law school. [40] |
| 5, Tier B conditional | Peter Salib | Excellent long-term law/AI intellectual fit: law as a mechanism for shaping AI behavior and AI rights/duties. | CBAI topics and desired qualifications remain "More information soon," so the actual Fall 2026 project is unknowable. [25] [14] |
| 6, Tier B | Nikola Jurkovic | Applicant's deployment-validity work and cyber background fit METR-style task design, execution, and risk assessment. | The mentorship page supplies no detailed qualifications or project list; law-policy fit is indirect. [30] [11] |
| 7, Tier B | Jonathan Zittrain | Strong policy, digital-law, writing, editing, and project-management fit; useful bridge to public communication. | Less obvious direct fit to the applicant's empirical evaluation work. [46] |
| 8, Tier B | Alexander Saeri | Strong fit through AI risk taxonomies, Delphi work, and evidence synthesis; good policy and communication bridge. | Individual mentoring style and project allocation within the MIT group are not documented. [23] [56] |
| 9, Tier B | Michael Noetel | Good fit to the repository, risk classification, human judgment, and empirical research infrastructure. | Same group-level uncertainty; little mentor-specific evidence. [23] |
| 10, Tier B | Peter Slattery | Good fit to AI risk infrastructure, classification, and governance mapping. | Less direct evidence about his individual current agenda or supervision. [3] [23] |
| 11, Tier B | Michael Chen | Cyber resilience, emergency response, loss of control, and critical infrastructure are plausible uses of the applicant's CTI background. | The applicant's law trajectory is not central to the listed topics. [34] |
| 12, Tier B/C | Sean McGregor | Incident database, incident learning, AVERI, and agentic-workgroup work fit validity and harm measurement. | Public material establishes his organizations and projects, not a CBAI-specific research plan or mentoring record. [36] |
| 13, Tier B/C | Charles Teague | Inspect AI and Meridian Labs offer practical evaluation-tooling exposure. | Less direct law/standards fit and no verified recent publication or mentor-specific pipeline in the reviewed material. [37] |
| 14, Tier B/C | Benno Krojer | Interpretability plus science communication is unusually relevant to the applicant's stated public-facing goals. | The technical center of gravity is mechanistic interpretability, not governance or standards. A 2026 public listing identifies work titled "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs," but I found no public CBAI mentee record. |
| 15, Tier C | David Bau | Strong interpretability, causal intervention, and tooling background; the applicant's attribution work gives a credible technical bridge. [52] | Law and policy are not the natural center of the lab; the applicant would need to commit to technical interpretability. |
| 16, Tier C | Hadas Orgad | Very good match to attribution, hidden failures, harmful behavior, and actionable interpretability. [59] [59] | The applicant's evaluation work is relevant, but the law-school trajectory could look like a near-term departure from interpretability. |
| 17, Tier C | James Mickens | Systems-level sandboxing and cyber experience are plausible technical matches. | No clear law or standards project in the reviewed material; applicant would need a systems-security project. |
| 18, Tier C | Jay Chooi | Benchmarks, model evaluation, disempowerment, and robotics/automation could use the applicant's evaluation instincts. [56] [56] | Robotics and labor-market work are not the applicant's strongest demonstrated areas. |
| 19, Tier C | Dylan Hadfield-Menell | Broad alignment, policy, incentives, and human-AI interaction agenda. [26] | The applicant has not shown a direct value-learning, recommendation, or embodied-intelligence project. |
| 20, Tier C | Shi Feng | Deception, intent misalignment, model organisms, and scalable oversight are important technical topics. [69] | No direct match to the applicant's law/standards goal unless he pivots to deception evaluations. |
| 21, Tier C | Samuel Gunn | Data attribution and watermarking could connect to the applicant's attribution audit. | I found no independently verified 2025-26 publication list, CBAI project detail, or mentoring evidence for this specific mentor. |
| 22, Tier C | Adam Tauman Kalai | Safe-by-design and theoretical alignment could benefit from a technically mature applicant. | The applicant has not demonstrated a theoretical alignment agenda, and law-policy fit is indirect. |
| 23, Tier C/D | Dan Braun | Parameter decomposition and interpretable-by-design networks could use attribution experience. | Requires a substantial interpretability pivot; no public mentor-specific 2025-26 project or mentoring record found. |
| 24, Tier C/D | Oliver Clive-Griffin | Targeted parameter decomposition is technically adjacent to attribution. | Narrower technical fit and no evident law/standards channel. |
| 25, Tier D conditional | Adam Shai | Representation science, belief-state geometry, and interpretability are intellectually interesting. [57] [57] | The applicant would need to make a serious technical pivot; law school is not an obvious asset. |
| 26, Tier D conditional | Paul Riechers | Computational mechanics and belief-state geometry offer a principled theory of intelligence. [57] [57] | Strong mathematical/physics orientation, weak direct fit to the applicant's governance path. |
| 27, Tier D | Zach Furman | Singular learning theory and training dynamics could suit a mathematically oriented applicant. | No demonstrated mathematical learning-theory background or policy connection. |
| 28, Tier D conditional | Hidenori Tanaka | Persona mechanics, collective intelligence, and representation work are scientifically interesting. [39] | A law-school trajectory and evaluation-validity framing are unlikely to be the natural match without a major technical reframing. |
Recommended order:
This order is better than putting Casper third if the applicant wants CBAI to understand that law school is a destination for applied research rather than an escape from technical work. If the application form or essay is clearly optimized for the AI Safety technical track, use Wei, Paskov, Casper, and mention Weil as the law-policy alternative.
The central one-sentence positioning should be:
I study whether AI evaluations measure the property they claim to measure, how to audit the evidence against shortcuts and correlated monitors, and how valid evidence can become an enforceable safety and governance standard.
That sentence gives all three top mentors a reason to say yes without pretending that the applicant is already a lawyer, an auditor, or a frontier-model safety researcher.
Wei's public description is unusually close to the applicant's interests. The CBAI page names three areas: science of evaluations, legal-alignment evaluations, and technically informed AI governance, with a particular interest in US implementation and evaluation-related policy. [55] His broader profile describes work on evaluation methodology, legal AI safety/alignment, regulation, and liability, with affiliations including Oxford and RAND and previous work at the UK AI Security Institute. [5] [6]
The most relevant recent work publicly associated with him includes:
The important point is that Wei is not merely a lawyer who happens to work on AI. His agenda asks how to construct, validate, and govern evaluations, including evaluations of legal compliance.
In this context, "law-following AI" means an empirical research program for testing whether an AI system can and will comply with legal requirements. The CBAI page describes automated assessments of an agent's abilities and propensities to comply with legal requirements. [55] This is narrower and more operational than asking whether an AI is generally "aligned with the law."
The likely research questions include:
CBAI explicitly lists empirical work on model-spec and AI-constitution compliance. [55] That makes the applicant's shortcut audit especially relevant: a model that appears compliant because it uses a superficial cue is analogous to a CTI model that appears accurate because it keys on threat-actor names.
This is the cleanest match to all three of the applicant's evaluation findings:
The applicant should not pitch a project as "AI and cybersecurity law" first. He should pitch one of these projects instead:
Wei's requirements are unusually explicit. For evaluation work he wants Python, prior ML research or software/data-science experience, and ideally statistics, causal inference, and familiarity with Inspect, Inspect Scout, or HiBayes. For legal-alignment work, a candidate may qualify through the evaluation route or through legal training. Governance projects favor prior public-policy experience, especially technically informed or quantitative work. [55] [55]
The applicant appears to satisfy the evaluation route through first-author research, Python implementation, audit construction, and a concrete validity result. He should not claim to satisfy the law-degree alternative: he is planning law school, not already in law school. He can satisfy the technical route and present law school as the next stage of impact.
I found no public evidence that Wei supervised a named MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or previous CBAI cohort, and no reliable former-mentee account about his day-to-day style. His prior AISI and RAND work is evidence of relevant professional research, not proof of a fellowship-to-job pipeline. No Wei-specific CBAI-to-GovAI hiring pipeline is documented.
Paskov's current public agenda is broader than generic "AI evaluations." She describes herself as building methods and standards for evaluating and verifying the safety and security of frontier AI systems. [16] She is associated with RAND's frontier-evaluation and policy work, the 2026 International AI Safety Report's Resilience section, and the EvalEval Coalition. [19] Her profile also says she became an Oxford DPhil candidate in Engineering Science in Fall 2026. [16]
Recent listed work includes:
AVERI is a US nonprofit launched in January 2026 to make frontier-AI auditing effective and universal. Its premise is that safety and security claims are largely self-reported by AI companies and need independent verification and common standards. It conducts research and pilot audits intended to inform audit standards, policy, and tooling. [4] [4]
Its frontier-auditing report, published January 15, 2026 and available as arXiv:2601.11699, proposes a system with:
CBAI specifically lists "building a PCAOB/FINRA analogue for AI auditing" as a possible project. [47]
What that would involve is partly documented and partly an inference. The documented AVERI ingredients are standards, access governance, conflict-of-interest controls, assurance levels, independence, continuous monitoring, and public-facing audit results. [47] [50]
An analogue could therefore require:
This is not evidence that AVERI has already settled on that institutional design. It is the natural research space implied by the CBAI topic and AVERI's published audit architecture.
The AI Evaluators Forum appears in the CBAI topic list as an expansion and governance project. [47] I found no public, finalized charter that would allow a more specific description of its institutional design.
Paskov is probably the best mentor for the applicant's evaluation inflation plus attribution audit story. The applicant has already done the core intellectual move that assurance work needs: he questioned a widely used measurement, reproduced the result under a more realistic protocol, and investigated whether the model used a shortcut.
The strongest application evidence for Paskov would be:
Paskov's desired qualifications explicitly include analytical writing, structured knowledge systems, public dissemination, proactive communication, partnerships, frontier-evaluation knowledge, CS/ML or hands-on evaluation experience, and ideally auditing, assurance, standards, credentialing, or self-regulation experience. [47]
The applicant clearly has the ML/evaluation, writing, and public-output pieces. His gaps are formal audit/assurance practice, professional standards-body experience, and external partnership management. These are gaps, not stated disqualifiers.
I found no public evidence that Paskov supervised a named MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or CBAI cohort. Her AVERI role does create a plausible employment connection, but not a documented fellowship-to-full-time pipeline. AVERI did advertise a project-manager role involving pilot-audit operations and refinement of audit methodology, with requirements including three or more years of project, audit, assurance, consulting, or comparable execution experience. [41] [41] [41] That is evidence of organizational hiring, not evidence that CBAI fellows are recruited into AVERI.
The MIT project is often described informally as LLM-plus-human taxonomy classification. The public methodology is more careful. The AI Risk Repository is a living database containing roughly 1,725 risks extracted from 74 frameworks and classifications. It uses two complementary taxonomies: a causal taxonomy and a domain taxonomy. [23] [23]
The causal taxonomy asks how, when, and why a risk occurs, including the responsible entity, intentionality, and lifecycle timing. The domain taxonomy organizes impacts into seven domains and 24 subdomains. [17] [17] [23] Each risk is coded according to the definitions of the relevant taxonomy, with risks retained as presented by the source authors rather than being silently rewritten into the researchers' preferred theory. [23] [23]
The workflow is primarily human evidence synthesis:
The authors explicitly state that they have not conducted a formal validation study of whether independent users can reliably classify novel risks with the taxonomies. [23]
For the separate AI Risk Mitigation Database, the team manually extracted 831 mitigations from 13 frameworks and iteratively developed a four-part, 23-subcategory taxonomy:
They tried LLM assistants for document extraction and classification, but found that they missed mitigations, generated spurious entries, and made classification errors. The team then audited the extractions manually, had a team member review classifications, and cross-checked classifications. The final draft classified 815 of 831 mitigations, or 98 percent. The database is available through Airtable. [43] [58] [58]
So the accurate answer is:
The project uses AI-assisted evidence work in some places, but its published classification method is not an autonomous LLM classifier certified by humans. The central claim is human-reviewed, reproducible evidence synthesis, with explicit warnings about validation limits.
That distinction is highly relevant to the applicant's finding that one model cannot reliably monitor another on the same failure distribution.
The main recent outputs are:
Saeri is the most obvious fit for evidence synthesis and Delphi work; Noetel for empirical measurement and human judgment; Slattery for risk infrastructure and governance mapping. That allocation is an inference from the group outputs, not a documented CBAI division of labor.
The applicant should approach the trio with a project such as:
This is a good second-tier choice because it uses the applicant's existing work without requiring him to abandon technical rigor or pretend to be a lawyer.
I found no reliable public evidence of individual supervision by Slattery, Saeri, or Noetel at MATS, SPAR, PIBBSS, ERA, LASR, Algoverse, or CBAI, and no individual hiring pipeline. CBAI's published cohort outcomes show that the inaugural cohort produced a NeurIPS Mechanistic Interpretability Workshop spotlight paper and other accepted work, but those are program-level outcomes and do not establish which mentor supervised which paper. [1] [1]
Weil's CBAI topics are liability and insurance, verification through insurance, punitive damages for uninsurable risks, administrative penalties, digital minds governance, and AI judges. His page says analytic quality matters more than credentials, a law degree is not required, and he expects a fellow to read cases, statutes, and law reviews and to move between AI failure modes and actual legal institutions. [42]
The principal public works I found are:
His current institutional role remains focused on liability as a tool for managing catastrophic AI risk, and he consults with policymakers. [15]
The documented argument is more nuanced than "liability solves AI risk." Weil says liability may be the centerpiece of AI governance for risk externalities: developers and customers capture benefits while third parties bear much of the risk. But he identifies several limits:
CBAI's reference to punitive damages and administrative penalties is important because those mechanisms could operate even where ordinary compensatory insurance cannot. But the public CBAI material does not itself supply the full argument or a finished 2025-26 article.
I did not find a published paper that directly refutes Weil's central claim that private compensatory liability cannot handle truly uninsurable catastrophic risk. The strongest relevant counterposition is instead a complement: "Insuring Uninsurable Risks from AI: Government as Insurer of Last Resort", arXiv:2409.06672, argues for a public backstop where private insurance and defendants' assets are insufficient.
That paper is not a clean rebuttal. It largely accepts the judgment-proof and insurance problem and responds with ex ante regulation, public insurance, or a government backstop. The best critique to raise with Weil is therefore not "tort liability is useless," but:
If a risk is both uninsurable and unlikely to produce a warning shot, why should the primary policy instrument be punitive liability rather than ex ante licensing, capital requirements, mandatory safety evidence, administrative sanctions, or a public backstop?
That question is especially suitable for the applicant because it connects evaluation validity to legal incentives: if the evidence supplied to regulators or insurers is invalid, every downstream liability mechanism is weakened.
Weil is the mentor most likely to regard law school as a positive signal rather than an exit. He expressly says a JD is not required and values people who can connect technical failures to legal institutions. [42]
The applicant should pitch an empirical legal-policy project, not a generic essay on AI ethics. Good examples are:
Weil serves on the PIBBSS board, but that is not evidence that he personally supervised PIBBSS fellows. [15] I found no verified public record of his supervision of the named fellowship programs, no reliable former-mentee account, and no documented CBAI-to-LawAI or university hiring pipeline.
Salib's framework is the most direct version of "law as an alignment technology" in this mentor list. His public description says legal institutions can reduce catastrophic and societal-scale risks from highly capable AI, and that legal and economic incentives can shape AI behavior in addition to technical safety training. [14]
His AI-rights argument is instrumental rather than merely a claim that AIs deserve moral status. The proposal is that granting advanced systems certain rights and duties, analogous in some respects to legal entities such as corporations, could give them institutional and economic reasons to behave safely. [14]
Recent or forthcoming work publicly listed includes:
The critical uncertainty is not Salib's intellectual fit. It is project availability. Both the Fall 2026 CBAI page's Mentor topics and Desired fellow qualifications still say "More information soon." [25] I found no later CBAI project description that resolves this. Treat Salib as a high-upside conditional choice, not as a known project match.
METR's time-horizon work measures the task duration, defined by human expert completion time, at which an agent is predicted to succeed with a specified probability. It is not a measure of how long an agent remains autonomous. METR fits a logistic curve to success as a function of human task duration and reports, for example, 50 percent and 80 percent time horizons. [11] [11] [11]
The current process is concrete:
METR's broader current work includes frontier capability evaluations, domain variation in time horizons, monitorability evaluations, risk reports, and reviews of developer sabotage or automated-R&D risk assessments. [38] [35]
Thus CBAI's shorthand "eval development, eval execution, and risk assessment" means, concretely:
This is a real fit for the applicant's evaluation-validity work and CTI/cyber background. It is weaker for his law-school plan unless he proposes a project on how time-horizon or agentic-risk measurements should inform policy, incident thresholds, or assurance standards.
CBAI lists only the topics and says "More information soon" for the rest of Jurkovic's page. [30] I found no public evidence of named fellowship supervision, former-mentee accounts, or a fellowship-to-METR pipeline. METR has a general careers presence, but that is not the same as a CBAI hiring channel. [38]
Casper's Fall 2026 topics are unusually explicit: open-weight safety and tamper resistance, predicting and preventing major AI incidents, and navigating technical ambiguities and loopholes in frontier AI laws. His desired qualification is demonstrated tenacity and self-directed execution; he gives independently writing a paper as an example of a strong signal. [40]
His current role is not simply interpretability. He works on safeguards, incidents, and governance; leads a MATS research stream; mentors for ERA and GovAI; and contributes to the International AI Safety Report and Singapore Consensus. [10]
Relevant recent work includes:
The applicant's tamper-resistance fit is not that he has already studied open-weight safeguards. It is that his attribution audit and monitor-failure result show the habit Casper explicitly values: testing whether a safety claim survives adversarial or non-ideal conditions.
Casper is the clearest documented mentor in the list for MATS, ERA, and GovAI. [10] His MATS stream expects academic research, writing and presentation, frequent group meetings, iterative project refinement, and a clear theory of impact. [31] I found no public list that reliably attributes particular published mentee papers to him, and no CBAI-to-full-time pipeline.
| Applicant finding | Best mentor | Why |
|---|---|---|
| Standard protocol inflated accuracy from 0.784 to 0.253 under deployment-like conditions | Kevin Wei; Patricia Paskov | It is a construct-validity and external-validity problem, then an auditing and standards problem. Wei's science-of-evals agenda is the direct match; Paskov can turn the result into reporting and assurance requirements. [55] [47] |
| Attribution audit showed no reliance on threat-actor-name shortcuts | Patricia Paskov; Kevin Wei; Hadas Orgad | It tests whether an evaluation result reflects the intended capability rather than a spurious cue. Paskov supplies the assurance frame; Wei supplies evaluation methodology and possible legal-compliance analogues; Orgad supplies technical interpretability and causal intervention. [47] [59] |
| One model monitoring another fails on the same inputs where the generator fails | Paskov; Casper; Jurkovic/METR | This is evaluator dependence, monitorability, and correlated failure. Paskov's independence and audit-governance agenda is the closest standards match; Casper's safeguards agenda is the closest technical-governance match; METR provides the evaluation-execution discipline. [47] [40] [11] |
The applicant should use the phrase correlated evaluator failure rather than only "LLM-as-a-judge failure." It connects directly to audit independence, monitorability, and the reliability of evidence used in policy.
This is an inference from the public agendas, not a statement by the mentors:
The applicant should never say, "I am doing technical work until law school." He should say, "I am developing the empirical and engineering foundation I will use in law and policy work."
The following table records what was publicly verifiable in the reviewed material. "No verified item found" means I did not find a reliable 2025-26 paper or preprint tied to that mentor, not that the person has done no work.
| Mentor | What appears current / recent public work | Mentoring, preference, background, and pipeline assessment |
|---|---|---|
| Jonathan Zittrain | CBAI describes ethics and governance of AI, agents, digital property, privacy, intermediaries, and education. [46] | Strong writing/editing and digital-law fit; CBAI wants a current student or relevant worker and favors digital law/policy. [46] No verified individual fellowship pipeline or former-mentee account found. Biology is unnecessary; law and communication are central. |
| Michael Chen | CBAI topics are emergency response preparedness, loss of control, critical infrastructure risk, and generative/agentic AI. [34] | Applicant's cyber background is useful. CBAI explicitly wants an excellent writer, frontier-safety knowledge, technical judgment, and reliability. No verified recent publication or mentor-specific fellowship record found. |
| Sean McGregor | AVERI cofounder, AI Incident Database executive director, and leader in the ML Commons Agentic Workgroup; public work emphasizes incident learning and AI risk. [36] | Good fit for incident reporting and public communication. PIBBSS or AVERI institutional association should not be confused with proof of direct supervision. No verified CBAI-to-job pipeline found. |
| Charles Teague | Meridian Labs is a nonprofit building open tools for understanding, evaluating, and testing models and agents; Inspect AI is among its projects. [37] | Good practical tooling fit. No verified 2025-26 publication list or mentor-specific supervision evidence. Law school is secondary; Python/evaluation implementation would matter. |
| David Bau | Bau Lab studies the structure and interpretation of deep networks and operates or contributes to NNsight/NDIF-style interpretability infrastructure. [52] [52] | Applicant's attribution work is a plausible bridge. Public lab material advertises PhD/research recruitment, not a CBAI fellowship pipeline. Law and biology are not prerequisites; deep-learning/interpretability work is. |
| Benno Krojer | A public 2026 listing identifies "LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs." His CBAI description also emphasizes science communication. | One of the better conditional choices for the applicant's communication goal. No verified named fellowship mentees or pipeline found. No biology/law prerequisite; technical interpretability plus communication are the likely requirements. |
| Dan Braun | Goodfire's public agenda centers on parameter decomposition and interpretable-by-design networks. | Attribution experience helps, but this requires an interpretability pivot. No verified 2025-26 paper, fellowship record, former-mentee account, or hiring pipeline found. |
| Oliver Clive-Griffin | Goodfire work is described around targeted parameter decomposition and related decomposition tools. | Narrow technical fit; no law or standards channel apparent from reviewed material. No verified mentor-specific publication or pipeline found. |
| Hadas Orgad | Current work concerns interpretable realistic, high-level behavior. Recent listed work includes harmful-response mechanisms, hidden failures in robustness, actionable interpretability, hidden factual knowledge, hallucination representations, and position-aware circuit discovery at ACL 2025. [59] [59] | The applicant's attribution and evaluator-failure work are technically relevant. No public CBAI-specific mentoring record found. Law school is likely a distraction unless the project is framed as actionable evaluation or safety intervention. |
| Shi Feng | Praxis/GWU work includes deception, intent misalignment, model organisms, scalable oversight, and recent belief-extrapolation or motivation-profile evaluation work. [69] | Applicant would need to pivot from validity auditing in CTI to deception/control evaluations. No verified named fellowship supervision or hiring pipeline found. |
| Adam Shai | Simplex studies representations and belief-state geometry as part of a principled science of intelligence. [57] [57] | Public MATS mentor page confirms a MATS mentoring role. [27] Strong technical/math pivot, weak law fit; no biology required. |
| Paul Riechers | Simplex work uses computational mechanics and theoretical-physics ideas to study belief states and intelligence. [57] [57] | No law or biology prerequisite, but substantial theory/physics orientation. No separate public fellowship-supervision record found. |
| Zach Furman | Iliad agenda is described around singular learning theory, training dynamics, and mathematical foundations. | Requires a stronger mathematical learning-theory profile than the applicant has shown. No verified 2025-26 title or mentoring pipeline found. |
| Hidenori Tanaka | Public group material lists work on emergent abilities, mathematical representations, persona/collective intelligence, and related neural dynamics; listed papers include "Forking Paths in Neural Text Generation" (ICLR 2025, arXiv:2412.07961) and "Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering" (arXiv:2511.00617). [39] | Public group material indicates a MATS mentor connection. [39] Strong technical pivot, little direct law or assurance fit. |
| Dylan Hadfield-Menell | Algorithmic Alignment Group research emphasizes conceptual understanding, algorithms, policy, value learning, incentives, recommendation, debugging, and human-AI interaction. [26] | Broad policy connection, but no direct public match to the applicant's existing evaluation project. No specific 2025-26 item or mentor-specific pipeline verified. |
| James Mickens | Publicly associated systems work includes "Guillotine: Hypervisors for Isolating Malicious AIs," HotOS 2025. | Cyber/systems applicant could be credible, but law-school fit is indirect. No verified fellowship supervision or CBAI-to-job pipeline found. |
| Samuel Gunn | CBAI bio identifies RESI work in watermarking and data attribution. | Applicant's attribution audit is relevant, but I found no independently verified 2025-26 paper list, detailed CBAI project page, or mentoring record. |
| Adam Tauman Kalai | RESI work is associated with safe-by-design and theoretical alignment, with prior OpenAI experience. | Applicant would need to show theory or safe-by-design research rather than only applied audit work. No verified 2025-26 publication or fellowship pipeline found. |
| Jay Chooi | Robocurve is building open-source tools and independent benchmarks around physical automation and labor-market disruption. Recent listed work includes "Covert Influence Between Language Models" (arXiv:2606.04071), "Efficient Ensemble Selection from Binary and Pairwise Feedback" (arXiv:2605.09588), and "Measuring AI-Induced Disempowerment: A Framework and Proposed Metrics." [56] [56] | Evaluation and benchmark experience fit, and his former MATS fellowship is documented, but that does not establish mentor supervision. [56] Robotics is not the applicant's strongest domain. |
The public AIxBio CBAI page names Active Site, SecureBio, the RAND AI-Bio Evals team, Protectome, and individual RAND/Harvard/Active Site people, but the individual descriptions remain "More Information." [18] The track-level description covers dangerous-capability evaluations, safeguards and standards, unlearning in biological models, biosurveillance, pathogen detection, DNA synthesis screening, access governance, and dual-use publication. [18]
AI Safety track:
AIxBio track:
The RAND AI-Bio Evals team is the best first AIxBio choice because the applicant brings evaluation methodology, deployment-validity concerns, attribution auditing, and some biological background. Active Site is a strong second choice for applied biosecurity evaluation and standards. SecureBio should be ranked highly if its project is evaluation, governance, or assurance; it is harder to rank precisely while the public page says only "More Information."
Use the same causal story in both tracks:
High-consequence AI decisions depend on evaluations. I want to make those evaluations resistant to shortcuts, deployment mismatch, and correlated evaluators, then translate the resulting evidence into standards and governance.
For AI Safety, the application can instantiate that story in legal alignment, model specifications, open-weight safeguards, and frontier evaluation. For AIxBio, instantiate it in biological capability evaluations, agentic bio-risk measurements, and independent assurance. Do not present the tracks as "I want to do law, interpretability, and wet-lab biosecurity." Present them as two application domains for the same evaluation-and-assurance problem.
The medical-school coursework in virology, immunology, pharmacology, and pathology is useful for reading bio-risk literature and recognizing domain assumptions. BlueDot training helps demonstrate serious interest. Neither establishes wet-lab competence, pathogen engineering expertise, or professional biosecurity experience. The applicant should be explicit that his comparative advantage is evaluation and audit methodology applied to biological-risk questions, not experimental biology.
I found no mentor-specific fellowship coding challenge or take-home screen for Wei, Paskov, Weil, the MIT trio, Jurkovic, Salib, or Casper.
The closest public signals are:
A reasonable preparation exercise, clearly labeled as preparation rather than a known requirement, would be a two-page audit protocol plus a small reproducible Python experiment showing benchmark inflation, shortcut sensitivity, and generator-monitor error correlation.
The applicant should optimize for a research identity that survives law school, not for a choice between "technical AI safety" and "law." His strongest evidence is not simply that he has published in cybersecurity or built an offline attribution pipeline. It is that he independently discovered a serious measurement failure, quantified it, investigated shortcut dependence, and found a failure mode in model-based monitoring itself.
That is why Kevin Wei is the single strongest mentor: the applicant already has a concrete science-of-evaluations result and wants to extend it into legal alignment and governance. Patricia Paskov is the strongest second choice because she can turn the same work into independent-audit and assurance methodology. Gabriel Weil is the strongest third choice for the applicant's stated law-school trajectory because he explicitly welcomes technically literate researchers who can reason about legal institutions without already having a JD.
The strongest alternative technical trio is Wei, Paskov, Casper. The strongest law-and-policy trio is Wei, Weil, Zittrain. The best AIxBio strategy is not to switch identities, but to apply the same evaluation-validity and assurance program to biological capability and biosecurity evaluations.
Name these three, in order: (1) Kevin Wei, (2) Patricia Paskov, (3) Stephen Casper.
Single strongest: Kevin Wei.
That ordering follows from the applicant’s unusually direct evidence of evaluation validity (deployment-vs-paper protocol gap, shortcut/attribution audit, and correlated monitor failure), his demonstrated ability to build reproducible offline tooling, and his explicit JD-to-AI-law/policy goal. Wei is the rare mentor whose published agenda joins all three: evaluation methodology, legal alignment, and technical governance. CBAI specifically names ecological validity, eval awareness, model-spec/constitution compliance, legal hierarchy, legal retrieval, deterrence, certification, and U.S.–China governance as his project space; it requires Python/ML/research experience for evals, which the applicant plainly has. His legal-alignment lane formally prefers a law degree in progress, so frame the applicant as an MS researcher deliberately building the empirical record before the JD—not as already qualified for a doctrinal-law project. CBAI—Wei RAND profile, updated Apr. 18 2026
Kevin Wei — best overall.
Patricia Paskov — best standards/auditing fit.
Stephen Casper — strongest technical-governance/bio bridge.
Nikola Jurkovic (METR): best choice if he wants a rigorous, code-heavy evaluation project and possibly a METR-adjacent research signal. His work is execution, threat modeling and forecasting—not primarily law. Emphasize reproducible experiments, human baselines, external validity, scanner failure analysis, and cyber-agent tasks. METR profile
Peter Slattery / Alexander Saeri / Michael Noetel (MIT AI Risk Initiative): strongest route to a policy-facing systematic review/database/taxonomy paper. The formal project asks for literature-review/qualitative-synthesis ability, AI governance familiarity, and says a PhD/equivalent is preferred. The applicant’s publication record, analytical audit and intended JD make him plausible, but his pitch must demonstrate careful coding protocol and synthesis—not just ML implementation. CBAI MIT AIRI project page
Gabriel Weil: strongest pure private-law/liability mentor. Good if the applicant wants to turn evidence from evaluation failures into a paper about standards of care, negligence, punitive damages, disclosure, or evidentiary use of evaluations. Lower than Wei/Paskov because this is a bigger pivot away from empirical evaluation and because no specific Fall project has been posted.
Peter Salib: excellent long-run intellectual match for a JD and AI governance, particularly law-as-alignment. Lower for this application because CBAI’s project information remains “More information soon” and there is no documented Fall-2026 task to target. CBAI roster
Sean McGregor / Charles Teague: strong alternatives for benchmark/incident infrastructure. McGregor’s publicly described CBAI direction is scaling the breadth and depth of incident databasing, and his recent work includes BenchRisk (NeurIPS 2025); Teague brings Inspect AI and scientific-publishing/tooling experience. They are especially attractive if the applicant frames STING as an auditable, offline evaluation/incident-analysis artifact. McGregor CBAI page Teague CBAI page
James Mickens / Samuel Gunn: credible cyber-security-adjacent technical safety alternatives. Mickens is particularly relevant to sandboxing and systems security; Gunn to watermarking/data attribution. They fit the applicant’s cyber profile but do not make the JD trajectory as central as Tier 1.
| Applicant finding | Closest current mentor agenda | Why |
|---|---|---|
| 3.1× apparent-performance inflation under the field-standard protocol | Kevin Wei (strongest); Paskov; Jurkovic; MIT AI Risk Initiative | Wei explicitly names ecological validity and evaluation methodology; Paskov works on proportional, credible evaluations and assurance; METR’s time-horizon work makes benchmark results interpretable in human task-time terms; MIT AIRI catalogues/evaluates risk mitigations. |
| Attribution audit finds no threat-actor-name shortcut | Benno Krojer/Bau/Orgad, then Wei | This is exactly a causal/shortcut-detection contribution. Krojer’s MVP benchmark exists because superficial cues can inflate evaluation scores. |
| One LLM monitoring another fails on the same inputs | Wei and Jurkovic, then Casper | This is a correlated-failure / judge-validity / oversight-robustness result. Wei’s science-of-evals work is the cleanest home; Jurkovic’s METR work includes failure analysis and scanner review; Casper works on audit and safeguard robustness. |
One unifying application thesis: “I study when AI-safety and cyber-AI evaluations look valid but fail under realistic conditions—and how to build causal, deployment-grounded evidence that can support standards, audits, and eventually law.” That thesis connects the cyber record, STING, validation engine, biology course background, and JD plan without pretending they are separate careers.
What “law-following AI” means. In the legal-alignment literature, it is not merely a model refusing illegal requests. It is a program of making AI systems comply with legitimate legal rules, use legal interpretation methods in reasoning, and use legal concepts as structural tools for reliability/trust/cooperation. The 2026 survey, coauthored by Wei, calls these the three research pathways of legal alignment. Legal Alignment for Safe and Ethical AI, arXiv:2601.04175, Jan./June 2026
Documented 2025–26 output: Position: Human Baselines in Model Evaluations Need Rigor and Transparency (ICML 2025); Methodological Challenges in Agentic Evaluations of AI Systems (ICML Technical AI Governance workshop, 2025); Infrastructure for AI Agents (TMLR 2025); Designing Incident Reporting Systems for Harms from General-Purpose AI (AAAI 2026); plus work on rigorous GPAI evaluations and RCT-style human-uplift studies. RAND
Model specs/constitutions. The public CBAI project menu explicitly asks how well models comply with model specs/AI constitutions, whether model and human interpretations differ, whether systems obey legal hierarchies, recognize legal implications, retrieve the right texts, and respond to deterrence-like penalties. This is project agenda, not necessarily a published Wei paper. The relevant adjacent empirical literature includes SpecEval (arXiv:2509.02464, Sept. 2025), which audits 16 models against provider behavior specifications and reports sizeable three-way specification/output/judge consistency gaps; Wei should not be credited as its author. CBAI—Wei SpecEval
Fit and prerequisite: perfect for evals; partial for legal alignment until JD begins. No documented public code screen. The published CBAI process says only that there is a mentor-specific task/screen after interview.
What AVERI is building. AVERI describes itself as building an independent third-party auditing layer for frontier AI. The January 2026 Frontier AI Auditing report defines this as third-party verification of developer safety/security claims and evaluation of their systems and practices against standards, with deep, secure access to non-public information. It proposes four AI Assurance Levels: AAL-1 as a present baseline and AAL-2 as the near-term objective for leading developers; higher levels require less reliance on company representations and more organization-wide scrutiny. AVERI report arXiv:2601.11699
“PCAOB/FINRA analogue” — inference, not an announced institution. A credible analogue would set auditor competence/independence rules, manage conflicts and cooling-off periods, specify access/security procedures for sensitive model/training/governance data, establish assurance-report formats, inspect or discipline audit providers, and make assurance levels intelligible to regulators and the public. That inference is grounded in AVERI’s explicit independence, deep-access, quality and anti-checkbox principles—not in a public claim that AVERI already is a PCAOB/FINRA-equivalent regulator. The AI Evaluator Forum is described as an emerging assessment-organization venue helping articulate access standards; it is not documented as a statutory regulator. AVERI legislative landscape, Apr. 20 2026
Fit: excellent. Biology is not presumed for her general auditing/standards work; statistics, measurement, writing, reproducibility and policy literacy matter more. There is no public evidence of a specific test task or an individual fellowship-to-AVERI hiring pipeline.
Documented CBAI project: systematic review of AI risk mitigations and systematic document review of organizational responses to AI risks. CBAI asks for strong literature review/qualitative synthesis and AI governance/policy familiarity; it says PhD/equivalent preferred. CBAI project page
Methodology and tooling. Their Mapping AI Risk Mitigations (arXiv:2512.11931, Dec. 12 2025) conducted a rapid evidence scan of 13 frameworks (2023–25), extracted 831 distinct mitigations, and created a four-category/23-subcategory draft taxonomy. It tested LLM assistance but found LLMs unreliable for fully automating extraction (confabulation, combining and omission); LLM classification suggestions were useful only with document-level manual comparison, author review, and multi-author consistency checks. The resulting data are public in an Airtable-backed database and interactive taxonomy. That is the correct sense of “LLM-plus-human-validated”—not an automatic taxonomy classifier. paper interactive taxonomy database
Taxonomy. The four top categories are Governance & Oversight, Technical & Security, Operational Process, and Transparency & Accountability controls; the 23 subcategories include risk management, model alignment/safety, testing/auditing, incident handling, disclosure and third-party assurance. Their related AI Risk Repository is a 1,725-risk meta-review and was published in Patterns in 2026. MIT FutureTech
Fit: strong if he reframes his evidence as a mini systematic-review/coding problem: classify evaluation failure modes and mitigations, preregister a codebook, quantify coder/LLM disagreement, and create a usable audit schema. No biology/cybersecurity prerequisite; limited direct evidence of junior-mentee outcomes or a full-time pipeline.
Current scholarship. The documented progression is: Tort Law as a Tool for Mitigating Catastrophic Risk from AI (SSRN 2024); Instrument Choice in AI Governance: Liability as the Indispensable Core (June 5 2025); Overcoming Judgment-Proofness: The Law & Economics of Insuring and Mitigating AI Risk (SSRN 2026); and Abnormally Dangerous Algorithms: The Case for Strict Liability at the AI Frontier (SSRN 2026). instrument-choice abstract strict-liability preprint insurance/judgment-proofness preprint
Punitive damages / uninsurable risk. Weil’s 2025 thesis is that strict, ex-post liability is comparatively calibrated to realized risk and incentivizes safety innovation; where compensatory damages cannot capture catastrophic stakes, punitive damages in compensable “near miss” cases associated with uninsurable risk could supply deterrence.
Counterargument—careful qualification. I did not locate a peer-reviewed, AI-specific published rebuttal squarely answering Weil’s punitive-damages proposal. The strongest documented general counterpoint is the mature audit/liability literature’s warning that assurance/liability systems have independence, expectations-gap, sensitive-information and box-ticking problems; a punitive regime additionally faces the classic under-deterrence problem if actors are judgment-proof and the pricing/information problem if insurers cannot observe or quantify tail risk. These objections are reasons to treat ex-post damages as complementary rather than sufficient, not evidence that Weil has been decisively refuted. Weil himself recognizes supportive, complementary and substitutionary non-liability policy. AVERI report Weil abstract
Fit: the applicant’s empirical work is valuable to Weil if positioned as evidence for standards of reasonable care, foreseeability, safety representations, or punitive-damages predicates. Law trajectory is an unequivocal asset. No specific Fall project, public screen, or hiring pipeline found.
Current agenda. Salib describes law as an alignment technology: rule-of-law systems already give powerful misaligned actors such as corporations/states incentives against harmful conduct; he is developing legal rights/duties frameworks for advanced AI that would similarly incentivize prosocial behavior. Forethought profile, Mar. 17 2026
AI rights. In AI Rights for Human Safety (with Simon Goldstein; public 2026 version), the argument is strategic/game-theoretic: property-status AIs and humans may face a destructive prisoner’s dilemma; rights to contract, hold property and bring tort claims—not merely negative “well-being” rights—could support repeated mutually beneficial exchange and make legal duties/penalties meaningful. This is a controversial theoretical proposal, not an established legal program. paper
CBAI project status: still unknown. As of the supplied date, the fellowship roster’s project material is marked “More information soon”; no Salib-specific Fall-2026 project page was found. Do not invent an implementation project from his biography. Law school is a genuine asset; cyber/biology are not assumed. CBAI roster
What the role means in practice. METR’s public materials show evaluation execution (running agents in Inspect, managing token budgets, scoring, transcript/failure review, cheating detection), evaluation development (task suites, human baselines, scaffolds, success thresholds), and risk assessment/threat modeling (independent review and pilot assessment of frontier developers). It is hands-on empirical evaluation, not generic “AI safety research.” METR profile
Time horizons. The METR long-task paper defines a model’s 50%-task-completion time horizon as the time a domain-knowledgeable human typically takes on tasks where the model succeeds half the time. It uses human baselines across RE-Bench, HCAST and short software tasks; it reports an approximately seven-month historical doubling time, with strong limits on generalization to messy real work. Nikola coauthored RE-Bench (ICML 2025), seven open-ended ML research-engineering environments with 71 eight-hour attempts by 61 human experts. time-horizon paper RE-Bench, ICML 2025
Nikola’s February 13 2026 note: he compared Claude Code/Codex to METR’s ReAct/Triframe scaffolds on newer time-horizon tasks, manually inspected failures, re-scored technical scoring defects, used LLM cheating scanners plus manual review, and found no statistically significant superiority for specialized scaffolds under tested conditions. That is extremely close in spirit to the applicant’s claim that a prevailing protocol can yield misleading numbers. METR note
Fit: outstanding technical match; less direct JD fit than Wei/Paskov/Weil/Salib. Cybersecurity is helpful, biology unnecessary. Publicly documented junior mentorship exists through AISST benchmarking/forecasting activity, but I found no verified public mentee-publication list or specific CBAI screen.
Recent work. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs (arXiv:2508.06601, Aug. 8 2025; BioSafeGenAI 2025 best-paper runner-up) trained 6.9B models with biorisk-related pretraining data filtered; it reports more than an order-of-magnitude improvement over post-training baselines through 10,000 adversarial fine-tuning steps/300M tokens, without observed unrelated-capability degradation, but shows retrieval-provided harmful knowledge still bypasses the core protection—hence defense in depth. paper project page
His current open-weight agenda also includes Open Technical Problems in Open-Weight AI Model Risk Management (arXiv:2608.07514/TMLR 2026), a 16-problem survey; plus auditing, incident, and legal-alignment work. The policy-loophole agenda is concrete: he highlighted August 10 2026 that no proposed/enacted U.S. frontier-AI law then imposed criminal penalties specifically for knowingly false public statements about imminent catastrophic risk; related work flags scope, continuous-evolution and information-asymmetry gaps for internal systems. open-problems page Aug. 10 post
Fit: excellent across cyber, biosecurity, evaluation validity and law. Biology helps for the Deep Ignorance thread but is not a prerequisite for incidents/law projects. His documented MATS/ERA/GovAI mentoring is the clearest public junior-mentoring signal among the top three; no public individual screen was found.
Use a single causal-evaluation-and-assurance narrative, not two narratives.
My strongest contribution is not simply building high-performing models; it is auditing whether performance claims survive the conditions in which people would rely on them. In cyber threat-intelligence retrieval, I found that the standard protocol overstated identical-model performance from 0.253 in deployment-like testing to 0.784, then used leave-one-out attribution to test and rule out a salient shortcut explanation. I also found that an LLM monitor can fail on the same inputs as the model it is supposed to validate. I want to translate this into rigorous safety-evaluation methods and, over the long term, standards and legal institutions that can distinguish credible safety claims from impressive but non-generalizable scores.
Final caveat: many individual Fall-2026 mentor project pages are still absent and the central fellowship page says “More information soon.” The tiering above relies on the published project pages where they exist and otherwise on current documented research—not an assumption that every mentor will offer the same project in October.
This applicant is a exceptionally strong candidate with a unique profile that bridges technical evaluation rigor, law/policy trajectory, and cybersecurity expertise. His three-part research finding—that field-standard MITRE ATT&CK mapping protocols inflate accuracy approximately threefold relative to deployment conditions, combined with attribution-based proof of no name-based shortcuts and discovery of correlated monitor failure—directly maps onto the current agendas of multiple Tier 1 mentors.
Tier 1 mentors (name these three as top choices, in this order):
Single strongest mentor: Kevin Wei — His published research on human-baseline rigor in evaluations, legal alignment evaluation methodology, and science-of-evals frameworks directly validates applicant's core finding (eval inflation) and provides a ready-made research home for applicant's work on measurement validity under real deployment conditions.
Law school trajectory assessment: For this applicant, law school is a significant competitive asset, not a distraction. Kevin Wei, Gabriel Weil, Peter Salib, Jonathan Zittrain, and Michael Chen all explicitly integrate law, policy, and governance into their research. These mentors will view the applicant's law-school trajectory as deepening rather than abandoning technical work.
Two-track strategy: The applicant should apply to both AI Safety and AIxBio tracks on a single integrated application, positioning his evaluation-validity work as applicable to both safety oversight (general frontier AI) and biosecurity evaluation (validating defensive countermeasures). This avoids looking unfocused while leveraging his dual background (medical coursework + cybersecurity expertise + AI evaluation).
Affiliation: Research Scholar, GovAI (Oxford Martin AIGI); Harvard JD; formerly UK AISI Science of Evaluations team
Current research (2025-2026): Wei has published four major papers directly relevant to the applicant's work:
Why he's the strongest match for this applicant:
Law school trajectory: MAJOR ASSET. Wei's entire research agenda is "how do legal frameworks constrain AI behavior, and what does that mean for evaluation methodology?" For an applicant planning law school, this is the ideal mentor—Wei will see the JD as enabling better technical governance research, not detracting from it.
Mentoring specifics: Wei is listed as a CBAI mentor for Fall 2026. No public code screen documented, but given his RCT/human-baseline methodology focus, applicant should expect questions on evaluation design, validity threats, and how to measure true vs. inflated performance.
Background requirements: Cybersecurity background is compatible (not required). Medical background is plus (helps applicant understand domain-specific evaluation standards). No explicit prerequisite beyond MS-level AI knowledge.
Affiliation: Director of Standards, AVERI (AI Verification and Evaluation Research Institute); Adjunct Researcher, RAND; Oxford DPhil candidate (multi-agent security)
Recent work (2025-2026):
Why she's a strong second choice:
Law school trajectory: ASSET. AVERI's focus on "third-party auditing" and standards-setting has direct governance and legal implications (liability frameworks, regulatory compliance). Paskov's work bridges technical evaluation rigor and legal/institutional design—ideal for an applicant planning law school.
Mentoring: RAND researcher with known mentoring experience. Applied to CBAI as mentor. No public code screen, but given standards/auditing focus, expect questions on evaluation methodology, audit design, and institutional scalability of rigor.
Background requirements: The applicant's cybersecurity background (MITRE ATT&CK mapping, threat intelligence) is directly applicable—auditing frameworks must work across threat domains. Medical background not prerequisite.
Affiliation: Assistant Professor of Public Policy, Harvard Kennedy School; Faculty Affiliate, Harvard School of Engineering and Applied Sciences; MATS TAIGR mentor
Current research (2025-2026):
Why he's Tier 1:
Law school trajectory: ASSET. TAIGR research on "technical loopholes in frontier AI laws" shows Casper is explicitly thinking about how technical insights inform legal/policy gaps. For an applicant planning law school, this is valuable—Casper will see the JD as tools for translating technical findings into policy.
Mentoring: Established mentor at MATS; note on CBAI: "Experience in conducting, writing, and presenting academic research; by default 2-3 meetings/week." Expects high research productivity.
Test tasks: Likely code-heavy or methodology-focused (TamperBench is a large-scale evaluation suite). Applicant should prepare to discuss evaluation design and how to measure tamper resistance rigorously.
Background requirements: Cybersecurity background is plus (tampering is an adversarial angle on cyber-resilience). Medical background not prerequisite.
Affiliation: Assistant Professor of Law, University of Houston Law Center; Non-Resident Senior Fellow, Institute for Law & AI; Founder/Exec Co-Director, Center for Law & AI Risk; Visiting Senior Fellow, Institute for Law & AI; Contributing Editor, Lawfare
Recent work (2025-2026):
Why he's Tier 1B (not 1A, but critical secondary):
Law school trajectory: CRITICAL ASSET. This is Weil's core research agenda. For an applicant saying "I want to go to law school then work on AI law/policy," Weil will see that as the applicant's strongest thesis, not a distraction.
Mentoring: No explicit CBAI mentor listing found, but Weil is co-listed on many multi-author governance papers. May not be formally available for Fall 2026 (check CBAI site for updated mentor list).
Affiliation: AI Science Advisor, California Governor's Office of Emergency Services (Cal OES); formerly METR
Recent work (2025-2026):
Why he's Tier 1B:
Law school trajectory: ASSET. Government service + policy expertise show clear pipeline for applicant.
Mentoring: MATS mentor (confirmed: Michael Chen listed as MATS Mentor for Summer 2026) [44]. Likely available for CBAI Fall 2026.
Affiliation: Member of Technical Staff, METR
Current focus: Eval development, eval execution, risk assessment [20]
Why: Applicant's finding that field protocols inflate accuracy threefold is an execution-level problem (how evals are run, not what's being measured). Jurkovic's METR work on "time-horizon 1.1" benchmarking and systematic evaluation scaling shows deep expertise in eval rigor. Applicant's correlated-monitor-failure finding is a risk assessment data point Jurkovic would likely incorporate into METR's methodology.
No major publications 2025-2026 found publicly, but METR's monthly frontier-risk reports (Feb-Mar 2026, May 2026) show systematic evaluation methodology that would be valuable collaboration ground.
Law school trajectory: Neutral. Pure evaluations technical focus with no explicit policy angle documented.
Affiliation: Goodfire (Dan Braun and Oliver Clive-Griffin, along with Lee Sharkey)
Current work:
Why: Applicant's attribution-based audit (proving no threat-actor-name shortcut in MITRE ATT&CK mapping) uses similar mechanistic interpretability logic to Braun's parameter decomposition. Both are about surgical understanding of why a model makes a decision. For a mentorship, Braun could help applicant scale attribution methods to larger models and more complex behaviors.
Test task risk: May have code screening (Goodfire is engineering-focused). Applicant should be prepared to discuss mechanistic interpretability and attribution computation.
Affiliation: Northeastern University (Bau Lab); Benno Krojer also active in science communication
Current work:
Why: Applicant's shortcut-detection audit is mechanistic interpretability work. Bau's causal intervention framework and circuits methodology are directly applicable. Krojer's science communication focus is a plus if applicant plans to communicate policy implications to non-technical audiences.
Science communication note: Krojer explicitly described as "active in science communication" — if applicant plans to communicate eval-validity findings to policymakers, Krojer could mentor the communication layer.
Law school trajectory: Neutral to slightly positive (Krojer's science communication + policy relevance). Bau is purely technical.
CBAI Note: Bau and Krojer are listed as a mentoring pair [25], so they come together.
Affiliation: Kempner Institute for Neuroscience, Harvard University
Recent work (2025-2026):
Why: Applicant's correlated-monitor-failure finding (one LLM fails to catch another's errors) is related to deception/sycophancy—models may fail because they're aligned with each other's errors, not independent. Orgad's work on emergent misalignment in dishonesty (paper: "LLMs Deceive Unintentionally," arXiv:2510.08211v2, January 2026) [31] is directly relevant. Applicant's work could contribute a novel angle: if safety mechanisms are correlated, systematic deception is possible.
Law school trajectory: Neutral. Pure interpretability/technical focus.
Affiliation: Assistant Professor, MIT CSAIL; Algorithmic Alignment Group
Recent work (2025-2026):
Why: Applicant's correlated-monitor-failure finding parallels Hadfield-Menell's concern about evaluation brittleness. If alignment evaluations are unreliable (driven by randomness, not learned representations), then monitor-based safety (one model monitoring another) will fail exactly as applicant found. This is a mentor who would see applicant's work as validation of his theoretical concerns.
Law school trajectory: Neutral. Algorithmic focus, no explicit policy angle.
Affiliation: Researcher/CSO, RESI (formerly OpenAI)
Recent work (2025-2026):
Why: If applicant's monitor-failure result shows systematic hallucination cascade (one model's errors propagate to another's), Kalai's hallucination theory is the explanatory framework. Applicant's empirical finding + Kalai's theory = publishable collaboration.
Law school trajectory: Neutral. Theoretical focus.
Affiliation: Co-founder AVERI; Founder, AI Incident Database; Berkman Klein Center
Current work:
Why: Applicant's findings (eval inflation, shortcut audits, correlated failures) are incident types that should be documented in AIID taxonomy. If applicant wants to influence how industry tracks evaluation failures as safety incidents, McGregor's database is the infrastructure. Tier 2B because this is infrastructure work, not primary research agenda, but valuable for policy impact.
Law school trajectory: Neutral. Incident tracking focus without explicit legal angle (though data could inform liability frameworks).
Affiliation: Harvard Law School; George Bemis Professor of International Law; Co-founder/Director, Berkman Klein Center for Internet & Society
Recent work (2025-2026):
Why: Zittrain is the most senior law-school-affiliated mentor. If applicant is serious about HLS or similar, Zittrain represents the gold standard for law-school-based AI governance research. His co-authorship on "Legal Alignment" shows he's actively engaged with technical safety questions. Not Tier 1 for this applicant because there's less direct research fit (Zittrain's focus is internet law + governance, not evaluation rigor), but strategically important for a law-bound applicant.
Law school trajectory: MAJOR ASSET. Zittrain IS the law school angle—his entire career shows that serious AI governance happens with legal expertise, not despite it.
Note: May not mentor Fall 2026 (co-director of Berkman Klein but no explicit mentor listing found).
Affiliation: Professor of Law, University of Houston Law Center; Founder/Exec Co-Director, Center for Law & AI Risk; CAIS Law & Policy Advisor
Recent work (2025-2026):
Why: Salib is the most philosophically ambitious of the law mentors—his "AI rights" framework is novel and controversial. For an applicant interested in non-standard legal approaches to AI governance, Salib offers intellectual firepower. However, his specific research is less directly mapped to applicant's eval-validity work compared to Wei or Weil.
Law school trajectory: ASSET. AI rights as alignment is exactly the kind of legal-philosophy-meets-technical-safety thinking applicant needs if heading to law school.
Note: CBAI mentor page for Salib still shows "More information soon" (as of July 2026) [19], so actual Fall 2026 project scope is TBD. This is a risk.
Shi Feng (Praxis Research / GWU) — Model organisms of deception, intent misalignment; relevant to applicant's correlated-failure finding
Zach Furman (Iliad) — Singular Learning Theory, training dynamics; relevant to understanding why evaluations inflate (SLT explains phase transitions in learning)
Adam Shai & Paul Riechers (Simplex / Astera) — Belief-state geometry, computational mechanics; relevant to interpretability of model disagreements
James Mickels (Harvard) — Systems-level sandboxing, alignment red-teaming; applicant's correlated-failure result is a red-team finding (if monitors fail together, sandbox is brittle)
Samuel Gunn (RESI) — Watermarking, data attribution; applicant's attribution-based audit is a data-attribution study
Charles Teague (Meridian Labs) — AI evaluation methodology (helped develop Inspect AI); RAND biological knowledge benchmarking for frontier LLMs
Hidenori Tanaka (Harvard Physics of Intelligence) — Persona mechanics, multi-agent collective intelligence; correlated-failure could be a multi-agent problem
Jay Chooi (Robocurve) — Robotics benchmarks; less directly relevant unless applicant interested in embodied AI safety
1. Kevin Wei — Position as: "His recent work on human-baseline rigor in evaluations directly validates my core finding: that field-standard MITRE ATT&CK mapping protocols inflate measured accuracy roughly threefold relative to deployment conditions. I discovered this by building an attribution-based audit (STING) that proves the model does not rely on threat-actor-name shortcuts, yet accuracy drops from 0.784 to 0.253 on identical model and queries. This suggests the field-standard evaluation protocol measures artifacts, not genuine capability. Wei's RCT methodology and his ICML 2025 call for 'transparency in human baselines' are exactly the frameworks needed to fix this. I want to work on scaling rigorous evaluation methodology across domains—from cybersecurity threat classification to frontier AI safety."
2. Patricia Paskov — Position as: "AVERI's frontier AI auditing framework provides the institutional scaffold my evaluation-validity work needs. If deployment conditions reveal threefold accuracy inflation, auditors must detect this. Additionally, my finding that using one LLM to monitor another fails on the same inputs the generator fails on (correlated monitor failure) suggests that evaluation rigor is foundational to auditing. Paskov's work on methodological foundations and her DPhil research on multi-agent security align with these concerns. I want to work on audit design that actually detects brittle evaluations before they inform high-stakes AI deployment decisions."
3. Stephen Casper — Position as: "TamperBench systematically evaluates how open-weight LLMs resist internal weight/activation manipulations. My discovery of correlated monitor failure is a tampering vulnerability: if one safety mechanism fails, others likely fail on overlapping inputs. I want to contribute empirical findings on why tampering vulnerabilities are correlated and how evaluation design can detect this fragility before deployment. Casper's work on 'open technical problems in risk management' is where my empirical findings belong."
How to present it strategically:
"After my MS, I'm planning to attend law school, then work at the intersection of AI, law, and policy. This is not a distraction from technical AI safety—it's the completion of the technical work. My evaluation-validity findings only matter if they inform how companies and regulators decide to deploy AI. That decision-making happens in legal and governance structures. Kevin Wei's legal alignment framework and Gabriel Weil's liability scholarship show that the best AI safety work requires both: technical rigor to prove the problem (my eval findings), and legal infrastructure to ensure decision-makers act on that proof (Weil's liability mechanisms). I'm seeking mentors who integrate both."
Mentors likely to see this positively:
Mentors likely neutral:
| Applicant's Finding | Conceptual Name | Primary Mentor(s) | Research Connection | Publication Venue |
|---|---|---|---|---|
| 0.784 → 0.253 accuracy drop on deployment | Evaluation Inflation | Kevin Wei, Patricia Paskov, MIT AI Risk Initiative trio | Wei's "Human Baselines Need Rigor" (ICML 2025); Paskov's methodology for calibrating eval confidence | ICML, AAAI, or TMLR |
| Attribution audit proves no threat-actor-name leakage | Mechanistic Eval Audit | Dan Braun, David Bau, Hadas Orgad | Braun's parameter decomposition; Bau's mechanistic interpretability; Orgad's actionable interpretability | Mechanistic Interpretability workshop, ICLM, or COLM |
| One LLM fails to catch another's errors on shared inputs | Correlated Monitor Failure | Stephen Casper, Dylan Hadfield-Menell, James Mickels | Casper's TamperBench framework for vulnerability correlation; Hadfield-Menell on randomness in alignment; Mickels on red-team fragility | ICLR, NeurIPS Safety Workshop, or Casper's "Open Technical Problems" venue |
Problem: Applicant is qualified for both tracks. How to apply without looking unfocused?
Solution: Frame it as domain-transfer of evaluation methodology.
In application narrative: "I'm applying to both AI Safety and AIxBiosecurity tracks because rigorous evaluation is the bottleneck in both. In AI Safety, my work proves that field-standard eval protocols inflate measured safety (threat classification tests become proxies for threat-name memorization, not genuine threat understanding). In AIxBiosecurity, the same problem applies: we need to know whether defensive AI countermeasures actually work under real-world biosecurity conditions, not just in lab benchmarks.
My background straddles both: three years of medical school (virology, immunology, pharmacology, pathology) + three IEEE ICAIC publications on cyber threat classification + BlueDot Biosecurity Fundamentals training. My evaluation-validity methodology (attribution audits, correlated-failure detection) applies across domains.
For the AI Safety track, I'm targeting Kevin Wei, Patricia Paskov, and Stephen Casper to work on generalizing evaluation rigor across frontier AI safety domains.
For AIxBiosecurity, I'm seeking mentors working on biosecurity-specific evals [list names once identified—likely includes RAND AI-Bio Evals team and others]. The evaluation methodology is the same; the domain is different."
Why this works: It's not "I'm interested in everything." It's "I've identified the core bottleneck (eval methodology), and I'm applying that focus across two safety-critical domains where evaluations are weaker than they should be."
Mentors with documented junior-researcher mentoring:
| Mentor | MATS | SPAR | PIBBSS | ERA | CBAI Previous | Evidence |
|---|---|---|---|---|---|---|
| Kevin Wei | ? | ? | ? | ? | ? | GovAI junior researchers; UK AISI mentees [4] [5] |
| Patricia Paskov | ? | ? | ? | ? | ? | RAND research mentorship pipeline |
| Stephen Casper | ✓ (TAIGR) | ? | ? | ✓ | ? | MATS TAIGR mentor; UK AISI residency [41] |
| Gabriel Weil | ? | ? | ? | ? | ? | Law school faculty (typical mentoring); policy-adjacent papers with co-authors |
| Michael Chen | ✓ (Summer 2026) | ? | ? | ? | ? | MATS Mentor listing [44] |
| David Bau | ? | ? | ? | ? | ? | Northeastern faculty; NEMI workshop organization suggests junior-researcher engagement |
| Dylan Hadfield-Menell | ? | ? | ? | ? | ? | MIT CSAIL faculty; multiple co-authored papers with junior researchers |
| James Mickels | ? | ? | ? | ? | ? | Harvard faculty; Berkman Klein affiliated |
Note: CBAI site does not archive past cohort mentors or mentee outcomes, limiting ability to verify historical mentoring success. Applicant should ask CBAI directly for references from previous fellows mentored by each candidate.
Publicly documented requirements:
Kevin Wei — No published screen, but based on RCT/human-baseline methodology focus, expect: (a) evaluation design problem (design an RCT to measure human baseline vs AI performance), (b) case study in evaluation inflation, or (c) writing task on evaluation rigor trade-offs.
Patricia Paskov — No published screen, but based on standards/auditing focus, expect: (a) audit design (how would you design a frontier AI audit?), (b) taxonomy/methodology problem, or (c) institutional scalability discussion.
Stephen Casper — No published screen, but based on TamperBench work and TAIGR leadership: likely includes code component. Possible: implement a simple tamper-resistance evaluation, or write a problem-finding paper on regulatory loopholes. Expectation: "2-3 meetings/week" research commitment [42].
David Bau / Benno Krojer — Pair requires: "Prior experience in interpretability, VLMs, cognitive science; high standards for scientific communication/rigor" [25].
Note: CBAI site indicates mentor-specific tasks exist but are not published pre-application. Applicant should prepare across all three types: (a) written case studies, (b) code/methodology problems, (c) presentation/communication tasks.
Explicitly stated requirements found:
Applicant's fit against prerequisites:
| Mentor | Stated Requirement | Applicant Has? | Gap? |
|---|---|---|---|
| Kevin Wei | Legal eval methodology knowledge | Partial (cybersecurity domain, not legal; but attribution/audit methods apply) | No gap; transferable |
| Patricia Paskov | Auditing/standards knowledge | Yes (audit design for MITRE ATT&CK; STING pipeline is audit tool) | No gap |
| Stephen Casper | Academic research skills | Yes (IEEE publications, first-author papers, under review at ACSAC, journal article in prep) | No gap; high fit |
| David Bau/Krojer | Interpretability + VLM experience | Partial (attribution-based audit is interpretability; LLM validation engine exists; VLM experience not documented) | Minor: applicant should emphasize LLM output validation work |
| Dan Braun | Parameter decomposition knowledge | No | Learnable on-the-job |
| Hadas Orgad | Actionable interpretability | Partial (applicant's audit findings are actionable—fix threat-classif eval) | No significant gap |
| Dylan Hadfield-Menell | Preference learning background | Unlikely | Not blocking; theoretical foundation learnable |
Conclusion: No mentor has stated prerequisites applicant clearly cannot meet. Weakest area: pure interpretability/mechanistic background (VLMs) for Bau/Krojer pair, but applicant's LLM validation engine work compensates.
What could NOT be verified (no public sources found):
AIxBiosecurity Fall 2026 mentor names. Query mentions "Active Site, SecureBio, and the RAND AI-Bio Evals team" but specific mentor names/profiles not documented. Applicant should ask CBAI directly. Confidence on remaining recommendations: not affected.
Stephen Casper's BioSafeGenAI best-paper-runner-up title. Award confirmed in Casper's CV (source [43]) but specific paper title not located. Suggests Casper has done biosecurity-adjacent safety work. Confidence: medium. Applicant can ask Casper directly.
Peter Salib's Cambridge UP book details. Announced in applicant query but CBAI page notes "More information soon" (as of July 2026). Book likely not yet published. Confidence impact: low. Salib's existing papers sufficient to assess fit; book details not critical.
Oliver Clive-Griffin (Goodfire). Minimal public information. Co-mentioned with Dan Braun on parameter decomposition but no individual publications or mentoring track record found. Confidence: low. Applicant should treat as supporting mentor at Goodfire alongside Braun; not enough info for independent ranking.
Charles Teague specifics. RAND biological benchmarking documented; Meridian Labs CEO + Inspect AI development mentioned but limited public detail on Inspect AI. Confidence: medium. Likely strong technical mentor but less research publication trail than others.
Jay Chooi (Robocurve). Robotics benchmarks mentioned but minimal documentation. Confidence: low. Skip unless applicant has robotics background.
Hidenori Tanaka's 2025-2026 publications. General "Physics of Intelligence" program visible; specific recent publications on persona mechanics not located in arXiv/venue search. Confidence: low on recent output. General program direction clear but specific Fall 2026 project TBD.
Angelo State University MS AI program specifics. Program exists (AI Center of Excellence, July 2026 announcement) but curriculum details not publicly documented. Confidence impact: low. Applicant's background is plausible; verification would require direct contact with Angelo State.
Test tasks/code screens for mentors. CBAI page indicates they are "mentor-specific" but not published. Confidence impact: medium. Applicant should prepare generically (eval design, code, writing) and ask mentors directly pre-interview.
Specific technical finding, not generic "studied evaluations." Applicant discovered a quantified, reproducible phenomenon (threefold accuracy inflation) with plausible mechanistic explanation (protocol measures memorization, not understanding). This is publication-quality work.
Methodological sophistication. Attribution-based audit (STING), leave-one-out token analysis, correlated-failure detection—these are advanced interpretability techniques applied to a novel problem. Mentors will see this as someone who thinks mechanistically.
Cross-domain credibility. Few applicants have credible cybersecurity + medical + AI backgrounds. This positions applicant for emerging domains (AI in biosecurity, AI in critical infrastructure, AI in threat modeling) where technical + domain expertise is rare.
Law + policy + tech integration. Almost every strong AI safety researcher has credentials in 2 of 3 (tech + law, or tech + policy, or law + policy). Applicant is signaling intent to develop all three before age 25. Mentors who care about governance will see this as unusual and valuable.
Government pipeline visible. DoD/Army RA + GovAI DC Winter Fellowship show applicant is already on the policy track, not pivoting last-minute. This signals seriousness about career trajectory.
Communication skills evidenced. Three IEEE ICAIC publications (first author on two) + Under review at ACSAC + Journal article in prep = ~4 papers on the trajectory. Few MS students have this output. Mentors know applicant can write.
Name in order:
Mention secondarily:
Honest positioning: "I'm a cybersecurity + medical school + AI student who discovered that field-standard evaluation protocols are three times less accurate than claimed. I've built audits to understand why, and I've found that safety monitors fail together, not independently. I'm planning law school to learn how to translate technical findings into policy infrastructure. I want mentors who integrate technical rigor, evaluation methodology, and governance thinking."
Most-cited sources for this analysis:
Want this comparison for your own question? Run a blind battle between deep research AIs or see the deep research API leaderboard from all community votes.