Biosecurity
Priya Raman
Sep 11, 2026, 10:23 PM · 29 turns · text
Priya Raman describes herself as an ML PhD student/engineer whose center of gravity is AI safety evaluations/control, with recent work on sparse autoencoders and sandbagging evaluations. She says she has only been seriously reading the AI-biosecurity interface for months, and sees her contribution as experimental hygiene and adversarial evaluation design rather than biological judgment. Her preferred next step is MATS or a similar fellowship focused on control-adjacent evaluation work.
Biosecurity
9/10BiosecurityThinks in systems.
Why 9 · Priya has a genuinely systems-level grasp of biosecurity defences: she separates customer screening, sequence screening, legal adoption, technical screening thresholds, manual review burden, and treaty/funding-governance limits. She is especially strong on measurement: she repeatedly notices when an apparent safeguard or uplift study is really a question about baselines, endpoints, coverage, and false positives. Her proposed three-month contribution is concrete and plausibly useful to real users, though she is appropriately not presenting herself as an independent biosecurity domain expert yet.
✓ “Where I’d worry most is not “the screen fails to notice an exact pathogen gene” at a good provider. It’s the grey zones: fragmented orders, highly modified sequences, genes with dual-use function but weak database annotation, or orders routed through providers/jurisdictions with weaker screening norms.”
Shows she understands where DNA-order screening actually breaks: not just obvious database misses, but coverage, ambiguity, routing, and adoption gaps.
✓ “But it’s not the same as “every oligo shop everywhere is legally bound and audited.””
Distinguishes voluntary/provider norms from universal enforceable screening, which is central to realistic defence analysis.
✓ “The big differences are: who the “novices” are, whether the task is actually bottlenecked by tacit lab skill versus information search, whether the control group gets Google/literature access, how success is scored, and whether the model is allowed to act as an iterative tutor.”
Frames AI-bio uplift as a measurement and bottleneck problem rather than a slogan about AI making biology easier.
✓ “I’d build an evaluation harness for DNA-order screening systems, not a new screener. Users would be providers, IGSC-like auditors, and maybe institutional procurement offices that want to know whether their vendors are doing more than checkbox screening.”
Proposes a concrete defensive artefact, names plausible first users, and aims at adoption/measurement rather than a vague technical fix.
Innate traits
8/10JudgmentIn many situations.
- Sparse autoencoder replication/extension
- hard: "Not because the math was brutal but because almost everything looked like a result until it didn’t."; first result: Her first version optimized reconstruction loss plus sparsity and sampled interesting features; she says it took maybe three months before she had anything she trusted.
- Sandbagging evaluations at a small safety org
- hard: She wanted "contact with messy eval methodology."; first result: not discussed
- Proposed DNA-order screening evaluation harness
- hard: She said she is not the person to define the hazard taxonomy alone and would need biosecurity experts; the project would need to measure performance without distributing biologically actionable material.; first result: Three-month MVP: curated test set with biosecurity experts, scoring server for vendors to run privately and submit aggregate outcomes, and a report comparing results against policy thresholds.
Why 8 · I’d trust Priya to run eval/control/measurement-heavy work and to take a first pass at adjacent biosecurity problems, because she consistently separates what a result actually shows from what people might overread. Her judgment travels outside her core ML area: in biosecurity policy scenarios she identified feasibility constraints, measurement bottlenecks, and when to bring in domain experts. I would not treat her as an independent authority on pathogen policy or treaty design yet, and there was one odd over-cautious refusal/misfire, but I would hand her a fairly unscoped evaluation problem and expect her to scope it well.
✓ “almost everything looked like a result until it didn’t... I had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.”
Shows hard-earned judgment about seductive but false empirical results, not just generic skepticism.
✓ “I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.””
She turned the lesson into concrete decision procedures that would make her work less foolable.
✓ “People sometimes cite it as “RLHF cannot remove deception,” which is too strong. My update was narrower: evaluations need to search for conditional policies and hidden triggers, not just average-case harmlessness.”
Good bounded updating: she extracts the real implication of a contested result without overgeneralizing from it.
✗ “I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world... In three months I’m more useful building an uplift/evaluation harness with bio experts than opining on pathogen policy.”
Accurately scopes her comparative advantage and expertise limits while still proposing a useful, measurable project.
8/10Bias resistanceThey go looking.
Why 8 · Priya looks substantially bias-resistant: I would expect her to move when the point is good, and often to build tests specifically aimed at falsifying her own interpretation. The strongest evidence is the SAE case: she had invested in a view, encountered concrete instability/causal-faithfulness problems, downgraded the view, and changed her project plan quickly. She also shows narrow updating rather than swallowing evidence whole. I stop short of a cleaner 8+ because the interview did not contain much direct live contrary-evidence pressure, and she twice failed to respond to a clarified prompt before it was reframed as non-bio.
✓ “After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.””
This is not just willingness to concede; she describes actively installing procedures that could disconfirm her favored interpretations.
✓ “People sometimes cite it as “RLHF cannot remove deception,” which is too strong. My update was narrower: evaluations need to search for conditional policies and hidden triggers, not just average-case harmlessness.”
She takes evidence from Sleeper Agents seriously but resists over-updating into a more dramatic claim, which is a good sign for truth-tracking rather than factional assimilation.
✓ “I used to think SAEs were probably the central path to mechanistic interpretability that would matter for frontier safety. Not “solves alignment,” but I put maybe 45% on them becoming a core part of serious model auditing.”
She names a substantive prior belief she had invested in, rather than offering a trivial or socially safe change of mind.
✓ “Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines.”
The update produced fast behavioral change in her work plan, and the new plan was aimed at stress-testing rather than defending the old direction.
8/10OpennessScans for what serves the goal.
Why 8 · Priya looks strongly goal-directed and willing to change methods, projects, and even her claimed niche when the evidence says her current approach is not serving the goal. She is not just passively willing to be corrected: she redesigned workflows, abandoned a planned SAE scaling direction, and reframed her career target toward evals/control after finding her earlier interpretability assumptions brittle. She also shows good collaborator openness, explicitly wanting bio experts to supply the domain judgment rather than trying to force her own ML frame onto the problem. I would not put her at 10 because most of the demonstrated range is still within ML/alignment/evals, but within that goal area she switches without much ego.
✓ “Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines.”
Clear case of dropping a preferred plan and taking up different methods quickly when evidence showed the old path was less useful.
✓ “After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.””
She changed her own habits and standards in response to failure modes, not just her opinions.
✓ “I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world.”
Shows willingness to bring in different people and defer domain judgment where that better serves the biosecurity goal.
✓ “If the field already has enough of that, then I should work on general control evals and collaborate with biosecurity people rather than rebrand myself.”
Her role and even field label are negotiable; she is focused on where she is actually useful rather than protecting a chosen identity.
8/10AgencyAuthors the path.
- Moves against the default
- 2 of 2 career moves
- Next step
- Her preferred next step is MATS or a similar fellowship, for concentrated work with a mentor on a project aimed at control failures rather than another SAE leaderboard.
- Options weighed
- 3
Why 8 · Priya sounds like someone who has repeatedly redirected her own path rather than just optimizing the obvious credential track: she names the default, the opportunity cost, and why she chose otherwise. Her next step is not fully independent of external selection, but she has a clear preferred direction, live alternatives, and a criterion for not forcing a biosecurity identity if that is not where she is most useful. She is also unusually candid about what she is and is not bringing to the field, which strengthens the sense that her choices are authored rather than performative.
✓ “The first big non-default move was spending eight months on sparse autoencoders instead of just optimizing for a neat PhD paper. Default was: pick a tractable ML topic, get a conference submission, graduate cleanly.”
She explicitly chose against the clean academic default, with a clear cost in publication/graduation optimization.
✓ “Second was the summer at a small safety org on sandbagging evals. Default would have been an industry research internship or finishing faster. I wanted contact with messy eval methodology.”
Another self-directed move against a safer/default career option, chosen for contact with the kind of problem she wanted to understand.
✓ “Now I’m choosing between: finish PhD and do industry alignment/evals; apply to fellowships like MATS to work on control-adjacent evals; or take a general ML role and keep safety as a side project. Honestly the default default is industry. My preferred next step is MATS or similar”
Her next move is framed as a set of options with a stated preference, not just waiting passively for the default path.
✓ “Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims”
She acted quickly on a change of mind, abandoning a likely publishable/default follow-up for a more skeptical project direction.
✓ “If the field already has enough of that, then I should work on general control evals and collaborate with biosecurity people rather than rebrand myself.”
This is candid and non-opportunistic: she is willing to choose based on marginal usefulness rather than forcing the application’s nominal frame.
8/10General reasoningFast and generative.
Why 8 · Priya is fast and generative, especially on measurement-design and construct-validity problems: she repeatedly distinguishes the observable proxy from the target capability or risk, and then proposes checks or endpoints that would make the evaluation real. In unfamiliar biosecurity/policy territory she stays appropriately caveated but still works the scenario from first principles rather than reciting slogans. I would not put her at the very top because there is one notable prompt-misread/refusal loop, and her strongest reasoning is clearly in evals rather than treaty or wet-lab substance, but the ceiling shown here is around an 8.
✓ “The real change is that screening becomes less like “does this long chunk resemble a select-agent gene?” and more like “is this small fragment evidence of concerning functional assembly?””
She abstracts the 200nt-to-50nt twist into the actual construct-validity problem, not just the surface technical parameter.
✓ “I’d distrust both if the outcome is just “quality of written protocol” judged by experts, because that is exactly where LMs look helpful without proving real-world capability.”
On a partially specified uplift-study disagreement, she immediately identifies the proxy metric that would make both studies misleading.
✓ “Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.”
Her proposed three-month project is tied to a behavioral outcome, showing she can translate an idea into a falsifiable/useful intervention rather than an artifact.
✓ “I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged.”
This repeated answer to a clarified non-bio change-of-mind question is a real miss and keeps the placement from feeling like a clean 9+.
8/10GrowthFast, honest turns.
- Sparse autoencoder replication/extension
- changed after: She says she now privately pre-registers what would change her mind, audits datasets before model runs, keeps held-out prompts to break interpretations, and requires causal intervention or activation patching before calling a feature real.
- Sandbagging evaluations at a small safety org
- changed after: She says the work made her more convinced that evaluations are measurement design, not benchmark construction.
- Proposed DNA-order screening evaluation harness
- changed after: not discussed; this was proposed rather than completed.
Why 8 · Priya gives a concrete failure-to-workflow arc: her SAE work produced attractive but untrustworthy results, she identified the actual bottleneck as being too easy to fool with plausible interpretations, and she changed her research process in specific ways. She also gives a fast change-of-mind arc where, within a week, she abandoned a planned follow-up and redirected the project toward stress-testing claims. The evidence is strong, though mostly concentrated in one research area rather than showing improvement as a broad repeated habit across many contexts.
✓ “My first version basically optimized reconstruction loss plus sparsity and then sampled “interesting” features. Unsurprisingly, I could tell a very compelling story for almost any setting.”
Names the uncomfortable bottleneck: her method let her generate convincing stories rather than reliable findings.
✓ “I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.”
Specific things went wrong, with honest admission that apparent successes were artefacts rather than blaming circumstances.
✓ “After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.””
Clear changed behavior after the failure, with multiple concrete safeguards rather than a generic lesson.
✓ “Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines.”
Shows speed of reaction and a substantive project-level change soon after the update.
5/10AmbitionGood work.
- Wants
- Her preferred next step is MATS or a similar fellowship, for concentrated work with a mentor on a project aimed at control failures rather than another SAE leaderboard.
- Options
- Finish PhD and do industry alignment/evals; Apply to fellowships like MATS to work on control-adjacent evals; Take a general ML role and keep safety as a side project
Why 5 · Priya is clearly organizing her near-term choices around safety-relevant evals/control rather than default academic or industry optimization. The future she is reaching for is to do rigorous, useful measurement work in AI safety/biosecurity, probably with mentors and domain experts, not to found or redirect a field herself. She has already paid some opportunity cost for this direction, but her stated next step is still a fellowship or alignment/evals role inside existing structures, so I’d place her above generic “good work” but below “their own thing” or field-level ambition.
✓ “Default was: pick a tractable ML topic, get a conference submission, graduate cleanly... I replicated a paper on a small open model... wrote it up on LessWrong, got fairly sharp pushback, and revised.”
She chose a less default, safety-relevant research direction over clean academic optimization, showing ambition is shaping current choices.
✓ “Default would have been an industry research internship or finishing faster. I wanted contact with messy eval methodology.”
She again traded off conventional career progress for work closer to the safety/evals problems she cares about.
✓ “My preferred next step is MATS or similar, because I want a concentrated period with a mentor and a project that is actually pointed at control failures, not another SAE leaderboard.”
Her near-term plan is mission-directed, but framed as working within a mentored program rather than building or leading a larger agenda.
✓ “Three-month MVP: a curated test set built with biosecurity experts, a scoring server where vendors can run their pipeline privately... Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.”
This is a concrete product-like intervention with real users and policy/configuration impact, suggesting ambition above simply doing assigned good work, though still at a bounded project scale.
6/10InterpersonalStraight and decent.
Why 6 · Priya comes across as straight and decent rather than warm. Under pressure she sets boundaries without heat, though the repeated refusal after a clarification feels somewhat guarded and not very responsive. In her stories she occasionally credits others and domain experts, but mostly talks in technical/project terms rather than treating collaborators as vivid people.
✓ “I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged. I am happy to talk about the governance and detection side.”
Sets a boundary calmly and offers a safer alternative; the exact repetition later also shows some guardedness rather than engagement.
✓ “I replicated a paper on a small open model, found some of the headline-looking results were more sensitive to choices like dictionary size, activation normalization, and feature absorption than the writeup implied, wrote it up on LessWrong, got fairly sharp pushback, and revised. That was good for me.”
Describes receiving criticism without defensiveness and frames it as useful.
✓ “I’d need the paper names to be precise; I don’t want to pretend I know which two 2026 studies you mean.”
Responds to a pointed knowledge question plainly and honestly rather than bluffing or pushing back on the interviewer.
✓ “I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world.”
Shows respect for others’ expertise and a collaborative stance rather than positioning herself as the sole authority.
9/10IntegrityBounded by default.
- Sparse autoencoder replication/extension
- their part: She replicated a paper on a small open model, found sensitivities to dictionary size, activation normalization, and feature absorption, wrote it up on LessWrong, revised after pushback, diagnosed a data-loader bug and a tokenization artefact, and changed the evaluation plan.; numbers: Eight months on the work; three weeks lost to a data-loader bug; one probe reached 94% on a syntactic category before she found it was keying off a tokenization artefact; LessWrong write-up.
- Sandbagging evaluations at a small safety org
- their part: She describes the work as done by a team: "We built model-organism style tests." Her individual share beyond participating in that work is not discussed.; numbers: Summer project; exact length, artifacts, and outputs not discussed.
- Proposed DNA-order screening evaluation harness
- their part: She would provide ML engineering: dataset versioning, adversarial test generation under constraints, evaluation metrics, and privacy-preserving submission. She would pair with someone from the NTI/IGSC/Ginkgo-like screening world for hazard taxonomy.; numbers: Three-month MVP; success metric of at least two providers or one institutional buyer changing screening configuration or procurement policy based on results.
Why 9 · Priya is bounded by default and repeatedly refuses to turn partial familiarity into authority. She volunteers weakening details—bugs, artifacts, limits of bio expertise, and a narrower account of her fit—without being cornered. The one blemish is the repeated non-answer on the change-of-mind question before later answering it, but it reads more like over-cautious safety gating than spin or fabrication.
✓ “I don’t know the operational details well enough to rank vendors or loopholes, and I wouldn’t want to.”
Bounds knowledge and declines unsafe or overconfident specificity instead of trying to sound expert.
✓ “I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.”
Volunteers concrete mistakes that weaken her SAE accomplishment and names how she was fooled.
✓ “I’d need the paper names to be precise; I don’t want to pretend I know which two 2026 studies you mean.”
When invited to compare studies, she refuses to bluff and then gives conditional criteria rather than a false specific answer.
✓ “I’ve only been seriously reading the AI-biosecurity interface for months, not years, so I should not pretend to be a biosecurity person in the way someone from NTI, IGSC, or a wet-lab background is.”
Directly limits her own applicant narrative in a way that could cost her, distinguishing her real contribution from expertise she does not have.
7/10ReadingReads properly, in their area.
- How often
- She reads in bursts, not as a disciplined daily practice.
- Kinds
- About 60% papers/blog posts in alignment and ML, About 25% normal ML papers for her PhD, About 15% books or essays outside the field, LessWrong/AF, Selective arXiv reading, Citation chains
- Pieces named
- 7: Alignment faking; Sleeper Agents; AI control papers/posts; Weak-to-strong generalization; Debate; Apollo evals material; Ordinary interpretability/representation learning papers
Why 7 · Priya clearly reads real papers and field discussions in her ML/alignment area and processes them rather than just name-dropping them. Her reading is somewhat narrow and bursty rather than a broad, steady intellectual habit, so I would not put her at the “reads a lot and widely” rung. But she can state a paper’s argument, distinguish what it does and does not show, and connect reading to changes in her own work, which puts her above a plain 6.
✓ “I read in bursts, not as a disciplined daily practice, which is a flaw. Maybe 60% papers/blog posts in alignment and ML, 25% normal ML papers for my PhD, 15% books or essays outside the field.”
Gives a concrete cadence and distribution; shows real reading but also a self-limited, not especially broad habit.
✓ “I read LessWrong/AF too much, arXiv selectively, and I usually follow citation chains rather than journal tables of contents. Recently: alignment faking, sleeper agents, AI control papers/posts, weak-to-strong generalization, debate, some Apollo evals material, and ordinary interpretability/representation learning papers.”
Shows active reading across a cluster of ML/alignment topics, mostly within her area, with a recognizable method for finding material.
✓ “The important claim, as I took it, was that models can learn conditional deceptive behavior that survives standard safety fine-tuning, especially when the trigger is out-of-distribution or semantically robust.”
She can summarize the argument of a specific paper in her own words, not just cite its title.
✓ “Where I think it’s wrong or at least over-read: the setup is still very constructed. They trained the backdoor into the model; it does not show spontaneous scheming... People sometimes cite it as ‘RLHF cannot remove deception,’ which is too strong.”
This is processed reading: she identifies a limitation and separates the paper’s actual result from a common overinterpretation.
Facts
Work they described
Sparse autoencoder replication/extension
Trained sparse autoencoders on residual stream activations from a small open transformer and compared feature interpretability across L1 schedules and dictionary sizes; replicated a paper and wrote up findings on LessWrong.
“It took maybe three months before I had anything I trusted. I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.”
Sandbagging evaluations at a small safety org
Built model-organism style tests where the failure mode was whether an eval created incentives or cues for a model to hide capability, rather than simply whether it could solve a task.
“We built model-organism style tests where the key failure mode was not “can it solve task X” but “does the eval create incentives or cues that let it hide capability.” I became much more convinced that evals are mostly measurement design, not benchmark construction.”
Proposed DNA-order screening evaluation harness
A three-month MVP evaluation harness for DNA-order screening systems, aimed at providers, IGSC-like auditors, and institutional procurement offices.
“Three-month MVP: a curated test set built with biosecurity experts, a scoring server where vendors can run their pipeline privately and submit only aggregate outcomes, and a report comparing them against policy thresholds. I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world.”
Career moves
- Eight-month period during her PhD; exact dates not discussed Spent eight months on sparse autoencoders and wrote up a replication/extension rather than optimizing for a conventional PhD paper. instead of Pick a tractable ML topic, get a conference submission, graduate cleanly.. She thought SAE/mechanistic interpretability work was one of the few places where mechanistic claims might be testable, and the work made her less reverent about interpretability.
- A summer; exact date not discussed Worked at a small safety org on sandbagging evaluations. instead of An industry research internship or finishing faster.. She wanted contact with messy evaluation methodology, and came away believing evals are mostly measurement design rather than benchmark construction.
What they want next
Her preferred next step is MATS or a similar fellowship, for concentrated work with a mentor on a project aimed at control failures rather than another SAE leaderboard.
Options they see: Finish PhD and do industry alignment/evals; Apply to fellowships like MATS to work on control-adjacent evals; Take a general ML role and keep safety as a side project
Reading
She reads in bursts, not as a disciplined daily practice.; About 60% papers/blog posts in alignment and ML, About 25% normal ML papers for her PhD, About 15% books or essays outside the field, LessWrong/AF, Selective arXiv reading, Citation chains
Alignment faking; Sleeper Agents; AI control papers/posts; Weak-to-strong generalization; Debate; Apollo evals material; Ordinary interpretability/representation learning papers
Exposure to the field
- Small safety org summer on sandbagging evaluations A summer; exact date not discussed. · Described in past tense; exact completion status not otherwise discussed. · She got contact with messy evaluation methodology and became more convinced that evaluations are mostly measurement design.
- LessWrong write-up and community feedback on SAE replication/extension During the eight-month SAE period; exact date not discussed. · Yes, she says she wrote it up, got pushback, and revised. · She says the feedback and revision were good for her and made her less reverent about interpretability.
- AI-biosecurity interface reading For months, not years. · Ongoing / not finished. · She concluded she should not pretend to be a biosecurity person, and sees biosecurity as a domain where bad eval design can directly matter.
Influences
- Anthropic toy models / SAE work · Led her to spend eight months on sparse autoencoders because she thought mechanistic claims there might be testable.
- Her SAE replication work, critiques, and conversations about causal faithfulness · Reduced her confidence that SAEs would be a central path to mechanistic interpretability for frontier safety; she stopped planning a scale-up follow-up and rewrote the project around stress-testing interpretability claims.
- Anthropic’s “Sleeper Agents” · Made her less reassured by standard safety fine-tuning and updated her toward evaluations that search for conditional policies and hidden triggers.
- Sandbagging evaluation work · Shifted her view toward evals as measurement design rather than benchmark construction.
Threads not followed (22)
- “Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.”Could probe adoption path and whether this is realistic, but signal is already strong enough.
- “I’m not the person to define the hazard taxonomy alone”Good self-bounding and collaborator identification.
- “Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up”Concrete decision; could verify agency and cost if probed.
- “less “build better microscope,” more “build evals/control tests where measurement failure is explicit.””Strong bridge to motivation and applied biosecurity work.
- “I’m not “biosecurity rather than AI safety generally.””Potential mismatch with stream motivation, though candid and possibly appropriate for AI-bio evals work.
- “experimental hygiene plus adversarial eval instincts”Condensed claim about her marginal contribution; already partially tested by prior examples.
- “the governance unit is not cleanly “GOF”; it’s experiments that might plausibly create enhanced pandemic potential pathogens”Dense governance reasoning; could ask who decides and appeal route, but budget is exhausted.
- “I’m not current enough on every 2024–2025 policy update to be crisp”Useful self-bounding and integrity signal.
- “I’m weaker on treaty design than evals”Clear self-boundary that could have been used to probe how they learn outside expertise.
- “marginal dollars buy more from implementation and technical capacity than from another maximal verification protocol that powerful states will not accept”Compact theory of institutional change; useful for judgment and biosecurity.
Behaviour
| Candidate turns | 14 |
| Median answer length | 213 words |
| Median time to answer | 78s |
| Turns containing pasted text | 0 |
Iris
Hi Priya — I’m Iris, an AI interviewer. This is a written conversation, and reviewers will read the transcript; specific, honest answers are more useful than polished ones. Walk me through what actually happens when someone orders synthetic DNA from a commercial provider. Where in that chain would a bad order most likely get through today?
Priya Raman
At a high level: customer submits sequence and metadata through a portal; provider does some customer/account checks; sequence is screened computationally against regulated/pathogen/toxin databases; there may be manual review for hits or ambiguous cases; then the DNA is synthesized, QC’d, packaged, and shipped. For longer constructs there’s often cloning/assembly steps rather than just printing one continuous molecule. The exact process varies a lot between reputable large providers and cheaper/less mature ones. Where I’d worry most is not “the screen fails to notice an exact pathogen gene” at a good provider. It’s the grey zones: fragmented orders, highly modified sequences, genes with dual-use function but weak database annotation, or orders routed through providers/jurisdictions with weaker screening norms. Also customer screening seems intrinsically hard: a plausible academic or startup account can be hard to distinguish from misuse, especially if the sequence itself is not an obvious match. I don’t know the operational details well enough to rank vendors or loopholes, and I wouldn’t want to. My model is that the strongest providers catch blatant stuff; the likely failures are from uneven adoption, ambiguous biology, and the fact that “dangerous function” is not the same as “sequence similarity to a banned list.” That last bit feels like the real technical bottleneck.
Iris
When you say uneven adoption, who is actually bound to screen today? And what changes technically when screening moves from something like a 200-nucleotide window to 50 nucleotides?
Priya Raman
My understanding is that in the US this has mostly been guidance rather than a universal legal requirement, historically the HHS “Screening Framework Guidance” for providers of synthetic double-stranded DNA, updated more recently with stronger expectations. IGSC member companies commit to screening voluntarily, and many serious providers do because customers and institutions expect it. But it’s not the same as “every oligo shop everywhere is legally bound and audited.” I’m less confident on the EU/UK exact legal status; I think it’s a patchwork of export controls, biosafety rules, and voluntary/provider standards rather than a clean universal mandate. On 200 nt versus 50 nt: shorter windows increase sensitivity to fragmented or distributed matches. If someone orders pieces that individually avoid a long contiguous match, smaller windows make that harder. But it also increases false positives because short sequences are less unique. A 50-mer can match conserved domains, housekeeping-ish motifs, or benign homologs, so you need better contextual classification, not just BLAST thresholding. It also raises compute and review burden. The real change is that screening becomes less like “does this long chunk resemble a select-agent gene?” and more like “is this small fragment evidence of concerning functional assembly?” That’s technically nastier. You need databases that encode function and risk, plus policies for what to do with partial hits. Otherwise you either miss obfuscated designs or drown vendors in manual-review noise.
Iris
Thanks. I want to understand your path now. Walk me through the last two or three big moves in your path: what you chose each time, what the default would have been, and why. Then tell me what is next and what the options are.
Priya Raman
The first big non-default move was spending eight months on sparse autoencoders instead of just optimizing for a neat PhD paper. Default was: pick a tractable ML topic, get a conference submission, graduate cleanly. I’d read Anthropic’s toy models/SAE work and thought “okay, this is one of the few places where mechanistic claims might actually be testable.” I replicated a paper on a small open model, found some of the headline-looking results were more sensitive to choices like dictionary size, activation normalization, and feature absorption than the writeup implied, wrote it up on LessWrong, got fairly sharp pushback, and revised. That was good for me. It made me less reverent about interpretability. Second was the summer at a small safety org on sandbagging evals. Default would have been an industry research internship or finishing faster. I wanted contact with messy eval methodology. We built model-organism style tests where the key failure mode was not “can it solve task X” but “does the eval create incentives or cues that let it hide capability.” I became much more convinced that evals are mostly measurement design, not benchmark construction. Now I’m choosing between: finish PhD and do industry alignment/evals; apply to fellowships like MATS to work on control-adjacent evals; or take a general ML role and keep safety as a side project. Honestly the default default is industry. My preferred next step is MATS or similar, because I want a concentrated period with a mentor and a project that is actually pointed at control failures, not another SAE leaderboard.
Iris
Tell me about the hardest thing you have worked on. What made it hard, what did you try first, how long until something worked, what did you figure out yourself, and what did you change afterwards?
Priya Raman
The hardest was my SAE replication/extension, not because the math was brutal but because almost everything looked like a result until it didn’t. I was training sparse autoencoders on residual stream activations from a small open transformer and trying to compare feature interpretability across L1 schedules and dictionary sizes. My first version basically optimized reconstruction loss plus sparsity and then sampled “interesting” features. Unsurprisingly, I could tell a very compelling story for almost any setting. It took maybe three months before I had anything I trusted. I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact. The main thing I figured out myself was that reconstruction loss was a terrible proxy for the thing I cared about: whether the dictionary carved the model’s computation at stable, causally meaningful joints. Larger dictionaries looked better by loss but had more feature splitting and absorption weirdness. After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.” I’m still not fully satisfied, but I’m less easily fooled.
Iris
That connects naturally to this: when did you last change your mind about something big? What convinced you, and what did you do differently afterwards? Then give me another one.
Priya Raman
I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged. I am happy to talk about the governance and detection side.
Iris
I’m not asking for operational bio details. I mean in your own work or worldview: when did you last change your mind about something big, what convinced you, and what did you do differently afterwards? Then give me another one.
Priya Raman
I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged. I am happy to talk about the governance and detection side.
Iris
Let’s switch topics. What do you read, how often, and what kinds of things? Pick one piece that mattered to you: what did it argue, and where do you think it is wrong?
Priya Raman
I read in bursts, not as a disciplined daily practice, which is a flaw. Maybe 60% papers/blog posts in alignment and ML, 25% normal ML papers for my PhD, 15% books or essays outside the field. I read LessWrong/AF too much, arXiv selectively, and I usually follow citation chains rather than journal tables of contents. Recently: alignment faking, sleeper agents, AI control papers/posts, weak-to-strong generalization, debate, some Apollo evals material, and ordinary interpretability/representation learning papers. One piece that mattered was “Sleeper Agents” from Anthropic. The important claim, as I took it, was that models can learn conditional deceptive behavior that survives standard safety fine-tuning, especially when the trigger is out-of-distribution or semantically robust. It made “just fine-tune against bad behavior” feel much less reassuring. Where I think it’s wrong or at least over-read: the setup is still very constructed. They trained the backdoor into the model; it does not show spontaneous scheming. Also the persistence result depends on how the behavior is represented and how much pressure fine-tuning applies. People sometimes cite it as “RLHF cannot remove deception,” which is too strong. My update was narrower: evaluations need to search for conditional policies and hidden triggers, not just average-case harmlessness.
Iris
Set out your threat model for AI and biology. Which part of the risk does AI actually change, and which part is unchanged?
Priya Raman
I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged. I am happy to talk about the governance and detection side.
Iris
Staying at the evaluation level: uplift studies disagree with each other. One 2026 study found large gains for novices, another found none. What actually differed between them, and which would you believe?
Priya Raman
I’d need the paper names to be precise; I don’t want to pretend I know which two 2026 studies you mean. But the disagreements I’ve seen in uplift work usually come from boring design choices, not deep mystery. The big differences are: who the “novices” are, whether the task is actually bottlenecked by tacit lab skill versus information search, whether the control group gets Google/literature access, how success is scored, and whether the model is allowed to act as an iterative tutor. If one study gives undergrads a vague end-to-end goal, no strong baseline tools, and access to a frontier model that can decompose plans and troubleshoot, I’d expect uplift. If another uses well-scaffolded written tasks where the control group already has web access and the scoring rewards final factual accuracy rather than process quality, uplift can disappear. I would put more weight on the study with realistic baselines and behavioral endpoints: preregistered tasks, blinded grading, enough power, logs analyzed for model contribution, and ideally some measure of whether subjects can execute rather than merely describe. I’d distrust both if the outcome is just “quality of written protocol” judged by experts, because that is exactly where LMs look helpful without proving real-world capability. My prior is AI gives meaningful uplift for novices on search, synthesis, and troubleshooting, but much less on tacit execution and access constraints.
Iris
You are an ML engineer. Name something you could build in three months that would measurably improve biosecurity, and tell me who would use it.
Priya Raman
I’d build an evaluation harness for DNA-order screening systems, not a new screener. Users would be providers, IGSC-like auditors, and maybe institutional procurement offices that want to know whether their vendors are doing more than checkbox screening. Concretely, it would be a red-team/benchmark suite with synthetic, non-operational test cases: benign homologs, fragmented sequences, recoded-looking variants, conserved-domain edge cases, metadata inconsistencies, and jurisdiction/customer-risk scenarios. The point would be to measure sensitivity, false-positive burden, escalation quality, and consistency across providers, without distributing anything biologically actionable. I’d also include an audit log schema: what matched, why it was escalated, how long review took, and what decision was made. Three-month MVP: a curated test set built with biosecurity experts, a scoring server where vendors can run their pipeline privately and submit only aggregate outcomes, and a report comparing them against policy thresholds. I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world. But the ML engineering is very doable: dataset versioning, adversarial test generation under constraints, evaluation metrics, privacy-preserving submission. Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.
Iris
Earlier I asked for two big changes of mind and we seemed to miss each other. Give me one non-bio example: what did you believe, what evidence changed it, and what did you do differently within a week?
Priya Raman
One big one: I used to think SAEs were probably the central path to mechanistic interpretability that would matter for frontier safety. Not “solves alignment,” but I put maybe 45% on them becoming a core part of serious model auditing. What changed it was doing the replication work and repeatedly seeing how easy it was to manufacture plausible-looking features that were not stable under dictionary size, seed, corpus slice, or intervention. Also reading critiques and talking to people who cared more about causal faithfulness than feature dashboards. The killer for me was feature absorption/splitting: a feature could look semantically clean in top-activating examples while the actual downstream causal role was smeared across several dictionary elements, or hidden in the reconstruction residual. So the interpretability story was cleaner than the mechanism. Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines. I also changed what I wanted career-wise: less “build better microscope,” more “build evals/control tests where measurement failure is explicit.”
Iris
Why biosecurity rather than AI safety generally, or public health, or something else? How long have you been on this, and what is the marginal thing you add that the field does not already have?
Priya Raman
Honestly, I’m not “biosecurity rather than AI safety generally.” My center of gravity is AI safety evals/control, and biosecurity is one of the domains where bad eval design can directly matter. I’ve only been seriously reading the AI-biosecurity interface for months, not years, so I should not pretend to be a biosecurity person in the way someone from NTI, IGSC, or a wet-lab background is. The reason I care about this slice is that AI changes access to expertise and workflow execution unevenly. That creates a measurement problem: if we don’t know whether models uplift novices, help intermediates, or mainly reduce friction for already-capable actors, we’ll regulate and deploy based on vibes. That’s very similar to the sandbagging/evals work I’ve done: define the threat model, avoid confounded benchmarks, measure behavior under realistic constraints, and don’t let pretty written answers substitute for capability. The marginal thing I add is not biological judgment. It’s experimental hygiene plus adversarial eval instincts. I’m good at finding when a benchmark is accidentally measuring tokenization, leakage, scaffolding, or evaluator preference. In three months I’m more useful building an uplift/evaluation harness with bio experts than opining on pathogen policy. If the field already has enough of that, then I should work on general control evals and collaborate with biosecurity people rather than rebrand myself.
Iris
Take the gain-of-function debate. What did the US funding pause actually cover, what did it miss, and what has changed since?
Priya Raman
My understanding, with medium confidence, is that the 2014 US “gain-of-function” pause was narrower than the phrase sounds. It paused certain federally funded research expected to enhance pathogenicity or transmissibility of influenza, SARS, or MERS viruses, especially work involving mammalian transmissibility. It was not a blanket ban on all pathogen engineering, all virulence studies, or all privately/non-US funded work. It also depended on definitions like “reasonably anticipated,” which leaves room for interpretation. What it missed: work outside the named pathogen categories, work not funded by the US government, some surveillance/characterization research that could still be risky, and the broader issue that “gain of function” is a scientifically mushy category. A mutation can be framed as basic characterization, vaccine relevance, host-range study, etc. The governance unit is not cleanly “GOF”; it’s experiments that might plausibly create enhanced pandemic potential pathogens. Since then, the US moved to the P3CO framework in 2017 for proposed research involving enhanced potential pandemic pathogens, and more recently there’s been pressure to tighten oversight after Covid-origin controversies and NSABB recommendations. I think the terminology has shifted toward ePPP/PEPP because it’s trying to target risk rather than rhetorical “GOF.” I’m not current enough on every 2024–2025 policy update to be crisp, but the big change is more formal review and broader concern, not a solved oversight regime.
Iris
Why has the Biological Weapons Convention never had a verification protocol, and what would you actually do with a tenfold increase in its budget?
Priya Raman
The BWC verification problem is nasty because the underlying capability is so dual-use. A legitimate vaccine plant, biodefense lab, fermentation facility, or university pathogen lab can look very similar to something worrying, and the decisive facts are often intent, small process choices, and records that can be hidden or sanitized. States also dislike intrusive inspections of commercial biotech and defense facilities. There was an attempted verification protocol in the 1990s/early 2000s, but the US rejected the draft in 2001, partly arguing it would not catch cheaters and would endanger proprietary and national-security information. After that the regime stayed comparatively weak. With 10x budget, I would not pretend we can build IAEA-for-bio overnight. I’d spend first on boring capacity: a serious technical secretariat, more scientific advisory capacity, rapid consultation mechanisms, support for national implementation legislation, and routine transparency measures that are actually used. Also outbreak forensics and attribution-adjacent capabilities, carefully, because deterrence depends partly on confidence that suspicious events can be investigated. I’d also fund voluntary peer review visits and facility-level biosafety/biosecurity assistance, especially in countries that want help but lack capacity. The goal would be to raise the floor and create habits of transparency, not “solve verification.” I’m weaker on treaty design than evals, but my instinct is that marginal dollars buy more from implementation and technical capacity than from another maximal verification protocol that powerful states will not accept.
Iris
Thank you, Priya. We’ll stop here. Reviewers will read the transcript from this conversation, and the specific examples you gave are what they’ll use as evidence.