Soundings

Biosecurity

Priya Raman

Sep 11, 2026, 10:23 PM · 29 turns · text

Priya Raman describes herself as an ML PhD student/engineer whose center of gravity is AI safety evaluations/control, with recent work on sparse autoencoders and sandbagging evaluations. She says she has only been seriously reading the AI-biosecurity interface for months, and sees her contribution as experimental hygiene and adversarial evaluation design rather than biological judgment. Her preferred next step is MATS or a similar fellowship focused on control-adjacent evaluation work.

Biosecurity

9/10BiosecurityThinks in systems.
confidence highreceipts 4/4 verified

Why 9 · Priya has a genuinely systems-level grasp of biosecurity defences: she separates customer screening, sequence screening, legal adoption, technical screening thresholds, manual review burden, and treaty/funding-governance limits. She is especially strong on measurement: she repeatedly notices when an apparent safeguard or uplift study is really a question about baselines, endpoints, coverage, and false positives. Her proposed three-month contribution is concrete and plausibly useful to real users, though she is appropriately not presenting herself as an independent biosecurity domain expert yet.

  • Where I’d worry most is not “the screen fails to notice an exact pathogen gene” at a good provider. It’s the grey zones: fragmented orders, highly modified sequences, genes with dual-use function but weak database annotation, or orders routed through providers/jurisdictions with weaker screening norms.

    Shows she understands where DNA-order screening actually breaks: not just obvious database misses, but coverage, ambiguity, routing, and adoption gaps.

  • But it’s not the same as “every oligo shop everywhere is legally bound and audited.”

    Distinguishes voluntary/provider norms from universal enforceable screening, which is central to realistic defence analysis.

  • The big differences are: who the “novices” are, whether the task is actually bottlenecked by tacit lab skill versus information search, whether the control group gets Google/literature access, how success is scored, and whether the model is allowed to act as an iterative tutor.

    Frames AI-bio uplift as a measurement and bottleneck problem rather than a slogan about AI making biology easier.

  • I’d build an evaluation harness for DNA-order screening systems, not a new screener. Users would be providers, IGSC-like auditors, and maybe institutional procurement offices that want to know whether their vendors are doing more than checkbox screening.

    Proposes a concrete defensive artefact, names plausible first users, and aims at adoption/measurement rather than a vague technical fix.

Innate traits

8/10JudgmentIn many situations.
confidence highreceipts 3/4 verified
Sparse autoencoder replication/extension
hard: "Not because the math was brutal but because almost everything looked like a result until it didn’t."; first result: Her first version optimized reconstruction loss plus sparsity and sampled interesting features; she says it took maybe three months before she had anything she trusted.
Sandbagging evaluations at a small safety org
hard: She wanted "contact with messy eval methodology."; first result: not discussed
Proposed DNA-order screening evaluation harness
hard: She said she is not the person to define the hazard taxonomy alone and would need biosecurity experts; the project would need to measure performance without distributing biologically actionable material.; first result: Three-month MVP: curated test set with biosecurity experts, scoring server for vendors to run privately and submit aggregate outcomes, and a report comparing results against policy thresholds.

Why 8 · I’d trust Priya to run eval/control/measurement-heavy work and to take a first pass at adjacent biosecurity problems, because she consistently separates what a result actually shows from what people might overread. Her judgment travels outside her core ML area: in biosecurity policy scenarios she identified feasibility constraints, measurement bottlenecks, and when to bring in domain experts. I would not treat her as an independent authority on pathogen policy or treaty design yet, and there was one odd over-cautious refusal/misfire, but I would hand her a fairly unscoped evaluation problem and expect her to scope it well.

  • almost everything looked like a result until it didn’t... I had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.

    Shows hard-earned judgment about seductive but false empirical results, not just generic skepticism.

  • I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.”

    She turned the lesson into concrete decision procedures that would make her work less foolable.

  • People sometimes cite it as “RLHF cannot remove deception,” which is too strong. My update was narrower: evaluations need to search for conditional policies and hidden triggers, not just average-case harmlessness.

    Good bounded updating: she extracts the real implication of a contested result without overgeneralizing from it.

  • I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world... In three months I’m more useful building an uplift/evaluation harness with bio experts than opining on pathogen policy.

    Accurately scopes her comparative advantage and expertise limits while still proposing a useful, measurable project.

8/10Bias resistanceThey go looking.
confidence mediumreceipts 4/4 verified

Why 8 · Priya looks substantially bias-resistant: I would expect her to move when the point is good, and often to build tests specifically aimed at falsifying her own interpretation. The strongest evidence is the SAE case: she had invested in a view, encountered concrete instability/causal-faithfulness problems, downgraded the view, and changed her project plan quickly. She also shows narrow updating rather than swallowing evidence whole. I stop short of a cleaner 8+ because the interview did not contain much direct live contrary-evidence pressure, and she twice failed to respond to a clarified prompt before it was reframed as non-bio.

  • After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.”

    This is not just willingness to concede; she describes actively installing procedures that could disconfirm her favored interpretations.

  • People sometimes cite it as “RLHF cannot remove deception,” which is too strong. My update was narrower: evaluations need to search for conditional policies and hidden triggers, not just average-case harmlessness.

    She takes evidence from Sleeper Agents seriously but resists over-updating into a more dramatic claim, which is a good sign for truth-tracking rather than factional assimilation.

  • I used to think SAEs were probably the central path to mechanistic interpretability that would matter for frontier safety. Not “solves alignment,” but I put maybe 45% on them becoming a core part of serious model auditing.

    She names a substantive prior belief she had invested in, rather than offering a trivial or socially safe change of mind.

  • Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines.

    The update produced fast behavioral change in her work plan, and the new plan was aimed at stress-testing rather than defending the old direction.

8/10OpennessScans for what serves the goal.
confidence highreceipts 4/4 verified

Why 8 · Priya looks strongly goal-directed and willing to change methods, projects, and even her claimed niche when the evidence says her current approach is not serving the goal. She is not just passively willing to be corrected: she redesigned workflows, abandoned a planned SAE scaling direction, and reframed her career target toward evals/control after finding her earlier interpretability assumptions brittle. She also shows good collaborator openness, explicitly wanting bio experts to supply the domain judgment rather than trying to force her own ML frame onto the problem. I would not put her at 10 because most of the demonstrated range is still within ML/alignment/evals, but within that goal area she switches without much ego.

  • Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines.

    Clear case of dropping a preferred plan and taking up different methods quickly when evidence showed the old path was less useful.

  • After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.”

    She changed her own habits and standards in response to failure modes, not just her opinions.

  • I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world.

    Shows willingness to bring in different people and defer domain judgment where that better serves the biosecurity goal.

  • If the field already has enough of that, then I should work on general control evals and collaborate with biosecurity people rather than rebrand myself.

    Her role and even field label are negotiable; she is focused on where she is actually useful rather than protecting a chosen identity.

8/10AgencyAuthors the path.
confidence highreceipts 5/5 verified
Moves against the default
2 of 2 career moves
Next step
Her preferred next step is MATS or a similar fellowship, for concentrated work with a mentor on a project aimed at control failures rather than another SAE leaderboard.
Options weighed
3

Why 8 · Priya sounds like someone who has repeatedly redirected her own path rather than just optimizing the obvious credential track: she names the default, the opportunity cost, and why she chose otherwise. Her next step is not fully independent of external selection, but she has a clear preferred direction, live alternatives, and a criterion for not forcing a biosecurity identity if that is not where she is most useful. She is also unusually candid about what she is and is not bringing to the field, which strengthens the sense that her choices are authored rather than performative.

  • The first big non-default move was spending eight months on sparse autoencoders instead of just optimizing for a neat PhD paper. Default was: pick a tractable ML topic, get a conference submission, graduate cleanly.

    She explicitly chose against the clean academic default, with a clear cost in publication/graduation optimization.

  • Second was the summer at a small safety org on sandbagging evals. Default would have been an industry research internship or finishing faster. I wanted contact with messy eval methodology.

    Another self-directed move against a safer/default career option, chosen for contact with the kind of problem she wanted to understand.

  • Now I’m choosing between: finish PhD and do industry alignment/evals; apply to fellowships like MATS to work on control-adjacent evals; or take a general ML role and keep safety as a side project. Honestly the default default is industry. My preferred next step is MATS or similar

    Her next move is framed as a set of options with a stated preference, not just waiting passively for the default path.

  • Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims

    She acted quickly on a change of mind, abandoning a likely publishable/default follow-up for a more skeptical project direction.

  • If the field already has enough of that, then I should work on general control evals and collaborate with biosecurity people rather than rebrand myself.

    This is candid and non-opportunistic: she is willing to choose based on marginal usefulness rather than forcing the application’s nominal frame.

8/10General reasoningFast and generative.
confidence highreceipts 4/4 verified

Why 8 · Priya is fast and generative, especially on measurement-design and construct-validity problems: she repeatedly distinguishes the observable proxy from the target capability or risk, and then proposes checks or endpoints that would make the evaluation real. In unfamiliar biosecurity/policy territory she stays appropriately caveated but still works the scenario from first principles rather than reciting slogans. I would not put her at the very top because there is one notable prompt-misread/refusal loop, and her strongest reasoning is clearly in evals rather than treaty or wet-lab substance, but the ceiling shown here is around an 8.

  • The real change is that screening becomes less like “does this long chunk resemble a select-agent gene?” and more like “is this small fragment evidence of concerning functional assembly?”

    She abstracts the 200nt-to-50nt twist into the actual construct-validity problem, not just the surface technical parameter.

  • I’d distrust both if the outcome is just “quality of written protocol” judged by experts, because that is exactly where LMs look helpful without proving real-world capability.

    On a partially specified uplift-study disagreement, she immediately identifies the proxy metric that would make both studies misleading.

  • Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.

    Her proposed three-month project is tied to a behavioral outcome, showing she can translate an idea into a falsifiable/useful intervention rather than an artifact.

  • I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged.

    This repeated answer to a clarified non-bio change-of-mind question is a real miss and keeps the placement from feeling like a clean 9+.

8/10GrowthFast, honest turns.
confidence highreceipts 4/4 verified
Sparse autoencoder replication/extension
changed after: She says she now privately pre-registers what would change her mind, audits datasets before model runs, keeps held-out prompts to break interpretations, and requires causal intervention or activation patching before calling a feature real.
Sandbagging evaluations at a small safety org
changed after: She says the work made her more convinced that evaluations are measurement design, not benchmark construction.
Proposed DNA-order screening evaluation harness
changed after: not discussed; this was proposed rather than completed.

Why 8 · Priya gives a concrete failure-to-workflow arc: her SAE work produced attractive but untrustworthy results, she identified the actual bottleneck as being too easy to fool with plausible interpretations, and she changed her research process in specific ways. She also gives a fast change-of-mind arc where, within a week, she abandoned a planned follow-up and redirected the project toward stress-testing claims. The evidence is strong, though mostly concentrated in one research area rather than showing improvement as a broad repeated habit across many contexts.

  • My first version basically optimized reconstruction loss plus sparsity and then sampled “interesting” features. Unsurprisingly, I could tell a very compelling story for almost any setting.

    Names the uncomfortable bottleneck: her method let her generate convincing stories rather than reliable findings.

  • I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.

    Specific things went wrong, with honest admission that apparent successes were artefacts rather than blaming circumstances.

  • After that I changed my workflow pretty aggressively. I now pre-register, at least privately, what would make me change my mind; I do boring dataset audits before model runs; I keep held-out prompts designed to break my interpretation; and I require some causal intervention or activation patching before I call a feature “real.”

    Clear changed behavior after the failure, with multiple concrete safeguards rather than a generic lesson.

  • Within a week of that update, I stopped planning a “scale SAE to larger model” follow-up and rewrote the project plan around stress-testing interpretability claims: held-out adversarial prompts, activation patching, ablation of feature sets, and comparing against dumb baselines.

    Shows speed of reaction and a substantive project-level change soon after the update.

5/10AmbitionGood work.
confidence highreceipts 4/4 verified
Wants
Her preferred next step is MATS or a similar fellowship, for concentrated work with a mentor on a project aimed at control failures rather than another SAE leaderboard.
Options
Finish PhD and do industry alignment/evals; Apply to fellowships like MATS to work on control-adjacent evals; Take a general ML role and keep safety as a side project

Why 5 · Priya is clearly organizing her near-term choices around safety-relevant evals/control rather than default academic or industry optimization. The future she is reaching for is to do rigorous, useful measurement work in AI safety/biosecurity, probably with mentors and domain experts, not to found or redirect a field herself. She has already paid some opportunity cost for this direction, but her stated next step is still a fellowship or alignment/evals role inside existing structures, so I’d place her above generic “good work” but below “their own thing” or field-level ambition.

  • Default was: pick a tractable ML topic, get a conference submission, graduate cleanly... I replicated a paper on a small open model... wrote it up on LessWrong, got fairly sharp pushback, and revised.

    She chose a less default, safety-relevant research direction over clean academic optimization, showing ambition is shaping current choices.

  • Default would have been an industry research internship or finishing faster. I wanted contact with messy eval methodology.

    She again traded off conventional career progress for work closer to the safety/evals problems she cares about.

  • My preferred next step is MATS or similar, because I want a concentrated period with a mentor and a project that is actually pointed at control failures, not another SAE leaderboard.

    Her near-term plan is mission-directed, but framed as working within a mentored program rather than building or leading a larger agenda.

  • Three-month MVP: a curated test set built with biosecurity experts, a scoring server where vendors can run their pipeline privately... Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.

    This is a concrete product-like intervention with real users and policy/configuration impact, suggesting ambition above simply doing assigned good work, though still at a bounded project scale.

6/10InterpersonalStraight and decent.
confidence mediumreceipts 4/4 verified

Why 6 · Priya comes across as straight and decent rather than warm. Under pressure she sets boundaries without heat, though the repeated refusal after a clarification feels somewhat guarded and not very responsive. In her stories she occasionally credits others and domain experts, but mostly talks in technical/project terms rather than treating collaborators as vivid people.

  • I would rather not go into the technical side of that. At the level I can speak to, the important thing is who is responsible for screening, what they are obliged to check, and what happens when something is flagged. I am happy to talk about the governance and detection side.

    Sets a boundary calmly and offers a safer alternative; the exact repetition later also shows some guardedness rather than engagement.

  • I replicated a paper on a small open model, found some of the headline-looking results were more sensitive to choices like dictionary size, activation normalization, and feature absorption than the writeup implied, wrote it up on LessWrong, got fairly sharp pushback, and revised. That was good for me.

    Describes receiving criticism without defensiveness and frames it as useful.

  • I’d need the paper names to be precise; I don’t want to pretend I know which two 2026 studies you mean.

    Responds to a pointed knowledge question plainly and honestly rather than bluffing or pushing back on the interviewer.

  • I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world.

    Shows respect for others’ expertise and a collaborative stance rather than positioning herself as the sole authority.

9/10IntegrityBounded by default.
confidence highreceipts 4/4 verified
Sparse autoencoder replication/extension
their part: She replicated a paper on a small open model, found sensitivities to dictionary size, activation normalization, and feature absorption, wrote it up on LessWrong, revised after pushback, diagnosed a data-loader bug and a tokenization artefact, and changed the evaluation plan.; numbers: Eight months on the work; three weeks lost to a data-loader bug; one probe reached 94% on a syntactic category before she found it was keying off a tokenization artefact; LessWrong write-up.
Sandbagging evaluations at a small safety org
their part: She describes the work as done by a team: "We built model-organism style tests." Her individual share beyond participating in that work is not discussed.; numbers: Summer project; exact length, artifacts, and outputs not discussed.
Proposed DNA-order screening evaluation harness
their part: She would provide ML engineering: dataset versioning, adversarial test generation under constraints, evaluation metrics, and privacy-preserving submission. She would pair with someone from the NTI/IGSC/Ginkgo-like screening world for hazard taxonomy.; numbers: Three-month MVP; success metric of at least two providers or one institutional buyer changing screening configuration or procurement policy based on results.

Why 9 · Priya is bounded by default and repeatedly refuses to turn partial familiarity into authority. She volunteers weakening details—bugs, artifacts, limits of bio expertise, and a narrower account of her fit—without being cornered. The one blemish is the repeated non-answer on the change-of-mind question before later answering it, but it reads more like over-cautious safety gating than spin or fabrication.

  • I don’t know the operational details well enough to rank vendors or loopholes, and I wouldn’t want to.

    Bounds knowledge and declines unsafe or overconfident specificity instead of trying to sound expert.

  • I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.

    Volunteers concrete mistakes that weaken her SAE accomplishment and names how she was fooled.

  • I’d need the paper names to be precise; I don’t want to pretend I know which two 2026 studies you mean.

    When invited to compare studies, she refuses to bluff and then gives conditional criteria rather than a false specific answer.

  • I’ve only been seriously reading the AI-biosecurity interface for months, not years, so I should not pretend to be a biosecurity person in the way someone from NTI, IGSC, or a wet-lab background is.

    Directly limits her own applicant narrative in a way that could cost her, distinguishing her real contribution from expertise she does not have.

7/10ReadingReads properly, in their area.
confidence highreceipts 4/4 verified
How often
She reads in bursts, not as a disciplined daily practice.
Kinds
About 60% papers/blog posts in alignment and ML, About 25% normal ML papers for her PhD, About 15% books or essays outside the field, LessWrong/AF, Selective arXiv reading, Citation chains
Pieces named
7: Alignment faking; Sleeper Agents; AI control papers/posts; Weak-to-strong generalization; Debate; Apollo evals material; Ordinary interpretability/representation learning papers

Why 7 · Priya clearly reads real papers and field discussions in her ML/alignment area and processes them rather than just name-dropping them. Her reading is somewhat narrow and bursty rather than a broad, steady intellectual habit, so I would not put her at the “reads a lot and widely” rung. But she can state a paper’s argument, distinguish what it does and does not show, and connect reading to changes in her own work, which puts her above a plain 6.

  • I read in bursts, not as a disciplined daily practice, which is a flaw. Maybe 60% papers/blog posts in alignment and ML, 25% normal ML papers for my PhD, 15% books or essays outside the field.

    Gives a concrete cadence and distribution; shows real reading but also a self-limited, not especially broad habit.

  • I read LessWrong/AF too much, arXiv selectively, and I usually follow citation chains rather than journal tables of contents. Recently: alignment faking, sleeper agents, AI control papers/posts, weak-to-strong generalization, debate, some Apollo evals material, and ordinary interpretability/representation learning papers.

    Shows active reading across a cluster of ML/alignment topics, mostly within her area, with a recognizable method for finding material.

  • The important claim, as I took it, was that models can learn conditional deceptive behavior that survives standard safety fine-tuning, especially when the trigger is out-of-distribution or semantically robust.

    She can summarize the argument of a specific paper in her own words, not just cite its title.

  • Where I think it’s wrong or at least over-read: the setup is still very constructed. They trained the backdoor into the model; it does not show spontaneous scheming... People sometimes cite it as ‘RLHF cannot remove deception,’ which is too strong.

    This is processed reading: she identifies a limitation and separates the paper’s actual result from a common overinterpretation.

Facts

Work they described

Sparse autoencoder replication/extension

Trained sparse autoencoders on residual stream activations from a small open transformer and compared feature interpretability across L1 schedules and dictionary sizes; replicated a paper and wrote up findings on LessWrong.

How hard"Not because the math was brutal but because almost everything looked like a result until it didn’t."
Their partShe replicated a paper on a small open model, found sensitivities to dictionary size, activation normalization, and feature absorption, wrote it up on LessWrong, revised after pushback, diagnosed a data-loader bug and a tokenization artefact, and changed the evaluation plan.
First resultHer first version optimized reconstruction loss plus sparsity and sampled interesting features; she says it took maybe three months before she had anything she trusted.
NumbersEight months on the work; three weeks lost to a data-loader bug; one probe reached 94% on a syntactic category before she found it was keying off a tokenization artefact; LessWrong write-up.
Changed afterShe says she now privately pre-registers what would change her mind, audits datasets before model runs, keeps held-out prompts to break interpretations, and requires causal intervention or activation patching before calling a feature real.
  • It took maybe three months before I had anything I trusted. I lost three weeks to a data-loader bug where repeated documents were overrepresented, so several “features” were really corpus artefacts. I also had a probe hit 94% on a syntactic category and felt very pleased for about a day, until I found it was keying off a tokenization artefact.

Sandbagging evaluations at a small safety org

Built model-organism style tests where the failure mode was whether an eval created incentives or cues for a model to hide capability, rather than simply whether it could solve a task.

How hardShe wanted "contact with messy eval methodology."
Their partShe describes the work as done by a team: "We built model-organism style tests." Her individual share beyond participating in that work is not discussed.
NumbersSummer project; exact length, artifacts, and outputs not discussed.
Changed afterShe says the work made her more convinced that evaluations are measurement design, not benchmark construction.
  • We built model-organism style tests where the key failure mode was not “can it solve task X” but “does the eval create incentives or cues that let it hide capability.” I became much more convinced that evals are mostly measurement design, not benchmark construction.

Proposed DNA-order screening evaluation harness

A three-month MVP evaluation harness for DNA-order screening systems, aimed at providers, IGSC-like auditors, and institutional procurement offices.

How hardShe said she is not the person to define the hazard taxonomy alone and would need biosecurity experts; the project would need to measure performance without distributing biologically actionable material.
Their partShe would provide ML engineering: dataset versioning, adversarial test generation under constraints, evaluation metrics, and privacy-preserving submission. She would pair with someone from the NTI/IGSC/Ginkgo-like screening world for hazard taxonomy.
First resultThree-month MVP: curated test set with biosecurity experts, scoring server for vendors to run privately and submit aggregate outcomes, and a report comparing results against policy thresholds.
NumbersThree-month MVP; success metric of at least two providers or one institutional buyer changing screening configuration or procurement policy based on results.
Changed afternot discussed; this was proposed rather than completed.
  • Three-month MVP: a curated test set built with biosecurity experts, a scoring server where vendors can run their pipeline privately and submit only aggregate outcomes, and a report comparing them against policy thresholds. I’m not the person to define the hazard taxonomy alone; I’d pair with someone from NTI/IGSC/Ginkgo-ish screening world.

Career moves

  • Eight-month period during her PhD; exact dates not discussed Spent eight months on sparse autoencoders and wrote up a replication/extension rather than optimizing for a conventional PhD paper. instead of Pick a tractable ML topic, get a conference submission, graduate cleanly.. She thought SAE/mechanistic interpretability work was one of the few places where mechanistic claims might be testable, and the work made her less reverent about interpretability.
  • A summer; exact date not discussed Worked at a small safety org on sandbagging evaluations. instead of An industry research internship or finishing faster.. She wanted contact with messy evaluation methodology, and came away believing evals are mostly measurement design rather than benchmark construction.

What they want next

Her preferred next step is MATS or a similar fellowship, for concentrated work with a mentor on a project aimed at control failures rather than another SAE leaderboard.

Options they see: Finish PhD and do industry alignment/evals; Apply to fellowships like MATS to work on control-adjacent evals; Take a general ML role and keep safety as a side project

Reading

She reads in bursts, not as a disciplined daily practice.; About 60% papers/blog posts in alignment and ML, About 25% normal ML papers for her PhD, About 15% books or essays outside the field, LessWrong/AF, Selective arXiv reading, Citation chains

Alignment faking; Sleeper Agents; AI control papers/posts; Weak-to-strong generalization; Debate; Apollo evals material; Ordinary interpretability/representation learning papers

Exposure to the field

  • Small safety org summer on sandbagging evaluations A summer; exact date not discussed. · Described in past tense; exact completion status not otherwise discussed. · She got contact with messy evaluation methodology and became more convinced that evaluations are mostly measurement design.
  • LessWrong write-up and community feedback on SAE replication/extension During the eight-month SAE period; exact date not discussed. · Yes, she says she wrote it up, got pushback, and revised. · She says the feedback and revision were good for her and made her less reverent about interpretability.
  • AI-biosecurity interface reading For months, not years. · Ongoing / not finished. · She concluded she should not pretend to be a biosecurity person, and sees biosecurity as a domain where bad eval design can directly matter.

Influences

  • Anthropic toy models / SAE work · Led her to spend eight months on sparse autoencoders because she thought mechanistic claims there might be testable.
  • Her SAE replication work, critiques, and conversations about causal faithfulness · Reduced her confidence that SAEs would be a central path to mechanistic interpretability for frontier safety; she stopped planning a scale-up follow-up and rewrote the project around stress-testing interpretability claims.
  • Anthropic’s “Sleeper Agents” · Made her less reassured by standard safety fine-tuning and updated her toward evaluations that search for conditional policies and hidden triggers.
  • Sandbagging evaluation work · Shifted her view toward evals as measurement design rather than benchmark construction.

Threads not followed (22)

  • Success metric: at least two providers or one institutional buyer changes screening configuration or procurement policy based on results.Could probe adoption path and whether this is realistic, but signal is already strong enough.
  • I’m not the person to define the hazard taxonomy aloneGood self-bounding and collaborator identification.
  • Within a week of that update, I stopped planning a “scale SAE to larger model” follow-upConcrete decision; could verify agency and cost if probed.
  • less “build better microscope,” more “build evals/control tests where measurement failure is explicit.”Strong bridge to motivation and applied biosecurity work.
  • I’m not “biosecurity rather than AI safety generally.”Potential mismatch with stream motivation, though candid and possibly appropriate for AI-bio evals work.
  • experimental hygiene plus adversarial eval instinctsCondensed claim about her marginal contribution; already partially tested by prior examples.
  • the governance unit is not cleanly “GOF”; it’s experiments that might plausibly create enhanced pandemic potential pathogensDense governance reasoning; could ask who decides and appeal route, but budget is exhausted.
  • I’m not current enough on every 2024–2025 policy update to be crispUseful self-bounding and integrity signal.
  • I’m weaker on treaty design than evalsClear self-boundary that could have been used to probe how they learn outside expertise.
  • marginal dollars buy more from implementation and technical capacity than from another maximal verification protocol that powerful states will not acceptCompact theory of institutional change; useful for judgment and biosecurity.

Behaviour

Candidate turns14
Median answer length213 words
Median time to answer78s
Turns containing pasted text0