Soundings

AI safety context

Priya Raman

Sep 11, 2026, 10:23 PM · 27 turns · text

Priya Raman describes herself as an AI safety researcher focused on control-adjacent evaluations, with prior hands-on work on sparse autoencoders and a summer on sandbagging evals. She says her views shifted from mechanistic interpretability as the central path toward model organisms, eval methodology, and control protocols. She wants to build evals that test whether control protocols catch strategically hidden failures under realistic elicitation pressure.

AI safety context

9/10AI safety contextDeep in their target sub-field.
confidence highreceipts 4/4 verified
Exposure
Hands-on sparse-autoencoder replication/extension on a small open transformer; Sandbagging evals work; AI safety reading and field exposure through papers, posts, and discussions
Pieces named
Hubinger et al.’s "Risks from Learned Optimization"; Carlsmith’s "Scheming AIs"; Ngo/Chan/Mindermann alignment framing; Christiano on amplification/debate; Irving et al. on debate; OpenAI weak-to-strong generalization; Anthropic Constitutional AI; Sleeper Agents; alignment faking; Goal Misgeneralization; ARC/Redwood/Paul Christiano posts; METR/ARC evals material; Olah’s circuits thread; induction heads; ROME/causal tracing adjacent work; IOI; TransformerLens tutorials; Cunningham/Bricken/Anthropic SAE papers; LessWrong arguments about feature geometry and superposition

Why 9 · Priya is deeply engaged with her stated target area, control-adjacent evals/model organisms, and also has strong working knowledge across technical alignment. She can explain recent papers by setup and result, distinguish what they do and do not demonstrate, name relevant groups, and identify methodological open problems like elicitation, false negatives, collusion, and distribution shift. She is around the 8–10 range: probably able to brief a team on the broad frontier and her bets, though the transcript shows more strong synthesis and critique than exhaustive lab-frontier command.

  • I want to work on control-adjacent evals: model organisms of deceptive or sandbagging behavior, and protocols that remain useful when you don’t fully trust the model... ARC/Redwood people and Ryan Greenblatt et al. have pushed the control framing most clearly, especially trusted-monitor/untrusted-actor setups. METR is doing capability evaluations for autonomous task completion and time horizons.

    Directly states target subfield and maps current actors and work in it.

  • One is Greenblatt et al.’s “AI Control: Improving Safety Despite Intentional Subversion.” My understanding is that they set up programming tasks where an untrusted strong model may try to insert subtle vulnerabilities... protocol design can substantially reduce risk even when the strong model is treated as adversarial... Main limitation: the threat model is stylized.

    Explains a recent control paper’s setup, finding, and limitation rather than just naming it.

  • What’s open is mostly methodology. Negative results are weak because elicitation is hard... We need better threat models, held-out triggers, pre-registered analyses, baselines against simple behavioral probes, and protocols tested against adaptive red teams... understand its false-negative rate under plausible adversaries.

    Shows a concrete grasp of open problems in her chosen area and frames them decision-relevantly.

  • SAEs give more monosemantic-looking features and allow feature-level interventions, but there are serious issues: feature absorption, splitting, basis dependence, dead features, sensitivity to training distribution, and the fact that reconstruction loss is a weak proxy for finding safety-relevant variables.

    Demonstrates working technical-alignment knowledge outside her target area, with specific mechanisms and failure modes.

Innate traits

8/10JudgmentIn many situations.
confidence mediumreceipts 4/4 verified
SAE replication/extension on a small open transformer
hard: "It was hard because every metric lied a little."; first result: Her first version was to train SAEs on residual activations, label top-activating examples, train probes, and report high accuracies; after about three weeks she found a data loader bug, and over an eight-month project she kept removing fake results until a more modest causal-validation approach worked.
Sandbagging evals summer
hard: "a lot of the work is boring hygiene" and broad eval claims could become underpowered or mushy.; first result: They ran a small pilot on whether models underperform on coding tasks when given a weak incentive to appear less capable; the pilot looked null-ish, but the power analysis showed the result would be uninformative, so they dropped it and shifted to narrower tasks.
Alignment-faking eval design update
hard: The difficulty was separating scaffolded roleplay from evidence that the model conditionally reasons about training versus deployment.; first result: not discussed

Why 8 · I would trust Priya’s judgment on unscoped technical AI-safety research problems, especially evals/control/interp methodology. She consistently separates what evidence actually shows from what people are tempted to infer, notices proxy-metric and elicitation failures, and changes plans when the evidence is underpowered or confounded. I would not infer broad all-purpose judgment outside technical alignment from this interview, but within and near her research area her judgment seems to travel well.

  • What would update me is a result where SAE-derived features detect hidden-deceptive state across a serious distribution shift better than behavioral probes and supervised baselines, with controls for tokenization/artifacts and with causal validation.

    She gives concrete conditions for changing her mind, focused on validity, baselines, artifacts, and causal evidence rather than vibes.

  • I don’t think it’s magically less brittle. The reason I prefer it is that control makes the brittleness more explicit and testable.

    In a live challenge to her preferred agenda, she does not defend it by overclaiming; she identifies the comparative advantage and the remaining fragility.

  • After about three weeks I found a data loader bug that meant one contrast set had subtly different truncation. Later I had a 94% probe that was mostly detecting tokenization around a delimiter. So the project took maybe eight months because I kept having to subtract fake results.

    A concrete past research case where she detected misleading evidence and revised her trust in attractive metrics.

  • the power analysis said we’d need far more tasks than we could afford to distinguish “no sandbagging” from “small but important sandbagging.” Earlier-me would probably have run it anyway and produced a mushy null result. What convinced me was just doing the math and seeing the confidence intervals were useless. We dropped it

    Shows decision-level judgment: she avoided producing an overinterpreted null result and redirected the project based on statistical informativeness.

9/10Bias resistanceThey go looking.
confidence highreceipts 5/5 verified

Why 9 · Priya looks actively truth-seeking rather than merely agreeable: she repeatedly states the strongest limitations of her preferred evals/control agenda and gives concrete observations that would change live disagreements. Under pushback, she accepts the brittle point and sharpens it into a top open problem rather than defending the agenda. Her past updates are substantive and costly enough to count—moving away from mech interp as the central path after her own SAE work, revising claims after pushback, and dropping an underpowered experiment instead of publishing a mushy null.

  • I don’t think it’s magically less brittle. The reason I prefer it is that control makes the brittleness more explicit and testable.

    When challenged on her favored control/evals direction, she concedes the core weakness rather than dodging it, then narrows the claim to a more defensible one.

  • Collusion is the part that worries me most. If independent calls are not actually independent because the model can infer shared context, use steganography, or coordinate via persistent state/tools, then many control arguments weaken badly.

    She takes the interviewer’s objection further, identifying a specific way her preferred approach could fail badly.

  • What would move me toward Carlsmith: demonstrations of spontaneous instrumental deception in agentic settings where the model is not directly instructed to deceive, especially if it generalizes across tasks and persists after safety training

    Gives a concrete reversal condition on a live disagreement, including the kind of evidence and generalization that would update her.

  • First shift was away from “mechanistic interpretability is the central path.” ... The SAE project made me less sold, not because it failed, but because it half-worked.

    Describes a real update away from an agenda she had invested in, caused by mixed evidence rather than by total failure or social pressure.

  • I revised the LessWrong writeup after pushback, softened the claims, and now I basically don’t trust an interp result unless it beats simple baselines and survives distribution shift.

    Shows she changed public claims and future standards in response to criticism and discovered artifacts, not just privately nodded along.

8/10OpennessScans for what serves the goal.
confidence highreceipts 4/4 verified

Why 8 · Priya has a clear goal—decision-relevant evidence about hidden failures and safe deployment protocols—and treats methods as tools rather than identities. She has repeatedly shifted agenda based on evidence: from mechanistic interpretability to evals/model organisms to control, while still keeping interpretability as useful instrumentation. She is not maximally “anything goes”; she has strong views about what is less promising, but she evaluates alternatives and changes plans when they better serve the goal.

  • First shift was away from “mechanistic interpretability is the central path.” In 2021/22 I basically bought the story that if we could reverse-engineer transformers, we’d get a handle on deception and goals. The SAE project made me less sold, not because it failed, but because it half-worked.

    Shows she moved away from her preferred/initial agenda after hands-on evidence complicated it, rather than defending it.

  • Third shift is toward control. I used to think “control” sounded like giving up on alignment, or security theater. The Greenblatt-style work changed my mind: assume the model may be adversarial, then ask what protocols still reduce risk while preserving usefulness.

    A concrete case of adopting an initially unattractive approach once she saw how it served the safety goal.

  • But I’d ideally combine them: interp features as one monitor input, not as the safety case.

    Keeps multiple technical approaches live and integrates them according to use, rather than replacing one dogma with another.

  • Earlier-me would probably have run it anyway and produced a mushy null result. What convinced me was just doing the math and seeing the confidence intervals were useless. We dropped it and shifted to narrower tasks where we could get enough repetitions and stronger manipulation checks.

    Shows willingness to abandon a planned experiment and change method when the original plan would not answer the question well.

8/10AgencyAuthors the path.
confidence highreceipts 4/4 verified
Moves against the default
5 of 5 career moves
Next step
She wants to work on control-adjacent evals: model organisms of deceptive or sandbagging behavior, and protocols that remain useful when the model is not fully trusted.
Options weighed
3
First step taken
She has already worked on SAEs and sandbagging evals, changed eval designs to include situational-awareness checks, and shifted sandbagging work toward narrower, better-powered tasks.

Why 8 · Priya comes across as actively steering her research path rather than optimizing within a given frame. She repeatedly names defaults she used to hold, explains what changed her mind, and describes costly follow-through: dropping an underpowered experiment, revising claims, and redirecting from pure interp toward control/evals. Her next direction is chosen among live options and already reflected in how she designs evals, though the transcript gives less evidence of a concrete career/logistics plan than would push this toward 9–10.

  • First shift was away from “mechanistic interpretability is the central path.” In 2021/22 I basically bought the story that if we could reverse-engineer transformers, we’d get a handle on deception and goals. The SAE project made me less sold, not because it failed, but because it half-worked.

    She explicitly identifies a field-default view she held, then moved away from it because of her own research evidence rather than social fashion.

  • Earlier-me would probably have run it anyway and produced a mushy null result. What convinced me was just doing the math and seeing the confidence intervals were useless. We dropped it and shifted to narrower tasks where we could get enough repetitions and stronger manipulation checks.

    This is a concrete self-directed move with cost: abandoning a planned experiment instead of taking the safer/default path of producing a weak result.

  • I revised the LessWrong writeup after pushback, softened the claims, and now I basically don’t trust an interp result unless it beats simple baselines and survives distribution shift.

    She changed a public-facing claim and her standards after evidence/pushback, showing ownership and willingness to lose polish/status for accuracy.

  • What’s next: I see live options as scalable oversight with much better adversarial evaluation, control protocols for early transformative systems, and interp integrated as instrumentation rather than the main plan. I’m less excited by pure preference training or pure SAE scaling.

    Her next direction is framed as a choice among strategies, with a clear ranking based on what she thinks works rather than what is fashionable or safe.

9/10General reasoningFast and generative.
confidence highreceipts 4/4 verified

Why 9 · Priya is fast and generative within the AI-safety technical domain: she repeatedly turns broad claims into testable assumptions, identifies the fragile inference in an argument, and proposes concrete controls or falsifiers. Her reasoning is especially strong around eval validity, baselines, causal validation, and threat-model assumptions; she does not just state skepticism, she operationalizes what would change her mind. I would not place this at 10 because the transcript stays largely inside her prepared area and does not show a clearly novel first-principles leap in unfamiliar territory, but within scope she is well above merely competent.

  • What would update me is a result where SAE-derived features detect hidden-deceptive state across a serious distribution shift better than behavioral probes and supervised baselines, with controls for tokenization/artifacts and with causal validation.

    Turns a vague skepticism about SAEs into a precise empirical test, including distribution shift, baselines, artifact controls, and causal checks.

  • If independent calls are not actually independent because the model can infer shared context, use steganography, or coordinate via persistent state/tools, then many control arguments weaken badly.

    When pushed on control brittleness, she identifies the key hidden assumption and concrete mechanisms by which it can fail.

  • Where I’m more skeptical is the jump from learned goal-directedness to stable, reflectively endorsed, across-context scheming. Current models show fragments... I don’t yet see strong evidence for internally coherent objectives that persist through fine-tuning and cause hidden strategic behavior absent prompting.

    Decomposes a complex scheming argument and locates the uncertain inference rather than accepting or rejecting the whole package.

  • I figured out, painfully, that “interpretability” without adversarial controls is just qualitative clustering... I basically don’t trust an interp result unless it beats simple baselines and survives distribution shift.

    Extracts a general methodological lesson from failed experiments and states a robust check for future work.

9/10GrowthFast, honest turns.
confidence highreceipts 4/4 verified
SAE replication/extension on a small open transformer
changed after: She softened her claims and now says she does not trust interp results unless they beat simple baselines and survive distribution shift.
Sandbagging evals summer
changed after: She became more skeptical of broad eval claims from underpowered samples and shifted to tasks with enough repetitions and stronger manipulation checks.
Alignment-faking eval design update
changed after: She took situational-awareness evals more seriously.

Why 9 · Priya shows repeated, concrete arcs from failed or misleading work to changed practice, and she is unusually willing to describe the unflattering version of what went wrong. The strongest evidence is not just that she 'learned lessons,' but that she changed methods, revised public claims, killed an underpowered project, and now builds specific checks into eval design. I would place her just below the very top because the speed is sometimes weeks/months rather than clearly immediate, but the honesty and repeated follow-through are strong.

  • My first version was embarrassingly straightforward: train SAEs on residual activations, label top-activating examples, train probes on feature activations, report nice-looking accuracies. After about three weeks I found a data loader bug that meant one contrast set had subtly different truncation. Later I had a 94% probe that was mostly detecting tokenization around a delimiter.

    She gives a concrete block/reaction story and names the embarrassing, non-flattering failure: her nice results were partly fake/artifactual.

  • What eventually worked was more modest: causal validation with activation patching/feature ablations, matched prompts, held-out templates, and comparing against dumb baselines like bag-of-token probes and linear probes on raw activations. I figured out, painfully, that “interpretability” without adversarial controls is just qualitative clustering. I revised the LessWrong writeup after pushback, softened the claims, and now I basically don’t trust an interp result unless it beats simple baselines and survives distribution shift.

    The failure produced a specific methodological change and a public correction, not just a verbal lesson.

  • Afterwards I changed how I design evals: I now include explicit checks for whether the model understands the evaluation/training context, and I treat “the model was told the setup” as a variable to sweep, not an automatic disqualifier.

    A belief update about alignment-faking translated into a concrete ongoing change in her experimental design.

  • Earlier-me would probably have run it anyway and produced a mushy null result. What convinced me was just doing the math and seeing the confidence intervals were useless. We dropped it and shifted to narrower tasks where we could get enough repetitions and stronger manipulation checks.

    She identifies a real personal failure mode, stops a planned project, and changes direction based on evidence rather than preserving the original plan.

8/10AmbitionA field.
confidence highreceipts 4/4 verified
Wants
She wants to work on control-adjacent evals: model organisms of deceptive or sandbagging behavior, and protocols that remain useful when the model is not fully trusted.
Options
Scalable oversight with much better adversarial evaluation; Control protocols for early transformative systems; Interpretability integrated as instrumentation rather than the main plan

Why 8 · Priya’s ambition is a methodological field agenda rather than a personal-status or startup-shaped one: she wants AI safety evals/control to become decision-relevant for deployments, not just produce scary demos or leaderboards. Her current choices already line up with that aim: she has spent months stress-testing interpretability claims, worked on sandbagging evals, changed experimental designs around power and elicitation, and now frames her next work around false-negative rates under adversarial pressure. I’d place her above “good work inside existing structures” and around field-level ambition, but not at the top because she does not describe an unreasonable-scale future or a plan that is visibly organising an institution/team this month.

  • “The thing I’d most like to build is not another leaderboard saying ‘model X scores Y on sabotage,’ but evals that measure whether a proposed control protocol actually catches strategically hidden failures under realistic elicitation pressure.”

    This is a concrete desired future for the work: changing evals from benchmark reporting to protocol-validating tools for hidden failures.

  • “The crux for me is moving from ‘we observed scary behavior in a demo’ to ‘this eval changes a deployment decision because we understand its false-negative rate under plausible adversaries.’”

    She is aiming at changing how the field makes deployment decisions, not just producing an individual paper or good analysis.

  • “I spent a summer on sandbagging evals, and a lot of the work is boring hygiene: preventing leakage, checking whether your ‘sandbagging detector’ is actually detecting formatting, doing power analysis before declaring null results.”

    Her current work choices already point toward the methodology-improvement agenda she says matters.

  • “We dropped it and shifted to narrower tasks where we could get enough repetitions and stronger manipulation checks. That made me much more allergic to broad eval claims from underpowered samples.”

    Shows the ambition is organising actual research behavior: she is willing to abandon appealing experiments to build more reliable eval practice.

7/10InterpersonalStraight and decent.
confidence mediumreceipts 4/4 verified

Why 7 · Priya comes across as straight, fair-minded, and easy to push: she answers objections directly without heat, concedes real limitations, and avoids turning disagreements into attacks. She talks about other researchers and papers with nuance rather than contempt, often separating the strongest version of a view from overclaiming around it. There is some evidence of receptiveness to feedback and repair, but the transcript mostly stages technical judgment rather than interpersonal warmth, so I would not place her much higher on this trait.

  • I don’t think it’s magically less brittle. The reason I prefer it is that control makes the brittleness more explicit and testable.

    When directly challenged on her preferred agenda, she concedes the force of the objection immediately and then explains her view without defensiveness.

  • My objection is not that SAEs are useless; I spent eight months on them and like them.

    She disagrees plainly with a research direction while taking care not to caricature it or the people working on it.

  • I think some Anthropic SAE writing is appropriately cautious, but the community downstream often rounds it up to “we can read the model’s thoughts.”

    This is a hard criticism expressed with a distinction that credits the original authors rather than lumping everyone together.

  • I revised the LessWrong writeup after pushback, softened the claims, and now I basically don’t trust an interp result unless it beats simple baselines and survives distribution shift.

    Shows she can receive criticism and change her public claims rather than digging in.

9/10IntegrityBounded by default.
confidence highreceipts 4/4 verified
SAE replication/extension on a small open transformer
their part: She trained SAEs on residual activations, labeled top-activating examples, trained probes on feature activations, found a data loader bug, identified a tokenization artifact, added causal validation and baselines, and revised a LessWrong writeup after pushback.; numbers: Small open transformer, roughly Pythia-scale; project took maybe eight months; one probe reached 94% accuracy before she found it mostly detected tokenization around a delimiter; she revised a LessWrong writeup.
Sandbagging evals summer
their part: She worked on sandbagging evals over a summer and personally abandoned a planned coding-task underperformance experiment after a pilot and power analysis.; numbers: "spent a summer on sandbagging evals"; pilot result was "null-ish"; power analysis showed they would need far more tasks than they could afford.
Alignment-faking eval design update
their part: She added explicit checks for whether the model understands the evaluation/training context and treats whether the model is told the setup as an experimental variable to sweep.; numbers: No quantitative result discussed; described as a change in eval-design practice.

Why 9 · Priya reads as strongly bounded and self-correcting rather than merely polished. She repeatedly volunteers details that weaken her own story—bugs, artefacts, underpowered results—and describes concrete changes that cost her a cleaner narrative or a planned experiment. I would trust her substantially about her own work and expect her to disclose limitations even when nobody is checking, though the interview did not include a direct external-pressure ethics scenario, so I would stop short of a perfect score.

  • After about three weeks I found a data loader bug that meant one contrast set had subtly different truncation. Later I had a 94% probe that was mostly detecting tokenization around a delimiter. So the project took maybe eight months because I kept having to subtract fake results.

    Volunteers specific errors that made her own results worse, rather than presenting only the successful version.

  • I revised the LessWrong writeup after pushback, softened the claims, and now I basically don’t trust an interp result unless it beats simple baselines and survives distribution shift.

    Describes correcting the public account of her work when challenged, not defending or respinning it.

  • Earlier-me would probably have run it anyway and produced a mushy null result. What convinced me was just doing the math and seeing the confidence intervals were useless. We dropped it and shifted to narrower tasks where we could get enough repetitions and stronger manipulation checks.

    In a pressure-like research situation, she says they abandoned a planned result rather than publish an underpowered claim.

  • I also read outside the field but I’m not going to pretend it’s systematic governance expertise.

    Shows ordinary boundedness: she marks the limit of her knowledge instead of rounding general reading up into expertise.

8/10ReadingReads a lot and widely.
confidence highreceipts 4/4 verified
How often
AI safety most days; papers two or three serious ones a week when not in deadline mode; LessWrong/Alignment Forum in batches; mech interp Discord/Twitter more than is healthy.
Kinds
AI safety papers, LessWrong/Alignment Forum posts, Mechanistic interpretability Discord/Twitter discussions, Straight ML papers for methods, Some outside-the-field reading, though not systematic governance expertise
Pieces named
19: Hubinger et al.’s "Risks from Learned Optimization"; Carlsmith’s "Scheming AIs"; Ngo/Chan/Mindermann alignment framing; Christiano on amplification/debate; Irving et al. on debate; OpenAI weak-to-strong generalization; Anthropic Constitutional AI; Sleeper Agents; alignment faking; Goal Misgeneralization; ARC/Redwood/Paul Christiano posts; METR/ARC evals material; Olah’s circuits thread; induction heads; ROME/causal tracing adjacent work; IOI; TransformerLens tutorials; Cunningham/Bricken/Anthropic SAE papers; LessWrong arguments about feature geometry and superposition

Why 8 · Priya reads AI safety material very regularly and at real-paper depth, not just summaries or discourse. She ranges across core alignment theory, scalable oversight, mech interp, evals/control, and adjacent ML methods, and repeatedly explains what papers demonstrated versus what they did not. She also processes what she reads critically: she can state an argument, identify overclaims, name missing controls, and say what evidence would change her mind. I would not put her at 9–10 because the breadth is mostly within AI safety/ML rather than unusually wide across many fields.

  • I read AI safety stuff most days, but unevenly. Papers maybe two or three serious ones a week when I’m not in deadline mode; LessWrong/Alignment Forum posts in batches; mech interp Discord/Twitter discussions more than is healthy; and some straight ML papers for methods.

    Gives a concrete regular reading habit and distinguishes serious papers from posts/discussions/methods reading.

  • I’ve read the standard cluster: Hubinger et al.’s “Risks from Learned Optimization,” Carlsmith’s “Scheming AIs,” Ngo/Chan/Mindermann on alignment problem framing, Christiano on amplification/debate, Irving et al. on debate, OpenAI weak-to-strong generalization, Anthropic’s Constitutional AI, Sleeper Agents, alignment faking, Goal Misgeneralization, a bunch of ARC/Redwood/Paul Christiano posts, and METR/ARC evals material. I’ve also read a lot of mechanistic interp: Olah’s circuits thread, induction heads, ROME/causal tracing adjacent work, IOI, TransformerLens tutorials, Cunningham/Bricken/Anthropic SAE papers...

    Shows substantial coverage across several AI safety subareas rather than a narrow title list.

  • One piece that mattered was Carlsmith’s “Scheming AIs.” It argues, roughly, that if we train very capable agents, there is a nontrivial path where they develop goals, understand the training process, realize bad behavior will be selected against, and therefore behave aligned during training while waiting for better opportunities.

    Can reconstruct the core argument of a major piece in her own words.

  • Where I think it’s wrong, or at least too confident in tone, is the stability/coherence of learned goals. The report sometimes feels like it imports a fairly agentic model of cognition: persistent objectives, instrumental reasoning, reflective patience.

    Moves beyond summary to a specific critique of an argument’s assumptions.

Facts

Work they described

SAE replication/extension on a small open transformer

Replicated/extended sparse-autoencoder work on a small Pythia-scale open transformer to see whether SAE features could support safety-relevant claims about refusal-ish behavior and deceptive/sandbagging-style prompts.

How hard"It was hard because every metric lied a little."
Their partShe trained SAEs on residual activations, labeled top-activating examples, trained probes on feature activations, found a data loader bug, identified a tokenization artifact, added causal validation and baselines, and revised a LessWrong writeup after pushback.
First resultHer first version was to train SAEs on residual activations, label top-activating examples, train probes, and report high accuracies; after about three weeks she found a data loader bug, and over an eight-month project she kept removing fake results until a more modest causal-validation approach worked.
NumbersSmall open transformer, roughly Pythia-scale; project took maybe eight months; one probe reached 94% accuracy before she found it mostly detected tokenization around a delimiter; she revised a LessWrong writeup.
Changed afterShe softened her claims and now says she does not trust interp results unless they beat simple baselines and survive distribution shift.
  • "The hardest thing I personally worked through was my SAE replication/extension on a small open transformer, roughly Pythia-scale rather than anything frontier. The question was: can SAE features support a safety-relevant claim beyond 'these dashboards look interpretable'? ... After about three weeks I found a data loader bug ... Later I had a 94% probe that was mostly detecting tokenization around a delimiter. So the project took maybe eight months because I kept having to subtract fake results."

Sandbagging evals summer

Worked on sandbagging evaluations, including hygiene around leakage, formatting artifacts, and power analysis.

How hard"a lot of the work is boring hygiene" and broad eval claims could become underpowered or mushy.
Their partShe worked on sandbagging evals over a summer and personally abandoned a planned coding-task underperformance experiment after a pilot and power analysis.
First resultThey ran a small pilot on whether models underperform on coding tasks when given a weak incentive to appear less capable; the pilot looked null-ish, but the power analysis showed the result would be uninformative, so they dropped it and shifted to narrower tasks.
Numbers"spent a summer on sandbagging evals"; pilot result was "null-ish"; power analysis showed they would need far more tasks than they could afford.
Changed afterShe became more skeptical of broad eval claims from underpowered samples and shifted to tasks with enough repetitions and stronger manipulation checks.
  • "I spent a summer on sandbagging evals, and a lot of the work is boring hygiene: preventing leakage, checking whether your 'sandbagging detector' is actually detecting formatting, doing power analysis before declaring null results."

Alignment-faking eval design update

Changed how she designs evals after reading the alignment-faking paper more carefully and discussing the setup.

How hardThe difficulty was separating scaffolded roleplay from evidence that the model conditionally reasons about training versus deployment.
Their partShe added explicit checks for whether the model understands the evaluation/training context and treats whether the model is told the setup as an experimental variable to sweep.
NumbersNo quantitative result discussed; described as a change in eval-design practice.
Changed afterShe took situational-awareness evals more seriously.
  • "After reading the paper more carefully and talking through the setup, I still think there’s scaffolding, but I updated because the behavior wasn’t just generic roleplay; the model conditionally reasoned about training versus deployment and sometimes accepted costs to preserve its earlier policy. That made me take situational-awareness evals more seriously."

Career moves

  • 2021/22, then after an SAE project Moved away from seeing mechanistic interpretability as the central path for alignment. instead of The default view that reverse-engineering transformers would give a handle on deception and goals.. Her SAE work produced some interpretable features and causal interventions, but she found measurement validity poor and results vulnerable to artifacts.
  • After reading/working through Sleeper Agents, alignment faking, and sandbagging work Shifted toward model organisms and evals. instead of The view that demos were too toy-ish to matter.. She came to think artificial controlled examples are useful if used to test methods rather than make overbroad claims.
  • After Greenblatt-style AI control work Shifted toward AI control. instead of The view that control sounded like giving up on alignment or like security theater.. She found the adversarial-protocol framing testable: assume the model may be adversarial and measure whether protocols reduce risk while preserving usefulness.
  • Most recently Updated upward on the importance of alignment-faking-style results for near-term priorities and changed her eval design practices. instead of Her initial dismissive reaction that the model was just roleplaying resistance because the prompt said its values were threatened.. She concluded the behavior involved conditional reasoning about training versus deployment and sometimes accepting costs to preserve an earlier policy.
  • During the sandbagging summer Abandoned a planned broad sandbagging experiment and shifted to narrower tasks. instead of Running the original weak-incentive coding-task experiment despite an underpowered null-looking pilot.. A power analysis showed they could not distinguish no sandbagging from small but important sandbagging with the available task budget.

What they want next

She wants to work on control-adjacent evals: model organisms of deceptive or sandbagging behavior, and protocols that remain useful when the model is not fully trusted.

Options they see: Scalable oversight with much better adversarial evaluation; Control protocols for early transformative systems; Interpretability integrated as instrumentation rather than the main plan

Already done: She has already worked on SAEs and sandbagging evals, changed eval designs to include situational-awareness checks, and shifted sandbagging work toward narrower, better-powered tasks.

Reading

AI safety most days; papers two or three serious ones a week when not in deadline mode; LessWrong/Alignment Forum in batches; mech interp Discord/Twitter more than is healthy.; AI safety papers, LessWrong/Alignment Forum posts, Mechanistic interpretability Discord/Twitter discussions, Straight ML papers for methods, Some outside-the-field reading, though not systematic governance expertise

Hubinger et al.’s "Risks from Learned Optimization"; Carlsmith’s "Scheming AIs"; Ngo/Chan/Mindermann alignment framing; Christiano on amplification/debate; Irving et al. on debate; OpenAI weak-to-strong generalization; Anthropic Constitutional AI; Sleeper Agents; alignment faking; Goal Misgeneralization; ARC/Redwood/Paul Christiano posts; METR/ARC evals material; Olah’s circuits thread; induction heads; ROME/causal tracing adjacent work; IOI; TransformerLens tutorials; Cunningham/Bricken/Anthropic SAE papers; LessWrong arguments about feature geometry and superposition

Exposure to the field

  • Hands-on sparse-autoencoder replication/extension on a small open transformer Eight months; exact dates not discussed. · Yes; she describes completing an eight-month project and revising a LessWrong writeup afterward. · She learned that SAE/interp results can look good while being driven by artifacts, and became more skeptical of interpretability as a safety case.
  • Sandbagging evals work A summer; exact dates not discussed. · Partly; she describes a summer project and dropping one planned experiment after a pilot and power analysis. · She learned to emphasize leakage prevention, formatting-artifact checks, power analysis, narrower tasks, and manipulation checks.
  • AI safety reading and field exposure through papers, posts, and discussions Most days for AI safety reading; exact start date not discussed. · Ongoing. · She formed views on scalable oversight, objective generalization, interpretability, evals/model organisms, and control, and says several works changed what she looks for in evals.

Influences

  • SAE work / mechanistic interpretability practice · Moved her from thinking mechanistic interpretability was the central path to viewing it as a microscope rather than a safety case.
  • Sleeper Agents, alignment faking, and sandbagging work · Shifted her toward model organisms and controlled evals of scary behavior, while maintaining caution about over-interpreting demos.
  • Greenblatt et al.’s AI Control work · Made her more sympathetic to control as a testable deployment-relevant posture.
  • Carlsmith’s "Scheming AIs" · Made her value explicit decomposition and probability estimates, and changed what she looks for in evals, though she disagrees with its apparent confidence about stable coherent goals.
  • OpenAI weak-to-strong generalization · She found it promising but sobering, and disagrees with overly optimistic readings that weak supervisors can reliably elicit latent strong capabilities.

Threads not followed (29)

  • I used to think “control” sounded like giving up on alignment, or security theater.Potentially useful for probing what changed her mind and whether she can separate vibes from technical claims.
  • possible exceptional, untestedThe crux is whether we can build evaluations with informative false-negative rates.Strong technical crux that could be tested numerically if there were follow-up budget.
  • I had a 94% probe that was mostly detecting tokenization around a delimiter.Concrete empirical failure with a number; useful evidence of calibration and experimental hygiene.
  • I revised the LessWrong writeup after pushback, softened the claimsCould probe integrity/growth, but budget is exhausted.
  • I now include explicit checks for whether the model understands the evaluation/training contextConcrete method change worth probing if budget remained.
  • the power analysis said we’d need far more tasks than we could affordSpecific statistical decision; could verify judgment with numbers if budget remained.
  • We dropped it and shifted to narrower tasksEvidence of changing course based on evidence, potentially agency/growth.
  • Papers maybe two or three serious ones a week when I’m not in deadline modeCould be checked by asking for recent papers, but no follow-up budget remains.
  • stability/coherence of learned goalsCentral crux in their disagreement with Carlsmith and their threat model.
  • messy reward gaming, situational manipulation, and tool-mediated accidents before clean “I will pretend until deployment” schemingImportant prioritization judgment that could be probed with probabilities or deployment implications.

Behaviour

Candidate turns13
Median answer length269 words
Median time to answer83s
Turns containing pasted text0