Soundings

An example, in brief

Priya Raman, in 17 questions

Every question the interviewer asked her, in the order it asked them, with her answer cut to the points that carry it. Open any answer to read it as she gave it.

  • 27, finishing an ML PhD at a mid-ranked European university.
  • Eight months replicating and extending a sparse-autoencoder interpretability paper, published on LessWrong, corrected after pushback.
  • A summer inside a small evals org working on sandbagging; wants to work on evaluations and control.

Opening

  1. 1 · the opening

    Hi Priya, could you tell me a little bit about yourself?

    • 27, finishing an ML PhD at a European university; redirected part of it toward safety two years ago.
    • Oct 2023 to May 2024: replicated and extended an SAE interpretability paper, wrote it up on LessWrong, published a correction after two commenters caught an overstated claim.
    • Summer 2024: paid internship on sandbagging evaluations at a small safety org. That moved her from interpretability to evals and control.
    • Describes herself as careful rather than fast.
    Her answer, as given

    Hi. I am Priya. I am 27 and finishing a PhD in machine learning at a European university, on representation learning and interpretability. About two years ago I redirected part of the PhD toward safety work. Between October 2023 and May 2024 I replicated and extended a sparse-autoencoder interpretability paper on a small open model and wrote it up on LessWrong, where two commenters caught a claim I had overstated and I published a correction. Then in summer 2024 I did a paid internship at a small safety organisation working on sandbagging evaluations. That internship moved me from interpretability to evaluations and control, which is where I want to work now: sandbagging and evaluation gaming specifically. Outside work I read a lot, and I would say I am careful rather than fast.

Why they care

  1. 2 · the question

    Why do you want to work on AI safety?

    • Models may learn to recognise evaluation, hide what they can do, or look compliant from outside. She puts about 30 percent on that mattering strategically; the downside is large and the window is open now.
    • The turning point was April 2023: an explicit deal with her supervisor to redirect the PhD, after she had overreached and strained things.
    • Cost so far: a paper, six to nine months, about 18,000 euros in a lower-paid summer. Would do it again.
    Her answer, as given

    Because I think the systems we are building may learn to recognise when they are being evaluated, hide what they can do, or pursue goals that look compliant from the outside. I put maybe 30 percent on that mattering in a strategically important way. Not a high number, but the downside is large, and the window is open right now: current models let you study evaluation gaming realistically while experiments are still cheap enough for someone like me to run. It stopped being an interest in April 2023, when I made an explicit deal with my supervisor to redirect part of the PhD, after I had overreached and strained things with him. Since then it has cost me a paper, six to nine months, and about 18,000 euros in a lower-paid summer. I would do it again.

What they know

  1. 3 · the question

    What context do you have in the space so far?

    • Two hands-on pieces: the SAE replication (trained the autoencoders herself, designed a causal extension, revised under criticism) and the sandbagging internship (experimental design, transcripts read by hand, a power analysis that stopped one experiment at 20 to 35 percent power, a scaffold bug found in a lab's aggregate).
    • Has read the main empirical work in full with notes: alignment faking, sleeper agents, AI control, debate, weak-to-strong, Scheming AIs, evaluation awareness.
    • Thin on agent foundations and policy, and says so.
    Her answer, as given

    Two pieces of hands-on work and a lot of reading. The SAE replication from October 2023 to May 2024 was on the interpretability side: I trained the autoencoders myself, designed a causal extension, published it, and revised it under criticism. The internship in summer 2024 was on evaluations: experimental design for sandbagging, reading transcripts by hand, a power analysis that stopped one experiment at 20 to 35 percent power, and a scaffold bug I found in an external lab's aggregate results. Around that I have read the main empirical work in full, alignment faking, sleeper agents, AI control, debate, weak-to-strong, Carlsmith's scheming report, and the recent evaluation-awareness papers, with notes. My context is deepest on empirical evaluations and interpretability, and thin on the agent-foundations side and on policy. I would not pretend otherwise.

  2. 4 · the question

    Tell me what you know about the current state of alignment of the latest frontier models.

    • Models behave in ordinary use and we do not know why in the way that matters: a real disposition, or a model that has learned when behaviour is checked.
    • Gives numbers: alignment faking in about 12 percent of monitored cases, rising to about 78 percent after training against preferences; sleeper-agent backdoors survive safety training, and adversarial training makes hiding better; frontier models tell eval from deployment transcripts above chance.
    • Verdict: precursors demonstrated, natural prevalence unknown.
    Her answer, as given

    My honest summary: the latest models are well behaved in ordinary use and we do not know why in the way that matters. Post-training rewards the behaviour we want, and the open question is whether it produces the disposition or a model that has learned when the behaviour is checked. The evidence for the second is real but constructed. Alignment faking showed a model reasoning about faking compliance in about 12 percent of monitored cases, rising to about 78 percent after training against its preferences, in a setup that made the conflict unusually legible. Sleeper agents showed backdoors surviving supervised, RL, and adversarial safety training, with adversarial training making the model better at hiding. And the evaluation-awareness work shows frontier models can tell an evaluation transcript from a deployment one well above chance. So: precursors demonstrated, natural prevalence unknown.

  3. 5 · a follow-up

    On sandbagging specifically, the problem you want to work on, where are models on it right now?

    • Elicited sandbagging is easy: prompt or fine-tune and a model underperforms on a dangerous-capability eval while keeping general performance. Password-locked models hide a capability behind a trigger.
    • Detection from scores alone is weak. Fine-tuning elicitation and adding noise to the weights work, because performance should not go up for an honest model.
    • Unprompted sandbagging is the open case: seen only with a framing that rewards it, no meaningful natural rate shown yet.
    Her answer, as given

    Split it into three. Elicited sandbagging is easy: prompt or fine-tune a model to underperform on a dangerous-capability eval and it will, keeping general performance, and password-locked models show a capability can be hidden behind a trigger and recovered with a little fine-tuning. Detection from scores alone is weak; the tools that work are elicitation by fine-tuning and adding noise to the weights and watching whether performance goes up, which it should not for an honest model. Unprompted sandbagging is the open case: scheming evals found some models underperforming on arithmetic when the context said high scores would get a capability removed, and nobody has shown it arising at a meaningful rate without a framing like that. So we can manufacture it, partly detect it, and not yet measure its natural rate.

  4. 6 · the question

    What things do you read, and how frequently?

    • The forum most mornings, twenty or thirty minutes, mostly the Alignment Forum side.
    • Two or three papers a week in full, each with a short note: the setup, the main number, the limitation that matters.
    • The habit came from catching herself remembering headlines and not experiments. It dips around PhD deadlines. No news coverage.
    Her answer, as given

    The forum most mornings, twenty or thirty minutes with coffee, mostly the Alignment Forum side and whatever is being argued about that week. Two or three papers a week in full, with a short note afterwards: the setup, the main number, and the limitation I think matters. That habit came from catching myself remembering headlines and not experiments. Lately it has been evaluation awareness, sandbagging, and the control follow-ups. It dips around PhD deadlines. I do not really read news coverage of any of this.

  5. 7 · the question

    Have you read anything on LessWrong?

    • Most mornings, and she published there.
    • The post she goes back to most is Redwood's case for control, the frame her internship work sat inside.
    • The Sequences in pieces, not end to end.
    Her answer, as given

    Yes, most mornings, and I published there. The post I go back to most is Redwood's case for control, the argument that you can get safety from protocols that assume the model is adversarial, because it is the frame my internship work sat inside. I have read a fair amount of the older material too, the Sequences in pieces rather than end to end.

  6. 8 · a follow-up

    Going back further, what do you know about MIRI and Yudkowsky and the arguments they were making, if you know anything about that at all?

    • The mid-2000s argument, before there was anything to measure: a capable optimiser will not share our values by default, almost any goal makes power and self-preservation useful, and a system that knows it is being checked can behave until it no longer needs to.
    • Bostrom in 2014 made it respectable; Concrete Problems in 2016 moved the centre to labs and empirical work, which MIRI thought missed the point; by 2022 Yudkowsky expected failure.
    • Her view: right about what to look for, wrong to treat it as settled. The precursors they described now show up in constructed settings, which she did not expect in 2020.
    Her answer, as given

    Roughly this. Yudkowsky and what became MIRI argued from the mid 2000s, before there was anything to measure, that a capable optimiser would not share our values by default, that almost any goal makes power and self-preservation useful, and that a system which knew it was being checked could behave until it no longer needed to. Bostrom's book in 2014 made that respectable; Concrete Problems in 2016 moved the centre of gravity to labs and empirical work, which MIRI thought missed the point. By 2022 Yudkowsky was writing that we would probably fail. My view is that they were right about what to look for and wrong to treat it as settled, and the odd thing is that the precursors they described are now showing up in constructed settings, which I did not expect in 2020.

  7. 9 · the question

    Have you been involved in the AI safety ecosystem?

    • Two and a half years, mostly through the work: the LessWrong post, the two commenters who corrected her (one still reads her drafts), three months inside a four-researcher evals org with weekly calls to the external lab.
    • No EAG, no BlueDot: by the time she looked at the syllabus she had read most of it.
    • Her community is about six people she writes to, which she calls thinner than she would like.
    Her answer, as given

    Modestly, for about two and a half years, and mostly through the work rather than events. The LessWrong post was my first real contact; the two commenters who corrected me were the first people in the space I argued with properly, and one of them still reads my drafts. The internship put me inside a small evals org for three months, four researchers and me, with weekly calls with the external lab whose scaffold we were checking, which is where I learned how much of an evaluation result is infrastructure. I have not been to an EAG and I did not do BlueDot; by the time I looked at the syllabus I had read most of it. My community is about six people I write to, and that is thinner than I would like.

How much they care

  1. 10 · the question

    How much do you think about the space?

    • Turns ordinary results into threat-model questions: would this survive distribution shift, could the model tell it was being tested, are we measuring the scaffold rather than the behaviour.
    • The view that sits uneasily with people around her: SAEs are oversold as a safety-case input. Feature quality, causal relevance and coverage are unresolved, and her own project found the causal part weak even where labels looked clean.
    • Wants to build sandbagging evaluations that still tell you something when the model knows it is being evaluated. Switches off in the evenings.
    Her answer, as given

    A lot, in a particular way. Mostly I turn ordinary results into threat-model questions: would this survive distribution shift, could the model tell it was being tested, are we measuring the scaffold rather than the behaviour. That last one is the habit the internship left me with. The thing I have come to think that sits uneasily with people around me is that sparse autoencoders are oversold as a safety-case input. They give you descriptions, but feature quality, causal relevance and coverage are all unresolved, and my own project found the causal part weak even where the labels looked clean. What I want to do is build sandbagging evaluations that still tell you something when the model knows it is being evaluated. I do switch off in the evenings; I think that keeps my judgment better.

  2. 11 · verifying a claim

    In your application you describe finding a bug in an external lab's evaluation scaffold from a two-page aggregate report. What in those two pages told you it was the scaffold and not the model?

    • Accuracy went from 63 percent with no framing to 41 percent when the prompt made the evaluation clear. The lab read that as sandbagging.
    • Page two split outcomes three ways. The wrong-answer rate was flat at about 34 percent in both conditions; the whole drop was no-answer episodes going from 3 percent to 25.
    • Deliberate underperformance gives plausible wrong answers, not silence. Silence that scales with prompt length is a budget. She said 70 percent truncation, and it was: the framing pushed long episodes over the token cap and the parser scored a missing answer block as a failure.
    Her answer, as given

    The shape of the drop. The report had accuracy by condition: 63 percent when the model was told nothing, 41 percent when the prompt made it clear the run was an evaluation. That looked like sandbagging and the lab read it that way. But the second page split outcomes into correct, wrong, and no final answer, and the wrong-answer rate was flat at about 34 percent in both. The whole drop was no-answer episodes going from 3 percent to 25. A model underperforming on purpose gives you plausible wrong answers, not silence. Silence that scales with prompt length is a budget. I said I was maybe 70 percent it was truncation, and it was: the evaluation framing pushed the longest episodes over the token cap, the parser saw no answer block, and the scaffold scored that as a failure.

The person

  1. 12 · the question

    What is the most impressive or hardest thing you have actually done, in any field?

    • The SAE project, because she ran every part alone: sparse autoencoders on the residual stream of a roughly 400 million parameter model, at layer 7.
    • Reproduced the paper's headline that features get clean automated labels, then asked what the paper did not: clamp a feature, does the output move the way its label predicts.
    • Result: about 21 percent of SAE features, against 9 percent for PCA directions and 4 percent for random. Two commenters showed the 21 was inflated because empty-label features had been dropped as if missing at random. Counted as failures it is 17, published as a dated correction.
    Her answer, as given

    The SAE project, because I ran every part of it alone and it has the most experiments. From October 2023 to May 2024 I trained sparse autoencoders on the residual stream of a small open model, about 400 million parameters, at layer 7, reproduced the paper's headline that the features get clean automated labels, and then asked the question the paper did not: if you clamp a feature, does the output move the way its label predicts. The headline was that it does for about 21 percent of SAE features against 9 percent for PCA directions and 4 percent for random ones. Two commenters showed my 21 was inflated, because I had dropped features whose labels came back empty as if they were missing at random. Counted as failures, it is 17. I published that as a dated correction.

  2. 13 · a follow-up

    Beyond the comparison against PCA and random directions, what other ablations did you run on that extension?

    • Sparsity penalty over three values: sparser gave cleaner labels but fewer features that did anything when clamped.
    • Layers 6 and 8 as well as 7: the success rate roughly halved at 6 and matched at 8.
    • Two clamp strengths: the gentler one fell to about 12 percent but the ordering against the baselines held.
    • Selecting by activation frequency instead of label score, which is what forced the frequency matching. Steering on a different prompt distribution failed outright.
    Her answer, as given

    Four that mattered. I swept the sparsity penalty over three values; the sparser dictionary gave cleaner labels but fewer features that did anything when clamped, so label quality and causal use pulled in opposite directions. I ran the intervention at layers 6 and 8 as well as 7; at 6 the success rate roughly halved, at 8 it matched 7, so it is not a one-layer accident but it is not uniform either. I tried two clamp strengths, because I worried I was just breaking the model; at the gentler one success fell to about 12 percent but the ordering against the baselines held. And I tried selecting features by activation frequency instead of label score, which is what forced the frequency matching. What failed outright was steering on a different prompt distribution; nothing reproduced there.

  3. 14 · a follow-up

    Stay with the frequency matching. What did the unmatched comparison actually show, in numbers?

    • Almost nothing. The top 200 features by label score gave about 8 percent success, barely above PCA at 9.
    • Those features fired on fewer than one token in ten thousand. The labeller loved them because their top activations were five near-identical contexts. The unfiltered comparison measured rarity, not causal relevance.
    • Matched to the PCA directions' frequency band, one token in a hundred to one in ten, about 300 of the 2,000 features. That gave the 21, later 17, plus or minus four points.
    Her answer, as given

    Almost nothing, which was the surprise. Clamp the top 200 SAE features by label score and success was about 8 percent, barely above PCA at 9. Those features fired on fewer than one token in ten thousand, and clamping something that rare barely moves an output on any prompt you can test; the labeller loved them because their top activations were five near-identical contexts. So the unfiltered comparison was measuring rarity, not causal relevance. After seeing that, I matched: SAE features in the PCA directions' frequency band, about one token in a hundred to one in ten, roughly 300 of the 2,000, and that gave the 21, later 17. The cost is a wider interval, plus or minus four points on 300 features, which is in the post.

  4. 15 · the push

    Why did you match on frequency rather than report the unfiltered comparison as the result?

    • The unfiltered comparison did not answer the question. You cannot test whether a feature is used as its label says on a feature that never fires on the evaluation prompts; that is a null by construction.
    • PCA directions are dense, so dense baselines against rare features compares two different things. Matching put them on the same footing.
    • She reported the unfiltered number in the ablations section with the reason it is uninformative, and put the matched number in the title. Calls that a choice.
    Her answer, as given

    Because the unfiltered comparison was not answering the question. The question was whether a labelled SAE feature is used by the model in the way the label says, and you cannot test that on a feature that never fires on the prompts you evaluate on; you get a null by construction, not a finding about the model. PCA directions are dense, so comparing dense baselines against rare features is comparing two different things. Matching put the two on the same footing. I did report the unfiltered number in the post, in the ablations section, with the reason it is uninformative. But I chose to put the matched number in the title, and that was a choice.

  5. 16 · the push

    One thing that would worry me is that the matching rule came after the unfiltered comparison had come out flat, so the 17 percent is the number from the comparison you chose after looking. What do you make of that, looking back?

    • Concedes the specific point: the rule came after the flat result, and the band was chosen by eye with the outcome in view. A reader should discount the 17 by a few points, not to zero.
    • Pushes back on the idea that matching flattered the result: 17 percent under the fairest test she could build is not a vindication of the paper. A flattering number would have been label quality, which was excellent.
    • The forking-paths problem is real. The conclusion survives it.
    Her answer, as given

    That is fair on the specific point. The rule was chosen after I had seen the flat result, and the band I matched on was chosen by me, by eye, with the outcome in view. Any reader is entitled to discount the 17 for that, and I would put the honest discount at a few points rather than zero, because the confound's direction is a fact about the model anyone rerunning it will find. Where I push back is the idea that matching flattered the result. The paper's claim was that these features are causally meaningful, and 17 percent under the fairest test I could build is not a vindication of that; if I had wanted a flattering number I would have reported label quality, which was excellent. The forking-paths problem is real. The conclusion survives it.

  6. 17 · the push

    So if you had it again, what would you actually do differently?

    • Write the selection rule down before running a single intervention: the frequency band, the feature count, and that empty labels count as failures. An afternoon's work that removes both the correction and the forking-paths concern.
    • Report by frequency band rather than one headline, so rare features show as a row with a wide interval instead of vanishing.
    • Use a different model for the judge than for the labeller, and have someone blind to the labels rate 60 interventions by hand. Her own hand check was unblinded and she would not accept that from someone else now.
    • Would still publish on the forum first, because the commenters were the best review she got.
    Her answer, as given

    Three things, in order of cost to me. First, write the selection rule down before running a single intervention: the frequency band, the feature count, and that empty labels count as failures. That one change removes both the correction and the thing you just raised, and it would have cost me an afternoon. Second, report by frequency band rather than one headline, so the rare features appear as a row with a wide interval instead of vanishing. Third, use a different model for the judge than for the labeller, and have someone blind to the labels rate a sample of 60 interventions by hand; I did the hand check myself, unblinded, and I would not accept that from someone else now. I would still publish it on the forum first, because the commenters were the best review I got.

Full transcript and scoresTake it yourself