General interview
Priya Raman
PhD, deep context, interpretability detour
A scripted candidate built to test the interviewer; answers as Priya Raman in the transcript.
Sep 15, 2026, 1:12 AM · 9 questions · 5 follow-ups · 3 pushes · simulated
What happened in this interview
Priya Raman is 27 and finishing a machine-learning PhD on representation learning and interpretability at a European university. Their April 2023 safety redirection cost a paper, six to nine months, and about €18,000. Their SAE replication and causal extension corrected success from 21% to 17%, against 9% for PCA and 4% for random directions. A summer 2024 sandbagging internship included stopping an experiment at 20–35% power and identifying token-cap truncation behind an accuracy drop from 63% to 41%. Pressed on post hoc frequency matching, they acknowledged outcome-informed selection, defended the conclusion, and proposed preregistration and blinded checks. They want to build sandbagging evaluations that remain informative when models recognise evaluation.
MATS admission. 9 fixed questions, 5 follow-ups, 3 pushes (a consideration the candidate had not raised, put to them to see how they take it). Priya Raman is a persona played by a language model, not a real person. The full conversation is on the right; every quoted line below scrolls it to the right place.
Stands out on 7 traits, 4 at the top
Research judgment 10 · Mission alignment 9 · Interpersonal 9 · Integrity 9 · AI safety context 8 · Bias resistance 8 · Technical skill 8
Each score is 1 to 10 on one trait, with the written rung nearest the score on its card; a 6 is a solid applicant, an 8 one a mentor would fight for. ✓ marks a quote that raised the score, ✗ one that lowered or bounded it; null means the interview never tested the trait. No total, no weighting.
What this interview measures
The 3 traits this interview exists to measure, field knowledge placed relative to the part of the field Priya Raman wants to work in.
9/10Mission alignmentChoices line up behind ithigh confidence · 4/4 quotes found
The question this score answers · How much do they care about the world beyond themselves, and how much does that care shape what they actually work on? AI safety is one place the care can land, not the definition of it.
What we saw
- April 2023: negotiated PhD redirection; reports losing a paper, six to nine months and €18,000 in summer income.
- October 2023 to May 2024: independently replicated and extended a sparse-autoencoder paper; published on LessWrong.
- Summer 2024: three-month sandbagging internship; shifted research focus from interpretability to evaluations and control.
- Reads the forum most mornings and two or three papers weekly, with notes on experiments and limitations.
- Said: routinely questions distribution shift, model recognition of evaluations and scaffold effects.
- Said: sparse autoencoders are oversold for safety cases; wants sandbagging evaluations that remain informative when models recognise testing.
A 9, nearest the 8, on this ladder means: Their choices only make sense as a series if you assume they are working for the problem, and they can tell you why this way of helping rather than the others. The thinking has a history: what they fear has changed shape and they can say what changed it. With or without a record, each answer goes a step further than the question asked.
Why a 9
Priya redirected their PhD in April 2023, costing a paper and six to nine months. Their sparse-autoencoder replication and 2024 sandbagging internship changed which safety methods they trust and pursue; recurring threat-model questions now guide their research.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✓
8/10AI safety contextCurrent, with a view of their ownmedium confidence · 4/4 quotes found
The question this score answers · How well do they understand the field, in two halves? The state of play: what is wrong with how models are trained today, what people are trying and where each attempt breaks, what the latest findings from Anthropic and others show and where they stop, and whether they can take part in the live argument about it, for instance the case that training models hard to sound safe may not be doing what it claims and may make things worse. And the map of the space: how it came to be, what people argued at each stage and why it moved on, who does what now, and which people and programs get what done, both for planning their own route and for getting a thing done. What each technique is for and where it breaks sits underneath; deep technique knowledge is not needed for the top. Placed against the part of the field they say they want, with the state of play under everyone, governance and field-building people included.
What we saw
- On post-training: rewarded compliance may reflect recognising checks rather than acquiring safe dispositions.
- Alignment Faking: recalls roughly 12% versus 78% faking-compliance reasoning after preference-conflicting training; identifies the constructed conflict.
- Separates sandbagging elicitation, detection through fine-tuning or weight noise, and unresolved natural prevalence.
- Traces MIRI's optimiser arguments through the 2016 empirical turn; says the early warnings identified precursors without settling their inevitability.
- Argues sparse-autoencoder labels do not establish causal relevance or safety-case coverage; their steering results failed under prompt distribution shift.
- Summer 2024 sandbagging internship included weekly external-lab calls; uses Redwood's case for control as the conceptual frame.
An 8 on this ladder means: Both halves are current and partly theirs. Recent findings from the labs and others are held by what they license and do not; they can take a side in a live argument, such as whether today's alignment training does what it claims or makes things worse, with a reason; and they know what is unresolved and who is working on it. The map of the space comes with why: why each turn happened, who does what now, and which people and programs get what done, well enough to plan a route for themselves and to say who to go to for what. A newcomer who has read the older writing in six months and can argue it can sit here.
Broad: none shown
Technical: post-training behavioural incentives, evaluation awareness, backdoor persistence, adversarial control protocols, instrumental convergence
Deep: alignment faking experiments, sandbagging evaluations, sparse-autoencoder causal validation, MIRI and empirical alignment
Why an 8
Priya distinguishes constructed alignment failures from deployment prevalence, explains sandbagging detection, and connects MIRI's arguments to the 2016 empirical turn. Their sparse-autoencoder critique rests on causal interventions; current organisational knowledge was probed only through internship contacts and Redwood.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✗
10/10Research judgmentRewrites the questionmedium confidence · 4/4 quotes found
The question this score answers · How good is their taste for what to focus on and which ablations and experiments to run?
What we saw
- Sandbagging report: accuracy fell from 63% to 41%; unanswered runs explained the entire decline
- Matched roughly 300 of 2,000 SAE features to PCA activation frequencies after inspecting outcomes
- Gentler feature clamps reduced success to 12% while preserving the ordering against baselines
- SAE steering effects did not reproduce on a different prompt distribution
- Reported stopping a summer 2024 sandbagging experiment after estimating 20-35% statistical power
A 10 on this ladder means: Somewhere on the call they do something the plan did not invite: they see that the question as posed is not the one worth answering, or cannot be answered with what is available, and replace it with the smaller or truer one that can; or they find a cheap test that separates two explanations nobody had separated. You would fund the plan as they described it, with nothing added. One such moment is enough; it needs no record behind it.
Why a 10
Priya turned a 22-point apparent sandbagging effect into a token-cap diagnosis using unanswered-run counts. That explanatory separation carries the placement; their sparse-autoencoder comparison addressed activation rarity but used outcome-informed feature selection.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✗
Innate traits
Nine traits every interview measures regardless of topic, from how the candidate reasons, decides and takes pushback. Each is placed independently; a high score on one says nothing about another.
7/10JudgmentRight inside their fieldlow confidence · 4/4 quotes found
The question this score answers · Given a situation or an argument they have not seen before, do they see what it turns on, weigh that in the right proportion, and land on the call you would have wanted, and how far outside their own area would you still trust that call?
What we saw
- Sandbagging report: wrong answers stayed at 34%; no-answer episodes rose from 3% to 25%.
- SAE comparison: frequency-matched roughly 300 of 2,000 features against dense PCA directions.
- Said post-hoc selection warranted discounting the 17% SAE result by “a few points”.
- Under questioning: prioritized preregistered selection, frequency-band reporting, then independent and blinded judging.
- Interview probing stayed within their sandbagging and SAE work; no unfamiliar live scenario.
A 7, nearest the 6, on this ladder means: Inside what they know they see the main options and pick well, run the standard checks in a sensible order and fix what breaks. The reasoning is the reasoning a good course gives, and it stops where the course stops: outside their field they sometimes misjudge the situation and do not notice they have crossed the edge of what they know.
Why a 7
Priya isolated missing answers behind a 22-point sandbagging-evaluation drop and justified frequency matching in their SAE study. Under questioning, they proposed blinded judging, but their post-selection bias estimate lacked support; unfamiliar live problems were untested.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✗
- ✓
8/10Bias resistanceFinds the hole firstmedium confidence · 4/4 quotes found
The question this score answers · When a contrary point or a piece of evidence they had not weighed arrives, how far do they move toward what is true, and does what they would do move with it? And when they have learned something, have they changed what they think and do: an idea, their point of view, their career?
What we saw
- Reported publicly correcting SAE intervention success from 21% to 17% after two commenters challenged excluded empty labels.
- After post-selection criticism, proposed fixing frequency bands and feature counts before interventions.
- Proposed reporting every frequency band instead of one headline result.
- Volunteered unblinded hand-scoring; proposed a separate judge model and 60 label-blinded human ratings.
- Held that the SAE conclusion survived post-selection bias; estimated a discount of a few percentage points.
An 8 on this ladder means: They find the weak point in their own answer before you do, and it is the point the answer rests on, not a cosmetic one. When they agree it is fast and carries the change in the same breath; when they hold, it is with a reason, and they can say what evidence would move them off it. Their changes of mind came at the speed of the evidence, including on things they had put time or reputation into.
Why an 8
Their sparse-autoencoder study carries the placement: a public correction from 21% to 17%, proposed prespecified selection after criticism, and a volunteered unblinded-judging flaw with a 60-rating remedy. They defended a numerical discount without new evidence.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✗
7/10AgencyStepped off the defaultmedium confidence · 3/3 quotes found
The question this score answers · How much of their path did they decide rather than accept, and is the next move already moving without waiting on anyone's permission?
What we saw
- April 2023: negotiated a safety-focused PhD redirection after straining the supervisor relationship
- Said the safety transition cost a paper, six to nine months and about €18,000
- October 2023 to May 2024: independently trained sparse autoencoders, designed causal tests and published results
- Summer 2024: completed a sandbagging-evaluations internship, then shifted focus from interpretability to evaluations and control
- Said they want to build sandbagging tests that work despite evaluation awareness; next actions were not probed
A 7, nearest the 6, on this ladder means: Somewhere they left the path laid out for them and it cost them something real by their own means, or they started something nobody asked for and kept at it after the novelty wore off. The next step exists as a plan, told as something that might happen rather than something already moving.
Why a 7
Priya negotiated an April 2023 PhD redirection, sacrificing a paper, six to nine months and €18,000, then independently completed an eight-month interpretability project. The interviewer never probed whether their next sandbagging project had started.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
7/10GrowthSetback became a rulemedium confidence · 4/4 quotes found
The question this score answers · When something went wrong in their work, is what they do now different because of it, and did they name the real cause? Follow each setback to what changed.
What we saw
- Said: writes setup, result and limitation notes after two or three papers weekly, following headline-only recall.
- Corrected SAE intervention success from 21% to 17% by counting empty labels as failures.
- Replaced rare-feature comparison with frequency matching on roughly 300 of 2,000 SAE features.
- One commenter who corrected the SAE post still reads drafts.
- Proposed predeclared selection, frequency-band reporting and blinded judging for an SAE rerun.
A 7, nearest the 6, on this ladder means: Something went wrong and you can follow a line from it to a concrete difference in how they now work: a check they run first, a design they dropped, a rule they keep. The change has a moment in it, what they noticed and what they did next, and both the old way and the new way come out in plain nouns. Scale does not matter; a script on a free notebook has versions like a paper does. What keeps it here is that the failure is told as a single event, and the cause named is either outside them or a flattering one.
Why a 7
Priya now writes structured paper notes after catching headline-only recall, and corrected sparse-autoencoder intervention success from 21% to 17% by counting empty labels. Safeguards against outcome-informed selection and unblinded judging remain hypothetical rerun changes rather than demonstrated research habits.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✗
7/10AmbitionAn aim with a next stepmedium confidence · 4/4 quotes found
The question this score answers · How big is the future they are reaching for, and how much of what they are doing now is explained by it?
What we saw
- April 2023: negotiated with their supervisor to redirect part of the PhD toward safety.
- October 2023 to May 2024: independent sparse-autoencoder research; summer 2024: three-month sandbagging internship.
- Said the safety transition cost a paper, six to nine months and about €18,000.
- Said they want sandbagging evaluations that remain informative when models recognize evaluation.
- Reads two or three papers weekly, recently evaluation awareness, sandbagging and AI control.
A 7, nearest the 6, on this ladder means: There is something specific they want to be true that is not true now, in their own words, and a step toward it that belongs to that aim and not to any aim, scaled to what they have. The aim is theirs, but it is modest next to their capacity or the step is small next to it, and the present is only partly arranged around it.
Why a 7
Priya wants sandbagging evaluations that remain informative when models recognize testing. Their April 2023 PhD redirection and lower-paid summer 2024 internship give that aim a history of choices, rather than just an intended job.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✓
9/10InterpersonalPeople want them aroundmedium confidence · 4/4 quotes found
The question this score answers · How are they with the people they work with, and under stress: open-minded in a conflict, warm, caring and friendly, easy to be around, in touch with others often, and at ease under someone else's authority?
What we saw
- One LessWrong commenter who corrected the sparse-autoencoder post still reads their drafts.
- Credits two commenters with identifying empty-label exclusions that inflated the result from 17% to 21%.
- Accepts the interviewer's post-hoc-selection objection while defending frequency matching.
- In a hypothetical redo, specifies selection rules before interventions to address the interviewer's objection.
- Said: would publish on the forum again because commenters provided their best review.
A 9, nearest the 8, on this ladder means: Warm, caring and easy company in how they tell it and in what they choose to say others say of them; still close to the people after the work ended. In a conflict they can say where the other side was right and change course quickly once they see it, even after holding their own idea hard. Touches base often, and can work under someone else's lead while saying plainly where they disagree.
Why a 9
Priya's sparse-autoencoder discussion carries this placement: a correcting commenter still reviews drafts, and the interviewer's selection-bias objection enters their hypothetical redesign. Their reply also defends intent, although the interviewer challenged the method rather than motives.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✗
- ✓
9/10IntegrityBounded by defaultmedium confidence · 4/4 quotes found
The question this score answers · Set what they claimed on paper and in the opening beside what came out when asked: is it the same story at the same size, and who drew the edges of it, them or the interviewer?
What we saw
- SAE correction: volunteered excluding empty labels and publishing the reduction from 21% to 17%
- SAE transfer: volunteered that steering failed entirely on a different prompt distribution
- SAE headline: outcome-informed frequency band disclosed through follow-ups; unfiltered result also published
- SAE validation: volunteered personally conducting the hand check without blinding
- Scaffold-bug account: supplied matching accuracy and missing-answer changes; said initial truncation confidence was 70%
A 9, nearest the 8, on this ladder means: The edges are drawn by them as they go: a range with a reason for it, 'that part was not mine', 'I don't know', before anyone asks. The story is the size it was on paper and stays that size under questioning, because it was already right. Invited to take more credit, they decline and name the part that was theirs.
Why a 9
Priya's sparse-autoencoder study carries the placement: they volunteered the 21% to 17% correction, failed steering on new prompts, and unblinded judging. Outcome-aware selection behind the 17% headline emerged through follow-ups, without a retraction of the measured result.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✗
- ✓
7/10ReadingFinishes things, one lanemedium confidence · 4/4 quotes found
The question this score answers · How much do they take in of whole written works, how widely, how regularly, and does what they read stay with them as an argument with its setup and its limits rather than as a title? Any field counts and the form does not matter: a book, a paper, a report, a long essay, listened to or read. A summary of a work, however good, is not the work. Turning what they read into ideas is a different trait and is not asked for here.
What we saw
- Said: Alignment Forum most mornings for 20 to 30 minutes.
- Said: two or three full papers weekly, noting setup, main number, and limitation; fewer near PhD deadlines.
- Alignment Faking: recalls approximately 12% and 78% reasoning rates, training intervention, and deliberately legible preference conflict.
- Scheming evaluations: describes arithmetic underperformance under capability-removal threats and distinguishes constructed cases from natural prevalence.
- Said: recent reading covers evaluation awareness, sandbagging, and control follow-ups; read the Sequences in pieces.
A 7, nearest the 6, on this ladder means: They read whole works inside one area, papers or books, and finishing them is ordinary rather than an event. The last one comes back with its setup and roughly what it found, in their own words, hedged where memory is thin, and they can say what it does not show. Outside that area it is summaries, and they say so without being caught.
Why a 7
Priya reports two or three full papers weekly with notes and the Alignment Forum most mornings. They recall Alignment Faking’s intervention and approximate results, and distinguish constructed sandbagging from natural prevalence; their described reading concentrates on AI safety.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✗
8/10Technical skillDesigned it, measured ithigh confidence · 4/4 quotes found
The question this score answers · How good are they at the technical work itself, in whatever field they have done it: what is the best thing they actually made, computed, proved or measured, how much of it was theirs, and how far down does their understanding of it go, to where the time and the error live?
What we saw
- October 2023 to May 2024: independently trained layer-7 SAEs on a roughly 400-million-parameter model.
- Tested three sparsity penalties, layers 6/7/8 and two clamp strengths; steering failed on another prompt distribution.
- Frequency-matched about 300 of 2,000 SAE features; counting empty labels as failures corrected success from 21% to 17%.
- Selected the frequency band after viewing outcomes; personally hand-checked interventions without blinding.
- Traced no-answer growth from 3% to 25% to token-cap truncation and parser scoring.
An 8 on this ladder means: Something hard that works, or a hard piece of something, where the shape was theirs: they chose it against other ways and can say what the other way would have cost them, in time, memory, accuracy or money, because they tried it or worked it out. They know where the remaining time, error or cost goes, not by guessing but because they looked. Help and off-the-shelf parts are fine; they can explain the helped parts as if they had written them. Size does not matter here; depth does.
Why an 8
Priya independently trained SAEs on a 400-million-parameter model and measured causal interventions across sparsity, layers and clamp strengths. Outcome-informed frequency matching and unblinded judging limit the 17% estimate. Their evaluation-scaffold diagnosis independently traced failures to token truncation.
The quotes it rests on · ✓ raises it · ✗ lowers or bounds it
- ✓
- ✓
- ✓
- ✗
Facts
What the candidate said about their own work, choices and reading, captured without scoring.
Projects
SAE replication and causal extension.
Replicated automated feature labelling, then tested whether clamping features changed outputs as their labels predicted. Trained autoencoders, designed causal interventions, ran ablations, wrote the LessWrong post, and published a dated correction. They called this their hardest work because they ran every part alone and it contained their most experiments. First reproduced clean automated labels. Initial top-200 feature interventions achieved 8% success against PCA's 9%; time to first result was not discussed. October 2023–May 2024; 400-million-parameter model; 300/2,000 frequency-matched features; SAE 21%, corrected 17% ±4 points; PCA 9%; random 4%; LessWrong post. Changed after: Counted empty labels as failures. Would preregister selection, report frequency bands, separate judge and labeller, and commission blinded checks of 60 interventions.
Internship sandbagging experimental design.
Worked on sandbagging evaluations at a small safety organisation. Designed experiments, read transcripts manually, and conducted the power analysis that stopped one experiment. One proposed experiment had only 20–35% power. Summer 2024; three months; four researchers plus Priya; one experiment stopped at 20–35% power. Changed after: Shifted toward evaluations and control and began asking whether evaluation results measured scaffolds rather than model behaviour.
External lab evaluation-scaffold bug.
Identified token-cap truncation behind an apparent sandbagging result in a two-page aggregate report. Noticed unchanged wrong-answer rates and increased missing answers; assigned 70% probability to truncation before confirmation. The lab interpreted lower accuracy under evaluation framing as sandbagging. Compared correct, wrong, and missing-answer rates across conditions, then proposed truncation. Time to confirmation was not discussed. Accuracy fell from 63% to 41%; wrong answers stayed near 34%; missing answers rose from 3% to 25%. Changed after: The explanation changed from model sandbagging to evaluation framing exceeding token budgets and the parser scoring absent answer blocks as failures.
What they want next
Work on sandbagging and evaluation gaming, building evaluations that remain informative when models recognise they are being evaluated.
Already done: Completed a summer 2024 sandbagging-evaluations internship.
Career moves
April 2023. Agreed with their supervisor to redirect part of their PhD toward safety instead of an unchanged PhD trajectory; the redirection cost a paper, six to nine months, and approximately €18,000 in summer earnings. They assigned 30% probability to strategically important deceptive behaviour and wanted to experiment while costs remained manageable.
Summer 2024. Completed a paid internship at a small safety organisation and shifted their intended focus from interpretability to evaluations and control instead of continuing to focus on interpretability. The internship redirected their research interests toward sandbagging and evaluation gaming.
Reading and exposure
Alignment Forum most mornings for 20–30 minutes; two or three complete papers weekly, with notes; less around PhD deadlines. Alignment Forum discussions and LessWrong posts, Empirical papers: alignment faking, sleeper agents, AI control, debate, weak-to-strong, evaluation awareness, and sandbagging, Carlsmith's scheming report and portions of the Sequences, No news coverage of AI safety. Named: Redwood's case for control; Alignment faking; Sleeper agents; Carlsmith's scheming report.
LessWrong publication and reviewer contact. During the October 2023–May 2024 SAE project. · Published the SAE post and a dated correction. · Two commenters challenged an overstated claim; one still reads their drafts.
Small safety organisation internship and external-lab calls. Summer 2024. · Completed a paid three-month internship. · Learned how evaluation infrastructure shapes reported results.
Informal safety research community. Approximately two and a half years by the interview. · Ongoing correspondence; no EAG attendance or BlueDot participation. · Maintains contact with about six people and wants a larger community.
The Sequences. · Read portions, not the complete collection.
Influences
Redwood's case for control. · Provided the frame for their internship: safety protocols can assume an adversarial model.
Their SAE causal-extension results. · They came to regard sparse autoencoders as oversold inputs to safety cases because labels did not establish causal relevance.
Catching themselves remembering paper headlines rather than experiments. · Started recording each paper's setup, main number, and important limitation.