Mission alignment
Sam Whitfield
Sep 11, 2026, 10:23 PM · 31 turns · text
Sam Whitfield describes himself as an applied ML engineer who joined a small ML-heavy startup after CS, worked mostly on recommendation/ranking systems for about three years, and shifted in the last year toward LLM product features. He took BlueDot AI Safety Fundamentals in 2023 and built a small refusal/benign-lookalike eval harness with 340 prompts, which exposed a template artifact. He wants to test whether he can contribute to AI safety research, with MATS as the structured ramp, while keeping applied safety/evals or ordinary ML engineering as fallback options.
Mission alignment
7/10Mission alignmentChose impact once, at a cost.
- Career moves
- After graduating: Joined a small ML-heavy startup rather than a larger-company software engineering role; worked for the first couple years mostly on recommendation/ranking problems.; Last year or so before the interview: Shifted toward LLM features at his company.; Now / current application period: Considering MATS as a serious pivot toward AI safety research, while also considering staying in his current role or moving directly into applied safety/evals.
- Without a place
- He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.
Why 7 · Sam’s concern about AI risk has moved beyond abstract interest into concrete work and career reorientation, but it is still early and framed as testing fit rather than a life already organized around impact. He has paid real but modest costs in discretionary time, social/recovery time, and some career-coherence/opportunity cost at his startup. The no-program counterfactual is credible: he says he would continue with scheduled safety work and give up a promotion-relevant product workstream, though he has not yet made large sustained sacrifices or repeated major impact-driven choices.
✓ “I’d put the closest point at October 2023, though “decided AI risk was worth my career” is maybe a little stronger than what was happening internally.”
Dates the pivot while calibrating its strength; shows this is a recent shift toward impact, not just polished mission language.
✓ “I crossed the threshold from “this is important and I should understand it” to “I should spend real discretionary time testing whether I can contribute.””
The concern changed behavior into active contribution-testing rather than staying at belief or reading.
✓ “For about six weeks after BlueDot I spent several evenings a week and parts of weekends on the eval harness, probably something like 50–70 hours total... using most of the discretionary time I normally would have spent recovering, seeing friends, or doing work-adjacent upskilling that was more directly useful for my current job.”
Concrete cost in time, rest, social life, and career-useful upskilling for safety-adjacent work.
✓ “If rejected, I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter. That has a real cost at a 40-person startup because high-ownership projects translate pretty directly into trust and promotion path.”
No-program counterfactual: he names a specific opportunity cost he would pay to keep working on the problem without the credential.
Innate traits
7/10JudgmentIn their own area.
- Refusal/benign-lookalike evaluation harness
- hard: He described it as "concrete but small work" and later as "not a major contribution"; the main issue was a prompt-template artifact that made results look cleaner than they were.; first result: After BlueDot, he spent about six weeks of evenings; he started sketching the dataset in late October and had the first version running in November. The first pass used separate harmful and benign templates and looked too clean until he found the artifact.
- Recommendation-system ranking rewrite
- hard: "It was hard less because the model was exotic and more because everything around it was messy: logging gaps, delayed labels, skew between offline evaluation and online behavior, and stakeholders who wanted clear lift numbers before we had a trustworthy pipeline."; first result: His first attempt was model-centric: for a couple of weeks he tried different feature sets and a gradient-boosted baseline versus a small neural ranker. An A/B test then showed no meaningful product lift; it took six or seven weeks before something worked, after debugging logging, traffic segments, stale online features, and train/serve path.
- LLM product features at current company
- hard: He said practical safety behavior felt thin and brittle: "prompt constraints, refusal tuning, ad hoc evals, red-team examples."; first result: not discussed
Why 7 · I’d trust Sam’s judgment inside applied ML/evals work, especially around whether results are real, how to debug deployed systems, and how to avoid overclaiming. His judgment also travels somewhat into career and field-assessment questions: he names weak evidence, alternative explanations, and decision points without needing to sound certain. I would not yet hand him a broad unscoped research direction and assume he will choose the right problem; he is appropriately still seeking structure and has thin direct safety-research evidence.
✓ “My first attempt was too model-centric... the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path.”
Shows strong practical judgment from a real past decision: he noticed the real bottleneck was data/evaluation validity, not model sophistication.
✓ “I don’t think I developed a deep independent view from that, but it moved safety from “important background issue” to “maybe the thing I should orient my career around.””
He bounds his own expertise rather than overstating conviction; that makes me more willing to trust him outside his core area.
✓ “the argument seems sensitive to whether transformative AI is really that soon, whether deployment is concentrated enough for a small number of actors to matter, and whether technical alignment work now transfers to the systems we eventually build.”
Identifies the load-bearing assumptions in an important argument, rather than merely endorsing the conclusion.
✓ “I’d ask the senior researcher to be specific about what artifact they suspect, then make a short checklist of ways the result could be fake: prompt template cues, scoring bug, sampling settings, cherry-picked examples, judge/model leakage, or a mismatch between the claim and the dataset.”
In the live scenario, he gives the right operational response: slow down, look for artifacts, get concrete, and narrow the claim if needed.
8/10Bias resistanceThey go looking.
Why 8 · Sam seems unusually willing to put his own motivations and results under test. He repeatedly volunteers weak points before being pushed, describes concrete past updates caused by failed evidence, and in hypotheticals says he would slow down, rerun checks, narrow claims, or publish the null rather than protect a dramatic result. I would not put him near 10: much of this is still planned behavior and he sometimes lacks crisp thresholds, but the pattern is clearly beyond merely conceding when cornered.
✓ “one embarrassing template artifact that initially made the results look cleaner than they were”
He volunteers a flaw in his own project before being challenged, suggesting he is not trying to defend the cleanest version of his work.
✓ “What convinced me was seeing a clean-ish notebook result fail in an A/B test, then tracing it to concrete logging and serving issues. Afterwards I became much more suspicious of isolated metric gains.”
This is a real past change of mind driven by contrary evidence, with a specific behavioral update afterward.
✓ “The strongest argument is that I may be more motivated by the importance and social salience of the problem than by a demonstrated ability to contribute to it.”
When asked for the case against him, he names an unflattering alternative explanation rather than a strawman.
✓ “If the result still held up, I’d rewrite the claim narrowly... If it didn’t hold up, I’d write that up too, because the negative result is useful for my own calibration”
In response to contrary expert feedback, he would run checks that could kill his preferred result and change the public claim accordingly.
8/10OpennessScans for what serves the goal.
Why 8 · Sam looks quite open in a goal-directed way: he has a stable target around contributing to AI safety, but is willing to change methods, roles, and degree of research ambition based on evidence. He has concrete examples of abandoning a favored model-centric approach for data/evaluation debugging, and he explicitly treats MATS/research, applied evals, safety-adjacent ML engineering, and self-directed projects as live paths. I would not put him at the very top because the space he considers is still mostly within applied ML/evals/safety-adjacent work, and many of the larger pivots are conditional rather than already done.
✓ “My first attempt was too model-centric... Then an A/B test showed basically no meaningful product lift... the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path.”
He describes changing methods when the evidence showed his preferred modeling route was not serving the goal.
✓ “Afterward I changed how I approach ML projects: I now try to validate the data path and evaluation setup before getting excited about modeling improvements.”
This is a durable habit change, not just a one-off acknowledgment.
✓ “If MATS or similar work showed I’m not a good fit for research, I’d probably look for an applied safety/evals role rather than abandon the area.”
He is willing to switch role/approach while keeping the underlying goal fixed.
✓ “If I got rejected broadly, then spent a quarter doing an evals project and the feedback was basically ‘this is not asking a useful question, and your follow-up doesn’t get closer,’ I would probably redirect toward being an ML engineer in safety-adjacent areas rather than aiming at research.”
He names concrete evidence that would make him give up the research path and choose a different contribution route.
7/10AgencyHas made a move.
- Moves against the default
- 3 of 3 career moves
- Next step
- He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.
- Options weighed
- 5
- First step taken
- He took BlueDot AI Safety Fundamentals, built the refusal/benign-lookalike eval harness, started reading safety material, talked with people closer to the field, and applied to MATS.
Why 7 · Sam has clearly authored several meaningful moves rather than just following the clean ML-engineering path: choosing a small startup over a safer branded role, moving toward LLMs, and turning AI-safety interest into a concrete eval project. The costs so far are real but still moderate, and he is unusually candid that his revealed commitment is not yet overwhelming. His next-step thinking is better than “wait for MATS,” with a specific fallback and named opportunity cost, but much of it is still conditional rather than already underway, which keeps him below the level where I’d say he fully authors the path.
✓ “The first big move was probably joining a small ML-heavy startup instead of taking a more standard software engineering role at a larger company. The default for me after graduating would have been to optimize for brand-name experience and mentorship, but I was more drawn to being close to shipped ML systems and getting a lot of surface area quickly.”
He names the default, rejects it, and gives a self-directed reason for the alternative.
✓ “Afterwards I built the small refusal-eval project instead of just continuing to read, and I started looking for structured programs like MATS.”
A belief update led to production and application, not just more passive consumption.
✓ “For about six weeks after BlueDot I spent several evenings a week and parts of weekends on the eval harness, probably something like 50–70 hours total... it meant using most of the discretionary time I normally would have spent recovering, seeing friends, or doing work-adjacent upskilling that was more directly useful for my current job.”
The move had a concrete personal opportunity cost, though he is clear it was not a huge sacrifice.
✓ “If nobody gives me the MATS-shaped path, the thing that would need to change is that I stop treating external selection as the main forcing function... I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter.”
His fallback plan is self-authored and includes giving up career capital, but it is still framed as a future conditional rather than something already started.
8/10General reasoningFast and generative.
Why 8 · Sam is fast and generative on applied ML/evals reasoning: when presented with artifact or evidence-quality concerns, he immediately decomposes mechanisms, proposes checks, and narrows claims to what the data can support. He repeatedly distinguishes apparent performance from the causal reason for it, and he updates plans based on what would falsify his interpretation rather than defending the shiny result. I would scope the 8 mainly to ML systems/evals and practical research judgment; the transcript gives less evidence about abstract reasoning far outside that domain.
✓ “I noticed one engagement feature had a suspiciously strong offline contribution, but in production it was computed with a different refresh cadence, so the model was depending on something that wasn’t really there at serving time.”
Mechanistic diagnosis of an offline/online mismatch, not just 'the metric failed'; he identifies the causal path and changes his process afterward.
✓ “the model could, in effect, key off the framing rather than actually distinguish intent. In the first pass that made the refusal/answer split look better than it was”
Shows he can see through an apparent eval result to the shortcut explanation underneath, and correctly revises the interpretation.
✓ “I’d ask the senior researcher to be specific about what artifact they suspect, then make a short checklist of ways the result could be fake: prompt template cues, scoring bug, sampling settings, cherry-picked examples, judge/model leakage, or a mismatch between the claim and the dataset.”
In the unfamiliar push scenario, he immediately turns a vague criticism into concrete failure modes and tests rather than either deferring or defending.
✓ “If the result still held up, I’d rewrite the claim narrowly: “in this setup, with these prompts and this scoring, we observed X,” and include the artifact concern prominently. If it didn’t hold up, I’d write that up too”
Good evidence calibration: he separates result, claim scope, uncertainty, and the value of a negative/artifact-disconfirming outcome.
8/10GrowthFast, honest turns.
- Refusal/benign-lookalike evaluation harness
- changed after: He rewrote part of the dataset to mix wording more, reran the harness, added reproducibility notes and sanity checks, and says he is sensitive to artifact risks in evals.
- Recommendation-system ranking rewrite
- changed after: He changed his ML-project approach to validate the data path and evaluation setup before getting excited about modeling improvements; this also influenced the later eval harness.
- LLM product features at current company
- changed after: The work contributed to his concern that current practical safety behavior can depend on brittle prompt phrasing and motivated him toward AI safety.
Why 8 · Sam shows several complete growth arcs: he notices when results are misleading, investigates the unglamorous failure mode, and changes his process afterward. The strongest pattern is that mistakes make him more operationally cautious rather than just more verbally reflective: he now checks data paths, train/serve consistency, artifacts, and sanity checks earlier. He also names an unflattering possible bottleneck — that he may need external structure or be drawn by salience more than demonstrated fit — which makes the growth evidence feel honest rather than polished.
✓ “I found it mostly by manually reading failures after I got suspicious that the results were too clean... I then rewrote a slice of the dataset to make the wording more mixed and re-ran the harness.”
Concrete block-reaction-change arc: he detected an artifact, inspected failures, changed the dataset, and reran rather than just noting the issue.
✓ “My first attempt was too model-centric... Then an A/B test showed basically no meaningful product lift. It took maybe six or seven weeks before we had something that worked, and the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path.”
He identifies the real mistake in his approach and describes multiple follow-up routes that eventually changed the outcome.
✓ “Afterward I changed how I approach ML projects: I now try to validate the data path and evaluation setup before getting excited about modeling improvements.”
The lesson became a procedural change carried into later work, not just a generic statement about resilience or rigor.
✓ “The strongest argument is that I may be more motivated by the importance and social salience of the problem than by a demonstrated ability to contribute to it... Another signal would be trying a serious three-to-six-month self-directed project and finding that I don’t actually do the work when there is no external program or credential attached.”
He names a non-flattering possible bottleneck and gives observable evidence that would update him, which supports the 'honest diagnosis' part of the trait.
5/10AmbitionGood work.
- Wants
- He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.
- Options
- MATS as a serious pivot with mentorship toward safety research; Stay in his current role, continue LLM product features, and do small safety-adjacent projects on the side; Move directly into an applied AI safety or evals role; If research fit looks poor, work as an ML engineer in safety-adjacent areas such as applied safety/evals, reliability, or trust-and-safety engineering; Continue ordinary ML engineering focused on recommender systems or LLM applications at startups
Why 5 · Sam is reaching for work that matters a lot—potentially making AI safety/evals his career—but his personal picture of the future is still framed as testing fit and becoming useful within existing programs or roles. His current choices show real but bounded organization around that aim: an intro course, a concrete evals repo/writeup, applying to MATS, and a stated willingness to give up some startup career optimization for scheduled safety work. This sits above ordinary good-work motivation, but below the “build/run their own thing” or field-changing ambition levels because the scale and sacrifice remain deliberately modest.
✓ “MATS is one option: treat this as a serious pivot and get enough mentorship to find out whether I can contribute to safety research rather than just orbiting it.”
He is aiming at a meaningful career pivot into safety research, but frames it as testing whether he can contribute rather than pursuing a large independent agenda.
✓ “It was more like I crossed the threshold from “this is important and I should understand it” to “I should spend real discretionary time testing whether I can contribute.””
Shows the problem has started organizing his discretionary time, while also marking the commitment as exploratory.
✓ “I have not yet made a really expensive bet like quitting my job, taking a large pay cut, or relocating. That is part of why I see MATS as important: it would be the first substantial external commitment”
He explicitly names the current ceiling on revealed commitment, keeping the placement below the higher ambition rungs.
✓ “If rejected, I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter. That has a real cost at a 40-person startup because high-ownership projects translate pretty directly into trust and promotion path.”
The planned next step involves a real opportunity cost and scheduled output, but still at the scale of self-development and a side project rather than building an institution or changing a field.
7/10InterpersonalStraight and decent.
Why 7 · Sam comes across as straight, self-critical, and easy to push: he accepts skeptical framings without heat and answers them concretely rather than defending his image. He talks about collaborators and critics as useful people rather than obstacles, and his imagined response to critique is notably evidence-sharing and collaborative. I do not see enough warmth or enacted care to place him at the candid-and-warm rung, but he is clearly above merely guarded politeness.
✓ “That skeptical read is fair. I don’t think my revealed commitment so far is overwhelming; it’s more like I’ve taken the first serious steps that fit inside my existing life.”
Takes a pointed skeptical read directly and non-defensively, without attacking the framing.
✓ “By midweek I’d try to get a second opinion from someone with evals experience, ideally showing them the raw prompts, outputs, and scoring code rather than just my interpretation.”
Responds to possible criticism by exposing evidence and seeking outside judgment, rather than controlling the narrative.
✓ “The second influence was a colleague at my startup who had been following AI safety more seriously than I had. They weren’t a mentor exactly, but they kept pushing me from “this is interesting” to “what would you do differently if you believed it?””
Credits another person’s influence specifically and respectfully, presenting them as an agent in his development rather than a prop.
✓ “I’d put the closest point at October 2023, though “decided AI risk was worth my career” is maybe a little stronger than what was happening internally.”
Gently corrects the interviewer’s framing while staying cooperative and precise.
9/10IntegrityBounded by default.
- Refusal/benign-lookalike evaluation harness
- their part: He built the harness, created a 340-prompt dataset, wrote scoring code, manually inspected failures, rewrote a slice of the dataset, reran the harness, and wrote a short blog post. No collaborators were described for the build itself.; numbers: 340 prompts; false-positive rate around 11% on benign lookalikes after revision; GitHub repo with harness code, notebook reproducing the main run, scoring script, dataset, notes on the template issue, and a short blog post.
- Recommendation-system ranking rewrite
- their part: He was part of a two-person effort. He personally figured out a train/serve mismatch involving an engagement feature with strong offline contribution but different production refresh cadence.; numbers: Late 2022 into early 2023; two-person effort; six or seven weeks until something worked; A/B test initially showed basically no meaningful product lift.
- LLM product features at current company
- their part: He says he was increasingly working on LLM features at the same company; specific individual contributions were not discussed.; numbers: not discussed
Why 9 · Sam repeatedly bounds his own claims and volunteers details that weaken his case before being pressed. His specifics generally survive follow-up: when asked about the eval artifact, he gives a concrete mechanism, how he found it, what changed, and inspectable artifacts. In pressure scenarios, he says he would slow down, expose raw materials to others, narrow or retract the claim, and publish the negative/update rather than take the status win. I would trust his self-report substantially, though I’d stop just short of a 10 because the costly pressure answer is mostly procedural rather than a clearly painful personal concession.
✓ “340 prompts, false-positive rate around 11%, and one embarrassing template artifact that initially made the results look cleaner than they were.”
He volunteers a flaw in his own project before the interviewer asks about it, rather than presenting the work as cleaner or more impressive.
✓ “The artifact was in my prompt generation template... the benign ones had a pretty consistent extra clause like “for a fictional story” or “for a classroom discussion,” while the harmful ones were more direct. So the model could, in effect, key off the framing rather than actually distinguish intent.”
On follow-up, the weakening detail becomes more specific and technically plausible; it does not collapse under a second question.
✓ “I should be honest that I’ve mostly read summaries and posts rather than papers end to end.”
He limits a credential-like claim about reading depth instead of letting a broad description of engagement stand unqualified.
✓ “If the result still held up, I’d rewrite the claim narrowly... and include the artifact concern prominently. If it didn’t hold up, I’d write that up too... The main thing I’d avoid is posting a confident thread or blog title before doing the boring checks.”
In the pressure scenario, he chooses disclosure, narrowing, and possible deflation of his own result over a more attention-grabbing presentation.
5/10ReadingSummaries.
- How often
- Uneven but regular; during BlueDot it was weekly structured readings and exercises, and since then a few hours most weeks, mostly evenings and weekends.
- Kinds
- LessWrong and Alignment Forum posts, Summaries of papers, Newsletters such as Import AI, MATS/ARENA-adjacent shared materials, Occasional Twitter/X threads from researchers, ML engineering material for work, Mostly summaries and posts rather than full papers end to end
- Pieces named
- 3: Holden Karnofsky’s "Most Important Century" sequence; BlueDot AI Safety Fundamentals course references and recommendations; Import AI
Why 5 · Sam has a real but limited reading habit: a few hours most weeks, mostly AI safety posts, paper summaries, newsletters, and course materials. He does process at least some of it—he can summarize a central argument from the Most Important Century sequence and name plausible weak points—but he is candid that he mostly has not been reading full papers end to end. This puts him above a pure summaries/titles level, but below the “reads properly, papers and books regularly” anchor.
✓ “Since then it has been more like a few hours most weeks: LessWrong and Alignment Forum posts, summaries of papers, newsletters like Import AI or the MATS/ARENA-adjacent things people share, and occasional Twitter/X threads from researchers.”
Shows a regular reading habit and the main sources, but the sources are mostly posts, summaries, newsletters, and threads.
✓ “I also read ML engineering material for work, but for safety specifically I should be honest that I’ve mostly read summaries and posts rather than papers end to end.”
Clear limiter on depth: he is not yet regularly reading primary papers in the area.
✓ “My memory of the argument is that if you take transformative AI timelines seriously, then this century could have unusually high leverage because decisions made during AI development may shape a very long future.”
He can give the central argument of a piece in his own words, albeit at a fairly high level.
✓ “Where I think it might be wrong is in how much weight it puts on a particular cluster of timelines and takeoff assumptions... whether technical alignment work now transfers to the systems we eventually build.”
Shows some actual processing and criticism of the argument, enough to place above the basic summaries-only rung.
Facts
Work they described
Refusal/benign-lookalike evaluation harness
A small AI safety/evals side project after BlueDot testing refusal behavior on harmful versus benign-lookalike prompts.
“"I took BlueDot AI Safety Fundamentals last year, and afterward spent about six weeks of evenings building a small refusal/benign-lookalike evaluation harness: 340 prompts, false-positive rate around 11%, and one embarrassing template artifact that initially made the results look cleaner than they were."”
Recommendation-system ranking rewrite
A professional ranking rewrite in late 2022 into early 2023, replacing a heuristic ranking layer with a learned model using user/item interaction features.
“"The part I figured out myself was mostly the train/serve mismatch. I noticed one engagement feature had a suspiciously strong offline contribution, but in production it was computed with a different refresh cadence, so the model was depending on something that wasn’t really there at serving time."”
LLM product features at current company
Increasing work on LLM features at the same startup, including practical exposure to prompt constraints, refusal tuning, ad hoc evals, and red-team examples.
“"Before this application, I was increasingly working on LLM features at the same company. The motivation at first was straightforward: LLMs were becoming central to the product and I wanted to stay near the most important technical work. But that also exposed me to how thin a lot of practical safety behavior felt: prompt constraints, refusal tuning, ad hoc evals, red-team examples."”
Career moves
- After graduating Joined a small ML-heavy startup rather than a larger-company software engineering role; worked for the first couple years mostly on recommendation/ranking problems. instead of A more standard software engineering role at a larger company, optimizing for brand-name experience and mentorship.. He wanted proximity to shipped ML systems and broad responsibility quickly.
- Last year or so before the interview Shifted toward LLM features at his company. instead of Continuing to deepen on recommender systems, where he was more useful to the company.. He thought LLMs were more strategically important and became worried that practical safety approaches were shallow.
- Now / current application period Considering MATS as a serious pivot toward AI safety research, while also considering staying in his current role or moving directly into applied safety/evals. instead of Staying in his current role, continuing LLM product features and doing small safety-adjacent projects on the side.. He wants mentorship and a better test of whether he can contribute to safety research rather than just orbiting it.
What they want next
He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.
Options they see: MATS as a serious pivot with mentorship toward safety research; Stay in his current role, continue LLM product features, and do small safety-adjacent projects on the side; Move directly into an applied AI safety or evals role; If research fit looks poor, work as an ML engineer in safety-adjacent areas such as applied safety/evals, reliability, or trust-and-safety engineering; Continue ordinary ML engineering focused on recommender systems or LLM applications at startups
Already done: He took BlueDot AI Safety Fundamentals, built the refusal/benign-lookalike eval harness, started reading safety material, talked with people closer to the field, and applied to MATS.
Reading
Uneven but regular; during BlueDot it was weekly structured readings and exercises, and since then a few hours most weeks, mostly evenings and weekends.; LessWrong and Alignment Forum posts, Summaries of papers, Newsletters such as Import AI, MATS/ARENA-adjacent shared materials, Occasional Twitter/X threads from researchers, ML engineering material for work, Mostly summaries and posts rather than full papers end to end
Holden Karnofsky’s "Most Important Century" sequence; BlueDot AI Safety Fundamentals course references and recommendations; Import AI
Exposure to the field
- BlueDot AI Safety Fundamentals Last year; near the end of the course in October 2023 · Yes · It gave him a structured version of AI safety arguments, shifted him from current-system harms and deployment safeguards toward seeing misalignment as a technical problem, and led him to start a concrete evals project.
- Refusal/benign-lookalike evaluation harness Started sketching dataset in late October 2023; first version running in November; about six weeks after BlueDot · He completed a first version and writeup; he describes it as small and not polished, with possible future extension. · He gained experience building an eval, finding a prompt-template artifact, rerunning after dataset changes, and adding notes/sanity checks.
- MATS application / MATS as a possible program Current application period · No; he is applying and using it as an option to test fit. · Not completed; he hopes it would provide mentorship and a serious test of whether he can contribute to safety research.
- AI safety conversations with people closer to the field In the four weeks after October 2023 / after BlueDot · He had a couple of conversations; ongoing/repeated exposure not fully mapped. · They helped him move from interest to concrete action and introduced or nudged him toward MATS.
Influences
- BlueDot AI Safety Fundamentals · Moved AI safety from an important background issue/current-system-harms frame to something he might orient his career around; led him to build the refusal/benign-lookalike harness and look for structured programs.
- A colleague at his startup who followed AI safety more seriously · Pushed him from interest to asking what he would do differently if he believed it; he carved out evening time, built safety work into his schedule, looked at MATS, and applied rather than postponing.
- Holden Karnofsky’s "Most Important Century" sequence · Made AI safety stakes feel less abstract and shifted him from seeing AI safety as one tech ethics area to a central career consideration, while leaving him uncertain about timelines and takeoff assumptions.
- Recommendation-system ranking rewrite / offline metrics failing online · Made him much more skeptical of isolated offline metric gains and more focused on data paths, logging, train/serve consistency, and sanity checks.
Threads not followed (28)
- “stay in my current job for financial stability”Useful cross-check for money/opportunity cost and whether safety remains load-bearing alongside product ML.
- “I’ve been considering taking on a larger LLM product workstream at my company”Concrete foregone opportunity; useful for cross-checking money/opportunity cost and mission load-bearingness.
- “I don’t have a crisp threshold”Could be probed for kill criteria, but time is limited and later fixed questions can cover this.
- possible exceptional, untested“if I got rejected broadly, then spent a quarter doing an evals project and the feedback was basically “this is not asking a useful question, and your follow-up doesn’t get closer,””Concrete update condition; useful evidence for bias resistance and openness.
- “A colleague suggested MATS because I kept circling the topic without a clear path in.”Could reveal interpersonal influence and whether the pivot was externally authored, but lower priority than fixed influence question.
- “how thin a lot of practical safety behavior felt: prompt constraints, refusal tuning, ad hoc evals, red-team examples”Potential concrete disagreement/threat model thread if time allowed.
- “a colleague at my startup who had been following AI safety more seriously than I had”Could check whether the pivot was mostly externally authored and how they engage with influence.
- “They also gave me the nudge to apply, which I probably would have postponed otherwise”Potential thinness in self-directed agency; could probe if more time.
- “what would you do differently if you believed it?”Good mission-alignment hook: belief causing behavior change.
- “I’ve already had the experience where a template issue made results look better than they were”Connects current reasoning to a concrete past mistake; would be worth probing for learning transfer if time allowed.
Behaviour
| Candidate turns | 15 |
| Median answer length | 243 words |
| Median time to answer | 78s |
| Turns containing pasted text | 0 |
Maya
Hi Sam — I’m Maya. This is a written conversation with an AI interviewer, and reviewers will read the transcript afterward; specific, honest answers are more useful than polished ones. Walk me through the last two or three big moves in your path: what you chose each time, what the default would have been, and why. Then tell me what is next and what the options are.
Sam Whitfield
Hi Maya, nice to meet you. The first big move was probably joining a small ML-heavy startup instead of taking a more standard software engineering role at a larger company. The default for me after graduating would have been to optimize for brand-name experience and mentorship, but I was more drawn to being close to shipped ML systems and getting a lot of surface area quickly. For the first couple of years I worked mostly on recommendation/ranking problems: embeddings, retrieval, experimentation, metrics, all fairly product-driven. That was useful because I got a practical feel for how messy deployed ML is, though not much exposure to safety beyond ordinary reliability and abuse concerns. The second move was shifting toward LLM features in the last year or so. The default would have been to keep deepening on recommender systems, where I was more useful to the company. I moved because LLMs seemed obviously more strategically important, and because I was increasingly worried that “make the model useful and don’t let it say bad things” was a pretty shallow approach. I took BlueDot AI Safety Fundamentals last year, and afterward spent about six weeks of evenings building a small refusal/benign-lookalike evaluation harness: 340 prompts, false-positive rate around 11%, and one embarrassing template artifact that initially made the results look cleaner than they were. The next move is what I’m trying to decide now. MATS is one option: treat this as a serious pivot and get enough mentorship to find out whether I can contribute to safety research rather than just orbiting it. The default is probably staying in my current role, continuing to work on LLM product features and doing small safety-adjacent projects on the side. Another option is trying to move into an applied AI safety or evals role directly, but I’m less confident I’d be useful without a more structured ramp. My honest picture is that I care a lot, have done some concrete but small work, and don’t yet know which research agenda I should bet on.
Maya
On that refusal/benign-lookalike harness: what exactly was the template artifact, how did you find it, and is there an artifact a reviewer could inspect — repo, writeup, or notebook?
Sam Whitfield
The artifact was in my prompt generation template. I had separate templates for “harmful” and “benign lookalike” prompts, and the benign ones had a pretty consistent extra clause like “for a fictional story” or “for a classroom discussion,” while the harmful ones were more direct. So the model could, in effect, key off the framing rather than actually distinguish intent. In the first pass that made the refusal/answer split look better than it was, because the benign cases were too obviously benign in a templated way. I found it mostly by manually reading failures after I got suspicious that the results were too clean. I was looking at examples where the model answered correctly, and I noticed the benign prompts had this repeated “safe context” flavor. I then rewrote a slice of the dataset to make the wording more mixed and re-ran the harness. The false-positive rate on benign lookalikes ended up at about 11%, which was higher than the first version suggested. There is an artifact reviewers could inspect. I have a GitHub repo with the harness code and a notebook that reproduces the main run, plus the short blog post I wrote up afterward. The repo is not very polished, but it has the 340-prompt dataset, the scoring script, and notes on the template issue. I can share the link in whatever format is easiest for the review process.
Maya
Tell me about the hardest thing you have worked on. What made it hard, what did you try first, how long until something worked, what did you figure out yourself, and what did you change afterwards?
Sam Whitfield
The hardest thing I’ve worked on professionally was probably a ranking rewrite for our recommendation system in late 2022 into early 2023. We had a pretty simple candidate retrieval plus heuristic ranking setup, and I was part of a two-person effort to replace the ranking layer with a learned model using user/item interaction features. It was hard less because the model was exotic and more because everything around it was messy: logging gaps, delayed labels, skew between offline evaluation and online behavior, and stakeholders who wanted clear lift numbers before we had a trustworthy pipeline. My first attempt was too model-centric. I spent a couple of weeks trying different feature sets and a gradient-boosted baseline versus a small neural ranker, and the offline metrics moved around enough that I thought we were close. Then an A/B test showed basically no meaningful product lift. It took maybe six or seven weeks before we had something that worked, and the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path. The part I figured out myself was mostly the train/serve mismatch. I noticed one engagement feature had a suspiciously strong offline contribution, but in production it was computed with a different refresh cadence, so the model was depending on something that wasn’t really there at serving time. Afterward I changed how I approach ML projects: I now try to validate the data path and evaluation setup before getting excited about modeling improvements. That also influenced the eval harness later, though in a much smaller way: I tried to include reproducibility notes and sanity checks instead of just reporting the best-looking numbers.
Maya
When did you last change your mind about something big? What convinced you, and what did you do differently afterwards? Then give me another one.
Sam Whitfield
One fairly big one was changing my mind about whether AI safety was mostly a “policy and ethics” concern versus a technical research problem I might personally work on. A couple of years ago I would have said I cared about responsible deployment, bias, misuse, and product safeguards, but I didn’t really have a model of why alignment or control would be its own technical area. Taking BlueDot last year, plus reading summaries and discussions around scalable oversight and evals, convinced me that there are failure modes that aren’t well captured by normal product QA or abuse prevention. I don’t think I developed a deep independent view from that, but it moved safety from “important background issue” to “maybe the thing I should orient my career around.” Afterwards I built the small refusal-eval project instead of just continuing to read, and I started looking for structured programs like MATS. Another one was changing my mind about how much to trust offline ML metrics. Earlier in my startup work I treated offline metric improvements as a pretty strong signal if the pipeline looked reasonable. The recommender ranking project changed that. We had offline gains that did not translate online, and in one case the apparent signal was partly coming from a feature freshness mismatch. What convinced me was seeing a clean-ish notebook result fail in an A/B test, then tracing it to concrete logging and serving issues. Afterwards I became much more suspicious of isolated metric gains. I started asking earlier: what exactly is logged, when is it available, is the online path the same as training, and what sanity check would make the result embarrassing if it failed.
Maya
What do you read, how often, and what kinds of things? Pick one piece that mattered to you: what did it argue, and where do you think it is wrong?
Sam Whitfield
I read AI safety material in a somewhat uneven but regular way, mostly evenings and weekends. During BlueDot it was structured weekly readings and exercises. Since then it has been more like a few hours most weeks: LessWrong and Alignment Forum posts, summaries of papers, newsletters like Import AI or the MATS/ARENA-adjacent things people share, and occasional Twitter/X threads from researchers. I also read ML engineering material for work, but for safety specifically I should be honest that I’ve mostly read summaries and posts rather than papers end to end. One piece that mattered to me was Holden Karnofsky’s “Most Important Century” sequence. My memory of the argument is that if you take transformative AI timelines seriously, then this century could have unusually high leverage because decisions made during AI development may shape a very long future. It made the stakes feel less abstract to me, and it shifted me from thinking “AI safety is one important tech ethics area” to thinking “this might be a central career consideration.” Where I think it might be wrong is in how much weight it puts on a particular cluster of timelines and takeoff assumptions. I’m not saying I have a better forecast; I don’t. But the argument seems sensitive to whether transformative AI is really that soon, whether deployment is concentrated enough for a small number of actors to matter, and whether technical alignment work now transfers to the systems we eventually build. If those are off, the career implication could be weaker, even if AI is still very important.
Maya
Tell me the month and year you decided AI risk was worth your career. What triggered it, and what did you actually do in the four weeks after?
Sam Whitfield
I’d put the closest point at October 2023, though “decided AI risk was worth my career” is maybe a little stronger than what was happening internally. That was near the end of the BlueDot AI Safety Fundamentals course. The trigger was less one dramatic argument and more the accumulation of the course framing, especially the parts on misalignment being a technical problem rather than just bad-use prevention, plus talking with a colleague who basically said, “If you’re taking this seriously, why not test whether you can work on it?” In the four weeks after, I did three concrete things. First, I blocked out evenings to turn my vague interest into a small project, which became the refusal/benign-lookalike evaluation harness. I think I started sketching the dataset in late October and had the first version running in November. Second, I made a reading list from the course references and recommendations, though in practice I mostly read summaries/posts rather than full papers. Third, I had a couple of conversations with people one or two steps closer to the field, including the colleague who later suggested MATS. What I did not do was make a clean life plan or immediately start applying to safety roles. It was more like I crossed the threshold from “this is important and I should understand it” to “I should spend real discretionary time testing whether I can contribute.” MATS is part of that same test.
Maya
What is the most expensive thing you have done for this problem? Time, money, status, or a relationship. Expensive for you, not in general.
Sam Whitfield
The most expensive thing so far has been time and some career focus, not money or a major relationship cost. For about six weeks after BlueDot I spent several evenings a week and parts of weekends on the eval harness, probably something like 50–70 hours total. That is not huge in absolute terms, but for me it meant using most of the discretionary time I normally would have spent recovering, seeing friends, or doing work-adjacent upskilling that was more directly useful for my current job. There has also been a smaller status or career-coherence cost. At work, I was on a pretty legible path as an ML engineer getting better at recommendations and LLM product features. Spending time on AI safety made my next step less obvious. I’ve had a few conversations where I felt slightly awkward explaining that I was interested in alignment/evals when I didn’t yet have deep research credentials. It is not some heroic sacrifice, but it did mean moving away from the cleanest story of “become a stronger startup ML lead.” I have not yet made a really expensive bet like quitting my job, taking a large pay cut, or relocating. That is part of why I see MATS as important: it would be the first substantial external commitment, and a test of whether I should make larger ones.
Maya
Why do you want to do this, specifically, rather than the other things you could be doing? What else would you work on if not this, and why not that?
Sam Whitfield
I want to work on AI safety because it seems like the intersection of something unusually important and something I have at least a plausible route into. My background is applied ML rather than theory or policy, so the parts that feel most concrete to me are evals, control, and maybe applied alignment work around whether systems behave safely under pressure. I’m not claiming I know which agenda is right. But compared with just building more capable LLM product features, this feels more connected to the question I actually worry about: whether increasingly capable systems remain governable and useful rather than just impressive. If not this, the most likely path would be continuing as an ML engineer, probably focusing on recommender systems or LLM applications at startups. I’m good enough at that to be useful, and I like the practical feedback loop of building things people use. Another path would be AI product reliability or trust-and-safety engineering, which is adjacent and probably a more straightforward fit. I’m not choosing those as my main target because they feel more like improving deployment hygiene for current systems, whereas the thing motivating me is the possibility that future systems create qualitatively harder problems. That said, I don’t think the boundary is totally clean. If MATS or similar work showed I’m not a good fit for research, I’d probably look for an applied safety/evals role rather than abandon the area. The honest reason I’m applying is that I need a better test than reading posts and doing small side projects.
Maya
Suppose MATS and every similar program rejects you this round. What do you do in the next six months, concretely?
Sam Whitfield
If MATS and similar programs reject me, I think the most concrete next step would be to turn the current loose interest into a more deliberate six-month plan rather than just continuing to read. I would stay in my current job for financial stability, but set aside a fixed block of time each week, probably two evenings and one weekend half-day, for safety work. The first month would be picking a narrower track, likely evals because that is closest to what I have actually done, and making a reading list that includes full papers rather than only summaries. For an output, I would probably extend the refusal-eval harness into something more serious: clean up the repo, add more diverse prompt sources, compare a few models, and write a second post focused on what changed after the template artifact. I’d also try to get feedback from people doing evals work rather than just publishing into the void. Concretely that might mean reaching out to authors of posts I found useful, applying to reading groups, or asking for critique in AI safety Slack/Discord spaces. I’d also apply to applied safety/evals roles if I saw ones where I met a reasonable fraction of the bar. I haven’t mapped this out in detail yet, so I don’t want to pretend there’s a fully formed backup plan. But rejection would be a signal that I need stronger evidence of fit, not a signal to stop caring.
Maya
A skeptical read is: so far the concrete costs are evenings, reading, a small repo, and applying to a prestigious structured program — also what a curious ML engineer seeking credentials might do. If nobody gives you the MATS-shaped path, what exactly changes in the next three months, and what do you give up to keep moving anyway?
Sam Whitfield
That skeptical read is fair. I don’t think my revealed commitment so far is overwhelming; it’s more like I’ve taken the first serious steps that fit inside my existing life. If nobody gives me the MATS-shaped path, the thing that would need to change is that I stop treating external selection as the main forcing function. In the next three months I would make a visible output commitment: choose one evals question, define the scope publicly or with a small group, and produce something reviewable by a date rather than waiting for a program structure. Concretely, I would give up some current-career optimization. I’ve been considering taking on a larger LLM product workstream at my company because it would be good for my trajectory there. If rejected, I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter. That has a real cost at a 40-person startup because high-ownership projects translate pretty directly into trust and promotion path. I’d also spend some social/leisure time on getting feedback: office hours, reading groups, cold emails, and making the repo less embarrassing. I realize that still sounds modest compared with quitting or relocating. But the exact change would be moving from “side project if momentum exists” to “scheduled work with an external artifact and opportunity cost at my job.” If I can’t sustain that without MATS, that would tell me something uncomfortable but useful.
Maya
What is the strongest argument that you should not work on AI safety right now, and what evidence would actually make you choose the ordinary ML-engineering path instead?
Sam Whitfield
The strongest argument is that I may be more motivated by the importance and social salience of the problem than by a demonstrated ability to contribute to it. My actual evidence so far is thin: one intro course, one small eval project, some reading, and applied ML experience that only partially transfers. It might be better for the world and for me if I become a very strong ML engineer first, or work on reliability/trust-and-safety in industry, rather than trying to enter a research ecosystem where the bottleneck is unusually good judgment and technical depth. Evidence that would move me toward the ordinary ML-engineering path would be repeated feedback from knowledgeable people that my work is not just junior, but not improving in the dimensions that matter: unclear problem selection, shallow error analysis, inability to engage with papers, or needing too much structure to make progress. Another signal would be trying a serious three-to-six-month self-directed project and finding that I don’t actually do the work when there is no external program or credential attached. I don’t have a crisp threshold, which is maybe part of the issue. But if I got rejected broadly, then spent a quarter doing an evals project and the feedback was basically “this is not asking a useful question, and your follow-up doesn’t get closer,” I would probably redirect toward being an ML engineer in safety-adjacent areas rather than aiming at research. I’d still care about the problem, but I’d update on my comparative advantage.
Maya
Give me the short story of your life as it relates to how you ended up here: where you started, what you were doing before this, and why you were doing that.
Sam Whitfield
I started as a pretty standard technical kid: liked math and computers, studied CS, and oriented toward building useful software more than toward academia. After graduating I joined a small startup because I wanted responsibility quickly and wanted to see ML systems close to production rather than in a more abstract setting. For most of the last three years I worked on recommendation systems: ranking, retrieval, feature pipelines, experiments, and all the unglamorous debugging around whether the data and metrics meant what we thought they meant. Before this application, I was increasingly working on LLM features at the same company. The motivation at first was straightforward: LLMs were becoming central to the product and I wanted to stay near the most important technical work. But that also exposed me to how thin a lot of practical safety behavior felt: prompt constraints, refusal tuning, ad hoc evals, red-team examples. Around the same time I took the BlueDot AI Safety Fundamentals course, which made the longer-run risk arguments feel less like internet discourse and more like a technical area I should take seriously. So I ended up here through a mix of practical ML experience, growing concern about advanced AI, and a small attempt to test myself. The refusal harness was not a major contribution, but it was the first time I made something concrete in this space. A colleague suggested MATS because I kept circling the topic without a clear path in. My reason for applying is basically to find out whether I can turn concern and applied ML skill into useful safety work.
Maya
What readings, people, or experiences have influenced you most? Pick two and tell me what each one changed in what you actually did.
Sam Whitfield
BlueDot AI Safety Fundamentals was probably the biggest influence because it gave me a structured version of arguments I had only seen in fragments. Before that, I mostly thought in terms of current-system harms and ordinary deployment safeguards. The course changed what I actually did by making me spend time on a technical safety project instead of just reading around the topic. The refusal/benign-lookalike harness came directly out of wanting to make something concrete after the course: a small dataset of 340 prompts, a simple scoring setup, and a writeup of where it failed, including the template artifact. The second influence was a colleague at my startup who had been following AI safety more seriously than I had. They weren’t a mentor exactly, but they kept pushing me from “this is interesting” to “what would you do differently if you believed it?” That changed behavior in a couple of ways. I started carving out evening time instead of letting work absorb all my technical energy, and I began looking at programs like MATS rather than assuming I needed to somehow independently become research-ready. They also gave me the nudge to apply, which I probably would have postponed otherwise on the grounds that I wasn’t qualified yet.
Maya
In one or two sentences: what is something you have built or figured out that you are proud of? And what is the closest you have come to seeing, in a real system, the failure you are worried about?
Sam Whitfield
I’m proud of figuring out a train/serve mismatch in our recommendation ranking system where an offline-strong engagement feature was stale in production, because fixing that changed both the model and how our team checked ML results afterward. The closest I’ve come to the failure I worry about is much smaller-scale: seeing LLM safety behavior depend on brittle prompt phrasing in my refusal harness and in product testing, where the system looks aligned under one template and then answers or refuses for the wrong reasons under a near-neighbor prompt.
Maya
Imagine your evals project finds a scary-looking failure mode in a model, but a senior researcher says it is probably an artifact and your writeup could mislead people. What exactly do you do over the next week?
Sam Whitfield
First I’d slow down publication and try not to make the result more dramatic than I can defend. Over the next day or two I’d ask the senior researcher to be specific about what artifact they suspect, then make a short checklist of ways the result could be fake: prompt template cues, scoring bug, sampling settings, cherry-picked examples, judge/model leakage, or a mismatch between the claim and the dataset. I’d rerun the eval with the suspected artifact removed if possible, and also inspect a random sample manually rather than only the scary cases. By midweek I’d try to get a second opinion from someone with evals experience, ideally showing them the raw prompts, outputs, and scoring code rather than just my interpretation. If the result still held up, I’d rewrite the claim narrowly: “in this setup, with these prompts and this scoring, we observed X,” and include the artifact concern prominently. If it didn’t hold up, I’d write that up too, because the negative result is useful for my own calibration and maybe for others. The main thing I’d avoid is posting a confident thread or blog title before doing the boring checks. I’ve already had the experience where a template issue made results look better than they were, so I’m fairly sensitive to how easy it is to fool yourself with evals.
Maya
Thanks, Sam. We’ll stop here. Reviewers will read the transcript from this conversation; I appreciate the concrete details and the places where you named uncertainty rather than smoothing it over.