Soundings

Mission alignment

Sam Whitfield

Sep 11, 2026, 10:23 PM · 31 turns · text

Sam Whitfield describes himself as an applied ML engineer who joined a small ML-heavy startup after CS, worked mostly on recommendation/ranking systems for about three years, and shifted in the last year toward LLM product features. He took BlueDot AI Safety Fundamentals in 2023 and built a small refusal/benign-lookalike eval harness with 340 prompts, which exposed a template artifact. He wants to test whether he can contribute to AI safety research, with MATS as the structured ramp, while keeping applied safety/evals or ordinary ML engineering as fallback options.

Mission alignment

7/10Mission alignmentChose impact once, at a cost.
confidence highreceipts 4/4 verified
Career moves
After graduating: Joined a small ML-heavy startup rather than a larger-company software engineering role; worked for the first couple years mostly on recommendation/ranking problems.; Last year or so before the interview: Shifted toward LLM features at his company.; Now / current application period: Considering MATS as a serious pivot toward AI safety research, while also considering staying in his current role or moving directly into applied safety/evals.
Without a place
He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.

Why 7 · Sam’s concern about AI risk has moved beyond abstract interest into concrete work and career reorientation, but it is still early and framed as testing fit rather than a life already organized around impact. He has paid real but modest costs in discretionary time, social/recovery time, and some career-coherence/opportunity cost at his startup. The no-program counterfactual is credible: he says he would continue with scheduled safety work and give up a promotion-relevant product workstream, though he has not yet made large sustained sacrifices or repeated major impact-driven choices.

  • I’d put the closest point at October 2023, though “decided AI risk was worth my career” is maybe a little stronger than what was happening internally.

    Dates the pivot while calibrating its strength; shows this is a recent shift toward impact, not just polished mission language.

  • I crossed the threshold from “this is important and I should understand it” to “I should spend real discretionary time testing whether I can contribute.”

    The concern changed behavior into active contribution-testing rather than staying at belief or reading.

  • For about six weeks after BlueDot I spent several evenings a week and parts of weekends on the eval harness, probably something like 50–70 hours total... using most of the discretionary time I normally would have spent recovering, seeing friends, or doing work-adjacent upskilling that was more directly useful for my current job.

    Concrete cost in time, rest, social life, and career-useful upskilling for safety-adjacent work.

  • If rejected, I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter. That has a real cost at a 40-person startup because high-ownership projects translate pretty directly into trust and promotion path.

    No-program counterfactual: he names a specific opportunity cost he would pay to keep working on the problem without the credential.

Innate traits

7/10JudgmentIn their own area.
confidence mediumreceipts 4/4 verified
Refusal/benign-lookalike evaluation harness
hard: He described it as "concrete but small work" and later as "not a major contribution"; the main issue was a prompt-template artifact that made results look cleaner than they were.; first result: After BlueDot, he spent about six weeks of evenings; he started sketching the dataset in late October and had the first version running in November. The first pass used separate harmful and benign templates and looked too clean until he found the artifact.
Recommendation-system ranking rewrite
hard: "It was hard less because the model was exotic and more because everything around it was messy: logging gaps, delayed labels, skew between offline evaluation and online behavior, and stakeholders who wanted clear lift numbers before we had a trustworthy pipeline."; first result: His first attempt was model-centric: for a couple of weeks he tried different feature sets and a gradient-boosted baseline versus a small neural ranker. An A/B test then showed no meaningful product lift; it took six or seven weeks before something worked, after debugging logging, traffic segments, stale online features, and train/serve path.
LLM product features at current company
hard: He said practical safety behavior felt thin and brittle: "prompt constraints, refusal tuning, ad hoc evals, red-team examples."; first result: not discussed

Why 7 · I’d trust Sam’s judgment inside applied ML/evals work, especially around whether results are real, how to debug deployed systems, and how to avoid overclaiming. His judgment also travels somewhat into career and field-assessment questions: he names weak evidence, alternative explanations, and decision points without needing to sound certain. I would not yet hand him a broad unscoped research direction and assume he will choose the right problem; he is appropriately still seeking structure and has thin direct safety-research evidence.

  • My first attempt was too model-centric... the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path.

    Shows strong practical judgment from a real past decision: he noticed the real bottleneck was data/evaluation validity, not model sophistication.

  • I don’t think I developed a deep independent view from that, but it moved safety from “important background issue” to “maybe the thing I should orient my career around.”

    He bounds his own expertise rather than overstating conviction; that makes me more willing to trust him outside his core area.

  • the argument seems sensitive to whether transformative AI is really that soon, whether deployment is concentrated enough for a small number of actors to matter, and whether technical alignment work now transfers to the systems we eventually build.

    Identifies the load-bearing assumptions in an important argument, rather than merely endorsing the conclusion.

  • I’d ask the senior researcher to be specific about what artifact they suspect, then make a short checklist of ways the result could be fake: prompt template cues, scoring bug, sampling settings, cherry-picked examples, judge/model leakage, or a mismatch between the claim and the dataset.

    In the live scenario, he gives the right operational response: slow down, look for artifacts, get concrete, and narrow the claim if needed.

8/10Bias resistanceThey go looking.
confidence highreceipts 4/4 verified

Why 8 · Sam seems unusually willing to put his own motivations and results under test. He repeatedly volunteers weak points before being pushed, describes concrete past updates caused by failed evidence, and in hypotheticals says he would slow down, rerun checks, narrow claims, or publish the null rather than protect a dramatic result. I would not put him near 10: much of this is still planned behavior and he sometimes lacks crisp thresholds, but the pattern is clearly beyond merely conceding when cornered.

  • one embarrassing template artifact that initially made the results look cleaner than they were

    He volunteers a flaw in his own project before being challenged, suggesting he is not trying to defend the cleanest version of his work.

  • What convinced me was seeing a clean-ish notebook result fail in an A/B test, then tracing it to concrete logging and serving issues. Afterwards I became much more suspicious of isolated metric gains.

    This is a real past change of mind driven by contrary evidence, with a specific behavioral update afterward.

  • The strongest argument is that I may be more motivated by the importance and social salience of the problem than by a demonstrated ability to contribute to it.

    When asked for the case against him, he names an unflattering alternative explanation rather than a strawman.

  • If the result still held up, I’d rewrite the claim narrowly... If it didn’t hold up, I’d write that up too, because the negative result is useful for my own calibration

    In response to contrary expert feedback, he would run checks that could kill his preferred result and change the public claim accordingly.

8/10OpennessScans for what serves the goal.
confidence highreceipts 4/4 verified

Why 8 · Sam looks quite open in a goal-directed way: he has a stable target around contributing to AI safety, but is willing to change methods, roles, and degree of research ambition based on evidence. He has concrete examples of abandoning a favored model-centric approach for data/evaluation debugging, and he explicitly treats MATS/research, applied evals, safety-adjacent ML engineering, and self-directed projects as live paths. I would not put him at the very top because the space he considers is still mostly within applied ML/evals/safety-adjacent work, and many of the larger pivots are conditional rather than already done.

  • My first attempt was too model-centric... Then an A/B test showed basically no meaningful product lift... the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path.

    He describes changing methods when the evidence showed his preferred modeling route was not serving the goal.

  • Afterward I changed how I approach ML projects: I now try to validate the data path and evaluation setup before getting excited about modeling improvements.

    This is a durable habit change, not just a one-off acknowledgment.

  • If MATS or similar work showed I’m not a good fit for research, I’d probably look for an applied safety/evals role rather than abandon the area.

    He is willing to switch role/approach while keeping the underlying goal fixed.

  • If I got rejected broadly, then spent a quarter doing an evals project and the feedback was basically ‘this is not asking a useful question, and your follow-up doesn’t get closer,’ I would probably redirect toward being an ML engineer in safety-adjacent areas rather than aiming at research.

    He names concrete evidence that would make him give up the research path and choose a different contribution route.

7/10AgencyHas made a move.
confidence highreceipts 4/4 verified
Moves against the default
3 of 3 career moves
Next step
He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.
Options weighed
5
First step taken
He took BlueDot AI Safety Fundamentals, built the refusal/benign-lookalike eval harness, started reading safety material, talked with people closer to the field, and applied to MATS.

Why 7 · Sam has clearly authored several meaningful moves rather than just following the clean ML-engineering path: choosing a small startup over a safer branded role, moving toward LLMs, and turning AI-safety interest into a concrete eval project. The costs so far are real but still moderate, and he is unusually candid that his revealed commitment is not yet overwhelming. His next-step thinking is better than “wait for MATS,” with a specific fallback and named opportunity cost, but much of it is still conditional rather than already underway, which keeps him below the level where I’d say he fully authors the path.

  • The first big move was probably joining a small ML-heavy startup instead of taking a more standard software engineering role at a larger company. The default for me after graduating would have been to optimize for brand-name experience and mentorship, but I was more drawn to being close to shipped ML systems and getting a lot of surface area quickly.

    He names the default, rejects it, and gives a self-directed reason for the alternative.

  • Afterwards I built the small refusal-eval project instead of just continuing to read, and I started looking for structured programs like MATS.

    A belief update led to production and application, not just more passive consumption.

  • For about six weeks after BlueDot I spent several evenings a week and parts of weekends on the eval harness, probably something like 50–70 hours total... it meant using most of the discretionary time I normally would have spent recovering, seeing friends, or doing work-adjacent upskilling that was more directly useful for my current job.

    The move had a concrete personal opportunity cost, though he is clear it was not a huge sacrifice.

  • If nobody gives me the MATS-shaped path, the thing that would need to change is that I stop treating external selection as the main forcing function... I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter.

    His fallback plan is self-authored and includes giving up career capital, but it is still framed as a future conditional rather than something already started.

8/10General reasoningFast and generative.
confidence highreceipts 4/4 verified

Why 8 · Sam is fast and generative on applied ML/evals reasoning: when presented with artifact or evidence-quality concerns, he immediately decomposes mechanisms, proposes checks, and narrows claims to what the data can support. He repeatedly distinguishes apparent performance from the causal reason for it, and he updates plans based on what would falsify his interpretation rather than defending the shiny result. I would scope the 8 mainly to ML systems/evals and practical research judgment; the transcript gives less evidence about abstract reasoning far outside that domain.

  • I noticed one engagement feature had a suspiciously strong offline contribution, but in production it was computed with a different refresh cadence, so the model was depending on something that wasn’t really there at serving time.

    Mechanistic diagnosis of an offline/online mismatch, not just 'the metric failed'; he identifies the causal path and changes his process afterward.

  • the model could, in effect, key off the framing rather than actually distinguish intent. In the first pass that made the refusal/answer split look better than it was

    Shows he can see through an apparent eval result to the shortcut explanation underneath, and correctly revises the interpretation.

  • I’d ask the senior researcher to be specific about what artifact they suspect, then make a short checklist of ways the result could be fake: prompt template cues, scoring bug, sampling settings, cherry-picked examples, judge/model leakage, or a mismatch between the claim and the dataset.

    In the unfamiliar push scenario, he immediately turns a vague criticism into concrete failure modes and tests rather than either deferring or defending.

  • If the result still held up, I’d rewrite the claim narrowly: “in this setup, with these prompts and this scoring, we observed X,” and include the artifact concern prominently. If it didn’t hold up, I’d write that up too

    Good evidence calibration: he separates result, claim scope, uncertainty, and the value of a negative/artifact-disconfirming outcome.

8/10GrowthFast, honest turns.
confidence highreceipts 4/4 verified
Refusal/benign-lookalike evaluation harness
changed after: He rewrote part of the dataset to mix wording more, reran the harness, added reproducibility notes and sanity checks, and says he is sensitive to artifact risks in evals.
Recommendation-system ranking rewrite
changed after: He changed his ML-project approach to validate the data path and evaluation setup before getting excited about modeling improvements; this also influenced the later eval harness.
LLM product features at current company
changed after: The work contributed to his concern that current practical safety behavior can depend on brittle prompt phrasing and motivated him toward AI safety.

Why 8 · Sam shows several complete growth arcs: he notices when results are misleading, investigates the unglamorous failure mode, and changes his process afterward. The strongest pattern is that mistakes make him more operationally cautious rather than just more verbally reflective: he now checks data paths, train/serve consistency, artifacts, and sanity checks earlier. He also names an unflattering possible bottleneck — that he may need external structure or be drawn by salience more than demonstrated fit — which makes the growth evidence feel honest rather than polished.

  • I found it mostly by manually reading failures after I got suspicious that the results were too clean... I then rewrote a slice of the dataset to make the wording more mixed and re-ran the harness.

    Concrete block-reaction-change arc: he detected an artifact, inspected failures, changed the dataset, and reran rather than just noting the issue.

  • My first attempt was too model-centric... Then an A/B test showed basically no meaningful product lift. It took maybe six or seven weeks before we had something that worked, and the useful progress came from less glamorous debugging: checking whether impressions were logged consistently, segmenting by traffic source, finding a feature that was available offline but stale online, and tightening the train/serve path.

    He identifies the real mistake in his approach and describes multiple follow-up routes that eventually changed the outcome.

  • Afterward I changed how I approach ML projects: I now try to validate the data path and evaluation setup before getting excited about modeling improvements.

    The lesson became a procedural change carried into later work, not just a generic statement about resilience or rigor.

  • The strongest argument is that I may be more motivated by the importance and social salience of the problem than by a demonstrated ability to contribute to it... Another signal would be trying a serious three-to-six-month self-directed project and finding that I don’t actually do the work when there is no external program or credential attached.

    He names a non-flattering possible bottleneck and gives observable evidence that would update him, which supports the 'honest diagnosis' part of the trait.

5/10AmbitionGood work.
confidence highreceipts 4/4 verified
Wants
He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.
Options
MATS as a serious pivot with mentorship toward safety research; Stay in his current role, continue LLM product features, and do small safety-adjacent projects on the side; Move directly into an applied AI safety or evals role; If research fit looks poor, work as an ML engineer in safety-adjacent areas such as applied safety/evals, reliability, or trust-and-safety engineering; Continue ordinary ML engineering focused on recommender systems or LLM applications at startups

Why 5 · Sam is reaching for work that matters a lot—potentially making AI safety/evals his career—but his personal picture of the future is still framed as testing fit and becoming useful within existing programs or roles. His current choices show real but bounded organization around that aim: an intro course, a concrete evals repo/writeup, applying to MATS, and a stated willingness to give up some startup career optimization for scheduled safety work. This sits above ordinary good-work motivation, but below the “build/run their own thing” or field-changing ambition levels because the scale and sacrifice remain deliberately modest.

  • MATS is one option: treat this as a serious pivot and get enough mentorship to find out whether I can contribute to safety research rather than just orbiting it.

    He is aiming at a meaningful career pivot into safety research, but frames it as testing whether he can contribute rather than pursuing a large independent agenda.

  • It was more like I crossed the threshold from “this is important and I should understand it” to “I should spend real discretionary time testing whether I can contribute.”

    Shows the problem has started organizing his discretionary time, while also marking the commitment as exploratory.

  • I have not yet made a really expensive bet like quitting my job, taking a large pay cut, or relocating. That is part of why I see MATS as important: it would be the first substantial external commitment

    He explicitly names the current ceiling on revealed commitment, keeping the placement below the higher ambition rungs.

  • If rejected, I would likely not volunteer for that, and instead protect roughly 8–10 hours a week for safety work for at least a quarter. That has a real cost at a 40-person startup because high-ownership projects translate pretty directly into trust and promotion path.

    The planned next step involves a real opportunity cost and scheduled output, but still at the scale of self-development and a side project rather than building an institution or changing a field.

7/10InterpersonalStraight and decent.
confidence mediumreceipts 4/4 verified

Why 7 · Sam comes across as straight, self-critical, and easy to push: he accepts skeptical framings without heat and answers them concretely rather than defending his image. He talks about collaborators and critics as useful people rather than obstacles, and his imagined response to critique is notably evidence-sharing and collaborative. I do not see enough warmth or enacted care to place him at the candid-and-warm rung, but he is clearly above merely guarded politeness.

  • That skeptical read is fair. I don’t think my revealed commitment so far is overwhelming; it’s more like I’ve taken the first serious steps that fit inside my existing life.

    Takes a pointed skeptical read directly and non-defensively, without attacking the framing.

  • By midweek I’d try to get a second opinion from someone with evals experience, ideally showing them the raw prompts, outputs, and scoring code rather than just my interpretation.

    Responds to possible criticism by exposing evidence and seeking outside judgment, rather than controlling the narrative.

  • The second influence was a colleague at my startup who had been following AI safety more seriously than I had. They weren’t a mentor exactly, but they kept pushing me from “this is interesting” to “what would you do differently if you believed it?”

    Credits another person’s influence specifically and respectfully, presenting them as an agent in his development rather than a prop.

  • I’d put the closest point at October 2023, though “decided AI risk was worth my career” is maybe a little stronger than what was happening internally.

    Gently corrects the interviewer’s framing while staying cooperative and precise.

9/10IntegrityBounded by default.
confidence highreceipts 4/4 verified
Refusal/benign-lookalike evaluation harness
their part: He built the harness, created a 340-prompt dataset, wrote scoring code, manually inspected failures, rewrote a slice of the dataset, reran the harness, and wrote a short blog post. No collaborators were described for the build itself.; numbers: 340 prompts; false-positive rate around 11% on benign lookalikes after revision; GitHub repo with harness code, notebook reproducing the main run, scoring script, dataset, notes on the template issue, and a short blog post.
Recommendation-system ranking rewrite
their part: He was part of a two-person effort. He personally figured out a train/serve mismatch involving an engagement feature with strong offline contribution but different production refresh cadence.; numbers: Late 2022 into early 2023; two-person effort; six or seven weeks until something worked; A/B test initially showed basically no meaningful product lift.
LLM product features at current company
their part: He says he was increasingly working on LLM features at the same company; specific individual contributions were not discussed.; numbers: not discussed

Why 9 · Sam repeatedly bounds his own claims and volunteers details that weaken his case before being pressed. His specifics generally survive follow-up: when asked about the eval artifact, he gives a concrete mechanism, how he found it, what changed, and inspectable artifacts. In pressure scenarios, he says he would slow down, expose raw materials to others, narrow or retract the claim, and publish the negative/update rather than take the status win. I would trust his self-report substantially, though I’d stop just short of a 10 because the costly pressure answer is mostly procedural rather than a clearly painful personal concession.

  • 340 prompts, false-positive rate around 11%, and one embarrassing template artifact that initially made the results look cleaner than they were.

    He volunteers a flaw in his own project before the interviewer asks about it, rather than presenting the work as cleaner or more impressive.

  • The artifact was in my prompt generation template... the benign ones had a pretty consistent extra clause like “for a fictional story” or “for a classroom discussion,” while the harmful ones were more direct. So the model could, in effect, key off the framing rather than actually distinguish intent.

    On follow-up, the weakening detail becomes more specific and technically plausible; it does not collapse under a second question.

  • I should be honest that I’ve mostly read summaries and posts rather than papers end to end.

    He limits a credential-like claim about reading depth instead of letting a broad description of engagement stand unqualified.

  • If the result still held up, I’d rewrite the claim narrowly... and include the artifact concern prominently. If it didn’t hold up, I’d write that up too... The main thing I’d avoid is posting a confident thread or blog title before doing the boring checks.

    In the pressure scenario, he chooses disclosure, narrowing, and possible deflation of his own result over a more attention-grabbing presentation.

5/10ReadingSummaries.
confidence highreceipts 4/4 verified
How often
Uneven but regular; during BlueDot it was weekly structured readings and exercises, and since then a few hours most weeks, mostly evenings and weekends.
Kinds
LessWrong and Alignment Forum posts, Summaries of papers, Newsletters such as Import AI, MATS/ARENA-adjacent shared materials, Occasional Twitter/X threads from researchers, ML engineering material for work, Mostly summaries and posts rather than full papers end to end
Pieces named
3: Holden Karnofsky’s "Most Important Century" sequence; BlueDot AI Safety Fundamentals course references and recommendations; Import AI

Why 5 · Sam has a real but limited reading habit: a few hours most weeks, mostly AI safety posts, paper summaries, newsletters, and course materials. He does process at least some of it—he can summarize a central argument from the Most Important Century sequence and name plausible weak points—but he is candid that he mostly has not been reading full papers end to end. This puts him above a pure summaries/titles level, but below the “reads properly, papers and books regularly” anchor.

  • Since then it has been more like a few hours most weeks: LessWrong and Alignment Forum posts, summaries of papers, newsletters like Import AI or the MATS/ARENA-adjacent things people share, and occasional Twitter/X threads from researchers.

    Shows a regular reading habit and the main sources, but the sources are mostly posts, summaries, newsletters, and threads.

  • I also read ML engineering material for work, but for safety specifically I should be honest that I’ve mostly read summaries and posts rather than papers end to end.

    Clear limiter on depth: he is not yet regularly reading primary papers in the area.

  • My memory of the argument is that if you take transformative AI timelines seriously, then this century could have unusually high leverage because decisions made during AI development may shape a very long future.

    He can give the central argument of a piece in his own words, albeit at a fairly high level.

  • Where I think it might be wrong is in how much weight it puts on a particular cluster of timelines and takeoff assumptions... whether technical alignment work now transfers to the systems we eventually build.

    Shows some actual processing and criticism of the argument, enough to place above the basic summaries-only rung.

Facts

Work they described

Refusal/benign-lookalike evaluation harness

A small AI safety/evals side project after BlueDot testing refusal behavior on harmful versus benign-lookalike prompts.

How hardHe described it as "concrete but small work" and later as "not a major contribution"; the main issue was a prompt-template artifact that made results look cleaner than they were.
Their partHe built the harness, created a 340-prompt dataset, wrote scoring code, manually inspected failures, rewrote a slice of the dataset, reran the harness, and wrote a short blog post. No collaborators were described for the build itself.
First resultAfter BlueDot, he spent about six weeks of evenings; he started sketching the dataset in late October and had the first version running in November. The first pass used separate harmful and benign templates and looked too clean until he found the artifact.
Numbers340 prompts; false-positive rate around 11% on benign lookalikes after revision; GitHub repo with harness code, notebook reproducing the main run, scoring script, dataset, notes on the template issue, and a short blog post.
Changed afterHe rewrote part of the dataset to mix wording more, reran the harness, added reproducibility notes and sanity checks, and says he is sensitive to artifact risks in evals.
  • "I took BlueDot AI Safety Fundamentals last year, and afterward spent about six weeks of evenings building a small refusal/benign-lookalike evaluation harness: 340 prompts, false-positive rate around 11%, and one embarrassing template artifact that initially made the results look cleaner than they were."

Recommendation-system ranking rewrite

A professional ranking rewrite in late 2022 into early 2023, replacing a heuristic ranking layer with a learned model using user/item interaction features.

How hard"It was hard less because the model was exotic and more because everything around it was messy: logging gaps, delayed labels, skew between offline evaluation and online behavior, and stakeholders who wanted clear lift numbers before we had a trustworthy pipeline."
Their partHe was part of a two-person effort. He personally figured out a train/serve mismatch involving an engagement feature with strong offline contribution but different production refresh cadence.
First resultHis first attempt was model-centric: for a couple of weeks he tried different feature sets and a gradient-boosted baseline versus a small neural ranker. An A/B test then showed no meaningful product lift; it took six or seven weeks before something worked, after debugging logging, traffic segments, stale online features, and train/serve path.
NumbersLate 2022 into early 2023; two-person effort; six or seven weeks until something worked; A/B test initially showed basically no meaningful product lift.
Changed afterHe changed his ML-project approach to validate the data path and evaluation setup before getting excited about modeling improvements; this also influenced the later eval harness.
  • "The part I figured out myself was mostly the train/serve mismatch. I noticed one engagement feature had a suspiciously strong offline contribution, but in production it was computed with a different refresh cadence, so the model was depending on something that wasn’t really there at serving time."

LLM product features at current company

Increasing work on LLM features at the same startup, including practical exposure to prompt constraints, refusal tuning, ad hoc evals, and red-team examples.

How hardHe said practical safety behavior felt thin and brittle: "prompt constraints, refusal tuning, ad hoc evals, red-team examples."
Their partHe says he was increasingly working on LLM features at the same company; specific individual contributions were not discussed.
Changed afterThe work contributed to his concern that current practical safety behavior can depend on brittle prompt phrasing and motivated him toward AI safety.
  • "Before this application, I was increasingly working on LLM features at the same company. The motivation at first was straightforward: LLMs were becoming central to the product and I wanted to stay near the most important technical work. But that also exposed me to how thin a lot of practical safety behavior felt: prompt constraints, refusal tuning, ad hoc evals, red-team examples."

Career moves

  • After graduating Joined a small ML-heavy startup rather than a larger-company software engineering role; worked for the first couple years mostly on recommendation/ranking problems. instead of A more standard software engineering role at a larger company, optimizing for brand-name experience and mentorship.. He wanted proximity to shipped ML systems and broad responsibility quickly.
  • Last year or so before the interview Shifted toward LLM features at his company. instead of Continuing to deepen on recommender systems, where he was more useful to the company.. He thought LLMs were more strategically important and became worried that practical safety approaches were shallow.
  • Now / current application period Considering MATS as a serious pivot toward AI safety research, while also considering staying in his current role or moving directly into applied safety/evals. instead of Staying in his current role, continuing LLM product features and doing small safety-adjacent projects on the side.. He wants mentorship and a better test of whether he can contribute to safety research rather than just orbiting it.

What they want next

He wants to use MATS as a structured test of whether he can contribute to AI safety research. If rejected, he plans to stay in his current job for stability while protecting 8–10 hours a week for safety work, choosing an evals question, producing a reviewable artifact within three months, and extending the refusal-eval harness over six months.

Options they see: MATS as a serious pivot with mentorship toward safety research; Stay in his current role, continue LLM product features, and do small safety-adjacent projects on the side; Move directly into an applied AI safety or evals role; If research fit looks poor, work as an ML engineer in safety-adjacent areas such as applied safety/evals, reliability, or trust-and-safety engineering; Continue ordinary ML engineering focused on recommender systems or LLM applications at startups

Already done: He took BlueDot AI Safety Fundamentals, built the refusal/benign-lookalike eval harness, started reading safety material, talked with people closer to the field, and applied to MATS.

Reading

Uneven but regular; during BlueDot it was weekly structured readings and exercises, and since then a few hours most weeks, mostly evenings and weekends.; LessWrong and Alignment Forum posts, Summaries of papers, Newsletters such as Import AI, MATS/ARENA-adjacent shared materials, Occasional Twitter/X threads from researchers, ML engineering material for work, Mostly summaries and posts rather than full papers end to end

Holden Karnofsky’s "Most Important Century" sequence; BlueDot AI Safety Fundamentals course references and recommendations; Import AI

Exposure to the field

  • BlueDot AI Safety Fundamentals Last year; near the end of the course in October 2023 · Yes · It gave him a structured version of AI safety arguments, shifted him from current-system harms and deployment safeguards toward seeing misalignment as a technical problem, and led him to start a concrete evals project.
  • Refusal/benign-lookalike evaluation harness Started sketching dataset in late October 2023; first version running in November; about six weeks after BlueDot · He completed a first version and writeup; he describes it as small and not polished, with possible future extension. · He gained experience building an eval, finding a prompt-template artifact, rerunning after dataset changes, and adding notes/sanity checks.
  • MATS application / MATS as a possible program Current application period · No; he is applying and using it as an option to test fit. · Not completed; he hopes it would provide mentorship and a serious test of whether he can contribute to safety research.
  • AI safety conversations with people closer to the field In the four weeks after October 2023 / after BlueDot · He had a couple of conversations; ongoing/repeated exposure not fully mapped. · They helped him move from interest to concrete action and introduced or nudged him toward MATS.

Influences

  • BlueDot AI Safety Fundamentals · Moved AI safety from an important background issue/current-system-harms frame to something he might orient his career around; led him to build the refusal/benign-lookalike harness and look for structured programs.
  • A colleague at his startup who followed AI safety more seriously · Pushed him from interest to asking what he would do differently if he believed it; he carved out evening time, built safety work into his schedule, looked at MATS, and applied rather than postponing.
  • Holden Karnofsky’s "Most Important Century" sequence · Made AI safety stakes feel less abstract and shifted him from seeing AI safety as one tech ethics area to a central career consideration, while leaving him uncertain about timelines and takeoff assumptions.
  • Recommendation-system ranking rewrite / offline metrics failing online · Made him much more skeptical of isolated offline metric gains and more focused on data paths, logging, train/serve consistency, and sanity checks.

Threads not followed (28)

  • stay in my current job for financial stabilityUseful cross-check for money/opportunity cost and whether safety remains load-bearing alongside product ML.
  • I’ve been considering taking on a larger LLM product workstream at my companyConcrete foregone opportunity; useful for cross-checking money/opportunity cost and mission load-bearingness.
  • I don’t have a crisp thresholdCould be probed for kill criteria, but time is limited and later fixed questions can cover this.
  • possible exceptional, untestedif I got rejected broadly, then spent a quarter doing an evals project and the feedback was basically “this is not asking a useful question, and your follow-up doesn’t get closer,”Concrete update condition; useful evidence for bias resistance and openness.
  • A colleague suggested MATS because I kept circling the topic without a clear path in.Could reveal interpersonal influence and whether the pivot was externally authored, but lower priority than fixed influence question.
  • how thin a lot of practical safety behavior felt: prompt constraints, refusal tuning, ad hoc evals, red-team examplesPotential concrete disagreement/threat model thread if time allowed.
  • a colleague at my startup who had been following AI safety more seriously than I hadCould check whether the pivot was mostly externally authored and how they engage with influence.
  • They also gave me the nudge to apply, which I probably would have postponed otherwisePotential thinness in self-directed agency; could probe if more time.
  • what would you do differently if you believed it?Good mission-alignment hook: belief causing behavior change.
  • I’ve already had the experience where a template issue made results look better than they wereConnects current reasoning to a concrete past mistake; would be worth probing for learning transfer if time allowed.

Behaviour

Candidate turns15
Median answer length243 words
Median time to answer78s
Turns containing pasted text0