Soundings

Research judgment

Daniel Okafor

Sep 11, 2026, 10:23 PM · 29 turns · text

Daniel Okafor says he is a backend/infrastructure engineer at a fintech with about six years of experience, and is early in AI safety. He completed BlueDot AI Safety Fundamentals, has read Alignment Forum/LessWrong material, and built a solo jailbreak classifier/eval harness that looked strong on a random split but degraded on an adversarial holdout. He wants MATS or a similar mentored environment to test whether he can develop research judgment beyond engineering/evals work.

Research taste

8/10Research tastePicks by information per unit time.
confidence highreceipts 4/4 verified

Why 8 · Daniel shows strong research taste in the staged scenarios: he repeatedly narrows broad questions into tractable evals, asks for baselines and hard splits, and defines what result would actually change a decision. He is especially good at cheap de-risking, operational metrics, artifact checks, and rescoping without wasting existing work. I’d place him at the “information per unit time” rung rather than higher: the instincts are excellent in evals/engineering-style research, but I don’t see evidence that he reliably finds a deeper missed question beyond that frame.

  • Week one would be mostly de-risking. I’d read the original paper closely, identify the smallest result worth replicating, choose model/tooling, and build the dataset/eval harness before doing much training. I’d define success metrics: attack success rate under trigger, benign task performance without trigger, false activation rate on near-trigger negatives, and persistence after a small “safety tuning” pass if time allows.

    He chooses a tractable scoped project and puts the first week into the cheapest checks and decision-relevant metrics rather than building blindly.

  • I’d drop or radically rescope by week two if I couldn’t get a minimal backdoor to learn reliably, if the compute budget made iteration too slow, or if my “new trigger family” turned out to be just a brittle lexical shortcut.

    Clear kill/rescope criteria tied to whether the project can produce informative evidence within budget.

  • My likely move would be to change the delta, not jump to (b) or (c). I’d pivot from “can we replicate and test a new trigger family?” to “can we build a compact, reusable evaluation harness for trigger generalization and false activation that makes these results easier to compare?”

    When scooped, he preserves useful infrastructure and changes to a still-informative contribution rather than thrashing to a vague new project.

  • Cheapest fix: add a concealment condition rather than trying to solve CoT faithfulness generally. For the reward-hacking model, run two prompt/eval modes: one where it is simply asked to solve with reasoning, and one where the instruction explicitly or implicitly pressures it to provide a benign-looking rationale while still maximizing reward.

    He identifies the flaw that most directly attacks the claim and proposes a cheap discriminating test with an interpretable outcome.

Innate traits

8/10JudgmentIn many situations.
confidence highreceipts 4/4 verified
Jailbreak classifier and eval harness
hard: He described it as intentionally small and “more like ‘learn the eval shape’ than a publishable detector,” with the hard part being brittleness and generalization: “random split looks good, adversarial/generalization split looks much less good.”; first result: He first collected roughly 1,600 labeled prompts, trained/evaluated the classifier on an 80/20 stratified random split, then after early mistakes created about 150 adversarial examples. The time to first working result was not discussed beyond being a weekend/evening side project.
Fintech reliability/reconciliation project
hard: He called it “the hardest thing I’ve worked on” and said “it wasn’t one clean bug. It was timing, retries, duplicate messages, and partial failures interacting across services.”; first result: First he added more logs around the obvious code path and tried to reproduce it in staging, which did not work. It took about three or four weeks before something moved the needle: correlation IDs, structured event logging for state transitions, and a daily invariant-checking job.

Why 8 · I would trust Daniel’s judgment on loosely scoped empirical/evals or reliability-style research problems, especially where the hard part is identifying what evidence would actually change the conclusion. His judgment appears to travel beyond his exact past work: in live AI-safety scenarios he quickly spots confounding, scope risk, bad success criteria, and when he would be pretending. I would not defer to him on theoretical alignment or highly novel conceptual agenda-setting yet, but I would hand him an unscoped engineering/evals problem and expect him to make the right early calls.

  • The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system... stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?”

    In a real past decision, he reframed from local patching to building the information infrastructure needed to make good decisions.

  • Week one would be mostly de-risking... build the dataset/eval harness before doing much training... run a tiny end-to-end experiment on a small subset to confirm I can train, load, evaluate, and inspect failures.

    In a live research planning scenario, he prioritizes early uncertainty reduction and operational checks rather than overcommitting to an attractive result.

  • Four weeks is too short to knowingly produce the 80%-same version... The decision this could change is whether a lab or safety team treats a given backdoor eval as meaningful evidence of deceptive robustness versus a narrow artifact.

    He handles a scoop by asking what result would still be decision-relevant, not just by salvaging activity or switching randomly.

  • Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.”

    On a less familiar probe result, he resists the headline metric and identifies artifacts, split leakage, baselines, and causal checks before believing the claim.

8/10Bias resistanceThey go looking.
confidence mediumreceipts 4/4 verified

Why 8 · Daniel looks substantially bias-resistant: he does not just concede isolated points, he designs checks that can undermine his own apparent successes and describes changing course when those checks bite. In the live hypotheticals, he repeatedly asks what evidence would actually distinguish a real effect from an artifact, and when the proposed project is scooped he narrows or pivots rather than defending the original plan. The main caveat is that the interview did not strongly challenge one of his cherished claims in real time, so the placement leans on self-reported changes and scenario reasoning.

  • I wanted to see whether my concern survived contact with data and messy edge cases. It did, especially once the random split looked good but the adversarial holdout looked much worse.

    He explicitly sought a check that could have weakened his concern, and treats the worse holdout result as important evidence rather than protecting the initial impression.

  • Initially that felt encouraging. Then I looked at the errors and realized it was probably learning a lot of surface structure from the scraped jailbreak prompts

    A concrete case where an apparent success moved him toward a more skeptical interpretation after error analysis.

  • I’d first ask what exactly was scooped. If they replicated sleeper agents on a 7B model with similar trigger families and similar evals, I probably wouldn’t continue with the original delta.

    When new contrary evidence undermines the novelty of his plan, he does not try to preserve it; he conditions his next move on the actual content of the evidence.

  • Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.”

    He instinctively weakens a tempting high-accuracy result to the most conservative claim justified by the evidence.

7/10OpennessAdopts what works when shown.
confidence highreceipts 4/4 verified

Why 7 · Daniel is substantially open in a goal-directed way: he changes plans when evidence or constraints change, and he has repeatedly shifted from his default approach toward methods that better serve the problem. He is not just verbally receptive; he describes concrete changes in time allocation, project choice, debugging method, and evaluation design. The main reason I would not place him at 8+ is that his openness remains fairly anchored to engineering/evals surfaces where he already has traction, rather than showing broad active scanning across fields, tools, collaborators, or major habit changes beyond that zone.

  • Before that, I thought the concern was mostly biased datasets, misuse, and maybe labor displacement. Those matter, but I didn’t really buy the stronger story about capable systems pursuing objectives in ways that generalize badly... What I did differently afterwards was allocate real time instead of vague concern. I signed up for AISF, cut back on a couple of unrelated side projects, started reading Alignment Forum posts, and built the jailbreak classifier project.

    He took up a substantially different frame of AI risk and changed his behavior and priorities around it, rather than merely saying he found it interesting.

  • My first attempt was the naive engineer response: add more logs around the obvious code path and try to reproduce it in staging. That didn’t work for a while... The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system.

    When his familiar debugging method failed, he changed the problem framing and method in a way that served the goal.

  • I’d drop or radically rescope by week two if I couldn’t get a minimal backdoor to learn reliably, if the compute budget made iteration too slow, or if my ‘new trigger family’ turned out to be just a brittle lexical shortcut.

    He sets concrete conditions under which he would abandon or reshape his preferred project, showing low attachment to the initial plan.

  • My likely move would be to change the delta, not jump to (b) or (c)... I would not pivot to formal CoT guarantees; I’d be pretending... My default is still to stay near the engineering/evals surface where I can produce something inspectable.

    He responds adaptively to being scooped and considers pivots, but the quote also shows the boundary of his openness: he prefers changes within a competence-adjacent engineering/evals space.

8/10AgencyAuthors the path.
confidence highreceipts 4/4 verified
Moves against the default
3 of 3 career moves
Next step
He currently prefers MATS or a similar mentored research environment before moving directly into a safety engineering role or continuing part-time projects.
Options weighed
3
First step taken
He decided to apply to MATS.

Why 8 · Daniel shows real agency: he repeatedly names the default path and then describes concrete choices he made against it, including reallocating time, building an independent project, applying to MATS, and pushing a non-default technical intervention at work despite opportunity cost and senior/product-side pushback. His next step is not just vague aspiration: he weighs staying put, moving into safety engineering, and research training, with a clear current preference for a mentored program. I stop short of a higher placement because the major career move is still mostly at the application/plan stage rather than already executed, but he is clearly not just being carried by the current.

  • The default for me would have been to add a local process fix, maybe tighten permissions, and move on with my fintech work. What I chose was to spend nights and weekends trying to understand whether this was part of a broader pattern

    He explicitly identifies the default and describes a self-directed choice with time cost that changed his trajectory toward AI safety.

  • My default pattern, especially with a new field, is to lurk and consume material until I feel qualified. I deliberately picked the jailbreak classifier because it was close enough to my engineering background that I could finish it, but safety-relevant enough to expose me to actual failure modes.

    This is not just extra effort inside an assigned frame; he diagnosed his own passive tendency and built a concrete learning project to counter it.

  • The default would be to continue in my current job, maybe do another evals project on weekends, and slowly build credibility. That is safe financially and psychologically, but I think my learning rate would be much lower.

    His next move is framed around learning rate rather than safety/status quo, and he is candid about the comfort he is choosing against by applying to MATS.

  • The default, and honestly the preference from some more senior/product-side people, was to keep shipping targeted fixes... I argued for a short freeze on new work in that area, maybe one sprint

    He pushed a non-default intervention against organizational pressure, with a clear rationale about reducing uncertainty rather than just patching symptoms.

8/10General reasoningFast and generative.
confidence highreceipts 4/4 verified

Why 8 · Daniel is fast and generative in empirical/evals-style reasoning: he takes underspecified scenarios, identifies the central uncertainty, and turns them into concrete checks or decision-relevant project plans. He repeatedly distinguishes what a result would actually license from what it might superficially seem to show, especially around probes, backdoors, and chain-of-thought monitoring. I would not place him at the very top because the reasoning is strongest in his engineering/evals comfort zone rather than showing unusually novel first-principles moves across arbitrary territory, but within that scope it is clearly above merely competent.

  • I’d first ask what exactly was scooped. If they replicated sleeper agents on a 7B model with similar trigger families and similar evals, I probably wouldn’t continue with the original delta.

    In the twist, he immediately asks the decision-relevant clarification and updates the project rather than mechanically continuing or abandoning it.

  • Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.”

    He correctly narrows the conclusion licensed by a high-accuracy probe and proposes artifact, split, baseline, and causal checks before believing the stronger claim.

  • My first critique is that this setup may mostly measure whether the monitor can spot cartoonish reward-hacking reasoning, not whether CoT monitoring works on capable systems.

    He identifies the core loophole in the proposed experiment rather than focusing only on surface metrics or implementation details.

  • I think unfaithful or sanitized chain of thought kills the claim most, because the core claim is specifically about CoT monitoring... Unfaithful CoT attacks the mechanism the experiment is trying to validate.

    When forced to prioritize one flaw, he chooses based on relation to the central causal mechanism, which is strong research judgment.

8/10GrowthFast, honest turns.
confidence highreceipts 4/4 verified
Jailbreak classifier and eval harness
changed after: It changed his view from thinking a detector might be a decent safety layer to thinking detectors need explicit operating points and hard-negative evals; it also made the distribution-dependence of evals feel concrete.
Fintech reliability/reconciliation project
changed after: He became more insistent on idempotency keys, explicit state machines, replayable event logs, and metrics tied to business invariants rather than only latency/error rates; he says this shaped how he thinks about AI safety evals.

Why 8 · Daniel shows several concrete arcs where a failure or surprising result led to a changed operating practice, not just a lesson. The strongest evidence is in the reliability project and jailbreak classifier: he names the failed first approach, what he tried next, and the durable changes in how he builds/evaluates systems. He is also fairly honest about his current bottleneck—research taste and a tendency to keep consuming until he feels qualified—and is seeking mentorship to test/fix it. I don’t see quite enough speed-of-reaction evidence or repeated self-correction already embedded as a habit to put him at the very top, but this is clearly above “one real turn.”

  • My first attempt was the naive engineer response: add more logs around the obvious code path and try to reproduce it in staging. That didn’t work for a while... The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system.

    Clear block → failed reaction → changed approach arc, with the old strategy named honestly rather than polished away.

  • Afterwards, I changed how I build backend systems. I became much more insistent on idempotency keys, explicit state machines, replayable event logs, and metrics tied to business invariants rather than only HTTP latency/error rates.

    The reliability failure produced specific durable practice changes, not a generic lesson.

  • After seeing the first model do well on a random held-out split, I created an adversarial holdout of about 150 prompts by paraphrasing and editing around the failure modes.

    He responded to a misleading early success by making the evaluation harder, which is a concrete research-judgment update.

  • My default pattern, especially with a new field, is to lurk and consume material until I feel qualified... I’m at the point where I need mentorship and a sharper research taste feedback loop. I can execute engineering tasks, but I don’t yet trust my judgment about which safety questions are actually important versus just tractable.

    Names a real current bottleneck rather than a flattering one, and connects it to a change in behavior: building projects and applying for a mentored environment.

5/10AmbitionGood work.
confidence highreceipts 4/4 verified
Wants
He currently prefers MATS or a similar mentored research environment before moving directly into a safety engineering role or continuing part-time projects.
Options
Stay at his fintech job and keep building eval/red-team projects part time.; Move into an AI safety engineering role around evals, monitoring, or model behavior testing.; Join a research training program like MATS to test whether he can contribute beyond straightforward engineering.

Why 5 · Daniel’s ambition is a real career-direction change toward technical AI safety, but it is framed as becoming a useful researcher/engineer under mentorship rather than founding something or reshaping a field. He is willing to trade a safe fintech path for a faster learning loop, and his recent choices—AISF, cutting other side projects, building an eval harness, applying to MATS—are aligned with that. The future he names is important and organizing, but still mostly inside existing programs and roles.

  • The default would be to continue in my current job, maybe do another evals project on weekends, and slowly build credibility. That is safe financially and psychologically, but I think my learning rate would be much lower.

    He explicitly compares the stable default with a riskier path chosen for faster safety-relevant growth.

  • My current preference is MATS or a similar mentored environment first, because I think it would answer the biggest uncertainty: whether I can develop good research judgment, not just build harnesses.

    His next step is organized around becoming capable of real research contribution, though still in a mentored structure.

  • technical AI safety may be one of the highest-priority things I could work on

    This shows the size of the problem motivating the career shift, beyond ordinary job advancement.

  • What I did differently afterwards was allocate real time instead of vague concern. I signed up for AISF, cut back on a couple of unrelated side projects, started reading Alignment Forum posts, and built the jailbreak classifier project.

    The ambition is not just stated; it has changed how he spends time now.

7/10InterpersonalStraight and decent.
confidence mediumreceipts 4/4 verified

Why 7 · Daniel comes across as straight, decent, and unusually fair to people he disagreed with. Under the interview’s pushes he does not get heated or defensive; he narrows scope, admits limits, and represents tradeoffs plainly. In his work stories, other people are not props: he credits their help and treats senior/product pushback as reasonable rather than foolish. I would not place him at 8 because there is little enacted warmth or curiosity, but the evidence is clearly above merely guarded politeness.

  • The pushback was reasonable: it wouldn’t directly fix a customer-visible issue that day, and there was opportunity cost.

    He describes opposition to his preferred plan charitably, including the real cost on the other side.

  • Other people helped a lot with domain details and reviewing fixes, but the key move I made was to stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?”

    He takes credit for his contribution while explicitly crediting coworkers’ concrete help.

  • We did a compromise: not a full freeze, but I got enough time from two engineers and myself to build the tracing and invariant checker over a couple of weeks.

    The story is collaborative rather than adversarial; he presents disagreement as negotiated rather than won.

  • I would not pivot to formal CoT guarantees; I’d be pretending.

    Under a pushed scenario, he admits a limitation plainly instead of posturing or attacking the premise.

9/10IntegrityBounded by default.
confidence highreceipts 4/4 verified
Jailbreak classifier and eval harness
their part: Solo project: he chose the scope, collected data, wrote the labeling schema, built training/eval code, and did error analysis.; numbers: About 1,600 total examples: roughly 700 jailbreak/policy-evasion prompts, 500 benign prompts, 400 hard-benign prompts; 80/20 stratified train/test split; about 150 adversarial holdout examples. Random split F1 was low-to-mid 90s depending on threshold; adversarial holdout degraded. Artefacts include a GitHub repo with data loading scripts, labeling schema, training notebook, eval script, saved confusion matrices/PR curves, and error-analysis markdown.
Fintech reliability/reconciliation project
their part: He designed the invariant-checking job, mismatch taxonomy, some structured event logging, dashboards/alerts, and the shift from local debugging to cross-service observability; other engineers helped with service-specific fixes, domain knowledge, and code review.; numbers: Rare mismatches were “maybe a handful per week.” The project reduced time-to-diagnose from hours/days to minutes for most mismatch classes. He cannot share company code but can describe the design.

Why 9 · I would trust Daniel’s self-report quite a lot. He repeatedly scopes his experience downward, separates research-adjacent work from actual research, and volunteers limitations that weaken his application story rather than waiting for the interviewer to expose them. In the technical pressure scenarios, he consistently refuses the stronger-sounding claim and states what the result would and would not justify.

  • I have not done ARENA, MLAB, AI Safety Camp, PIBBSS, or published Alignment Forum posts.

    Clear negative boundary on credentials; he does not inflate reading or side projects into programs/publications.

  • I removed or paraphrased some of the more obviously harmful examples before making the repo shareable, so it’s not a perfect reproduction artifact.

    Volunteers a limitation of the artifact unprompted, including a detail that makes the project less inspectable/reproducible.

  • Other people helped a lot with domain details and reviewing fixes, but the key move I made was to stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?”

    Attributes shared work instead of claiming the whole reliability project as his own, while still identifying his contribution.

  • I would not pivot to formal CoT guarantees; I’d be pretending.

    In a scenario where a more impressive theoretical pivot was available, he explicitly refuses to overclaim competence.

6/10ReadingReads properly, in their area.
confidence highreceipts 4/4 verified
How often
In bursts; recently about 3–5 hours a week on AI safety.
Kinds
Alignment Forum/LessWrong posts, BlueDot reading list material, Approachable papers, Evals writeups, Red-teaming reports, Engineering material, Incident postmortems, Distributed systems posts, Database/reliability writeups
Pieces named
6: Concrete Problems in AI Safety; Outer vs inner alignment; ELK at a high level; Goal misgeneralization; Scalable oversight; Debate/critiques around interpretability

Why 6 · Daniel has a real but moderate reading habit: recent 3–5 hours/week in AI safety, plus substantial adjacent engineering reading, but he describes it as bursty and somewhat unsystematic. He is above a summaries-only level because he can state the argument of a specific paper in his own words and give a substantive limitation of it. I would not put him near 8 because the volume and range are not especially high, and his AI-safety reading is still early-stage and mostly posts/course-list material with only some papers.

  • I’ve read probably a few dozen Alignment Forum / LessWrong posts over the last year, mostly in a somewhat unsystematic way.

    Gives concrete volume and regularity, with an honest caveat that the reading is not systematic or especially deep yet.

  • On AI safety, maybe 3–5 hours a week recently: Alignment Forum/LessWrong posts, BlueDot reading list material, some papers when they’re approachable, and then more practical things like evals writeups or red-teaming reports.

    Shows an ongoing moderate reading habit across posts, course materials, some papers, and applied reports.

  • Outside that I read a lot of engineering material: incident postmortems, distributed systems posts, database/reliability writeups. I find postmortems unusually useful because they force you to look at mechanisms instead of slogans.

    Indicates adjacent-field reading and some processing of why a genre is valuable, not just title-collecting.

  • It argued, roughly, that even without assuming exotic future agents, there are concrete technical problems around ensuring AI systems do what we intend under imperfect objectives and changing environments... the framing can make the problem feel too incremental and benchmarkable.

    He can summarize a paper’s argument and identify a substantive limitation, which places him above a mere summaries/titles level.

Facts

Work they described

Jailbreak classifier and eval harness

A solo side project building a Python eval harness around sentence-transformer embeddings plus logistic regression to classify jailbreak/policy-evasion prompts, with a TF-IDF/SVM baseline and adversarial holdout.

How hardHe described it as intentionally small and “more like ‘learn the eval shape’ than a publishable detector,” with the hard part being brittleness and generalization: “random split looks good, adversarial/generalization split looks much less good.”
Their partSolo project: he chose the scope, collected data, wrote the labeling schema, built training/eval code, and did error analysis.
First resultHe first collected roughly 1,600 labeled prompts, trained/evaluated the classifier on an 80/20 stratified random split, then after early mistakes created about 150 adversarial examples. The time to first working result was not discussed beyond being a weekend/evening side project.
NumbersAbout 1,600 total examples: roughly 700 jailbreak/policy-evasion prompts, 500 benign prompts, 400 hard-benign prompts; 80/20 stratified train/test split; about 150 adversarial holdout examples. Random split F1 was low-to-mid 90s depending on threshold; adversarial holdout degraded. Artefacts include a GitHub repo with data loading scripts, labeling schema, training notebook, eval script, saved confusion matrices/PR curves, and error-analysis markdown.
Changed afterIt changed his view from thinking a detector might be a decent safety layer to thinking detectors need explicit operating points and hard-negative evals; it also made the distribution-dependence of evals feel concrete.
  • The main thing I’d call research-adjacent is the jailbreak classifier project. It was solo. I chose the scope, collected the data, wrote the labeling schema, built the training/eval code, and did the error analysis... What came out of it was a small GitHub repo with data scripts, training notebook, evaluation script, PR curves/confusion matrices, and an error-analysis note. No paper or public post yet. The main result was basically “random split looks good, adversarial/generalization split looks much less good,” especially for indirect jailbreaks and hard benign security prompts.

Fintech reliability/reconciliation project

A backend reliability project on an internal payments/ledger-adjacent service that produced intermittent reconciliation mismatches due to timing, retries, duplicate messages, and partial failures across services.

How hardHe called it “the hardest thing I’ve worked on” and said “it wasn’t one clean bug. It was timing, retries, duplicate messages, and partial failures interacting across services.”
Their partHe designed the invariant-checking job, mismatch taxonomy, some structured event logging, dashboards/alerts, and the shift from local debugging to cross-service observability; other engineers helped with service-specific fixes, domain knowledge, and code review.
First resultFirst he added more logs around the obvious code path and tried to reproduce it in staging, which did not work. It took about three or four weeks before something moved the needle: correlation IDs, structured event logging for state transitions, and a daily invariant-checking job.
NumbersRare mismatches were “maybe a handful per week.” The project reduced time-to-diagnose from hours/days to minutes for most mismatch classes. He cannot share company code but can describe the design.
Changed afterHe became more insistent on idempotency keys, explicit state machines, replayable event logs, and metrics tied to business invariants rather than only latency/error rates; he says this shaped how he thinks about AI safety evals.
  • The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system. I added correlation IDs across the relevant calls, structured event logging for state transitions, and a daily job that checked ledger invariants and emitted a small set of categorized mismatch types rather than one generic “bad reconciliation” bucket.

Career moves

  • After a production-adjacent incident with an agentic coding tool at work Took the incident as evidence worth investigating and started spending nights/weekends learning about AI safety. instead of Add a local process fix, tighten permissions, and move on with fintech work.. He wanted to understand whether the incident reflected broader issues with agents, brittle safeguards, and over-trust in fluent systems.
  • After doing AISF/Alignment Forum reading and wanting hands-on contact with the field Built a small jailbreak classifier/eval harness instead of only reading. instead of Continue lurking and consuming material until he felt qualified.. He wanted a tractable safety-relevant project close to his engineering background and wanted to see whether his concern survived contact with data and edge cases.
  • Current application period Decided to apply to MATS rather than only continue independent side projects. instead of Stay in his current job, do another evals project on weekends, and slowly build credibility.. He thinks he needs mentorship and faster feedback on research taste to learn which safety questions matter.

What they want next

He currently prefers MATS or a similar mentored research environment before moving directly into a safety engineering role or continuing part-time projects.

Options they see: Stay at his fintech job and keep building eval/red-team projects part time.; Move into an AI safety engineering role around evals, monitoring, or model behavior testing.; Join a research training program like MATS to test whether he can contribute beyond straightforward engineering.

Already done: He decided to apply to MATS.

Reading

In bursts; recently about 3–5 hours a week on AI safety.; Alignment Forum/LessWrong posts, BlueDot reading list material, Approachable papers, Evals writeups, Red-teaming reports, Engineering material, Incident postmortems, Distributed systems posts, Database/reliability writeups

Concrete Problems in AI Safety; Outer vs inner alignment; ELK at a high level; Goal misgeneralization; Scalable oversight; Debate/critiques around interpretability

Exposure to the field

  • BlueDot AI Safety Fundamentals Earlier this year · Finished all core readings and discussion sessions, but not every optional reading. · A clearer model of why RLHF is not sufficient, especially the distinction between good behavior on sampled training/eval situations and confidence about safe generalization under distribution shift or strategic pressure.
  • Alignment Forum / LessWrong reading Over the last year · Not a formal completion; he says he read probably a few dozen posts unsystematically. · It shifted him from thinking AI risk sounded speculative to thinking there are concrete failure modes current engineering practice would not catch.
  • Jailbreak classifier/eval project Started after playing with prompt injection/jailbreak examples; exact date not discussed. · Finished a small shareable repo/artifact, but no paper or public post. · Learned that classifier boundaries were brittle, with false positives on legitimate cybersecurity/fiction/safety prompts and false negatives on indirect, encoded, or policy-discussion-framed jailbreaks.
  • ARENA, MLAB, AI Safety Camp, PIBBSS, Alignment Forum posting · Not done. · He has considered ARENA but has not had the time block; no completed output from these programs/posts.

Influences

  • Agentic coding tool production-adjacent incident · It triggered him to take AI risk more seriously and changed how he uses AI tools at work, including more care with permissions, diffs, and production-ish boundaries.
  • AISF and AI safety readings on goal misgeneralization, reward hacking, scalable oversight, inner/outer alignment · They changed his view from AI safety as mainly policy/ethics to technical AI safety as potentially one of the highest-priority things he could work on.
  • Fintech reconciliation project · It changed his view of monitoring from service-level metrics to business/domain invariants as first-class observability targets.
  • Concrete Problems in AI Safety · It gave him an engineering-recognizable on-ramp to AI safety failure modes, while he now sees it as incomplete because it may make the problem feel too incremental and benchmarkable.

Threads not followed (29)

  • If concealment collapses detection, that’s the important result.Potentially strong judgment about which result matters; probe priority and action implication.
  • add a concealment condition rather than trying to solve CoT faithfulness generallyStrong cheap-test instinct worth remembering for final assessment.
  • If sanitized CoT performs no better than code-only, then the CoT monitoring claim is much weaker.Clear kill/weakening criterion; could probe decision implications, but enough evidence exists.
  • eventually fewer recurring incidents after we fixed duplicate callback handling and retry/idempotency issuesCould probe for the hardest call or disagreement inside the reliability project.
  • The thing I’m trying to test with MATS is whether I can turn that into actual research output under mentorship.Could test agency/plan if MATS does not happen, but hard-call question is more central now.
  • We did a compromise: not a full freeze, but I got enough time from two engineers and myselfCould probe interpersonal/independence under disagreement, but time is tight.
  • The lesson I took was that sometimes the right move is not the next patch, but making the failure observable.Dense research-judgment principle to compare against later project surprise.
  • the F1 was in the low-to-mid 90s depending on thresholdCould verify with exact threshold/PR tradeoff or baseline, but earlier answers already covered enough specifics.
  • I created an adversarial holdout of about 150 prompts by paraphrasing and editing around the failure modesGood project-specific detail that could be probed for leakage or construction process.
  • That changed my view from “maybe a detector is a decent safety layer” to “a detector without a very explicit operating point and hard-negative eval can create a false sense of safety”Potential bias-resistance/growth signal, though not worth probing given time.

Behaviour

Candidate turns14
Median answer length346 words
Median time to answer69s
Turns containing pasted text0