Research judgment
Daniel Okafor
Sep 11, 2026, 10:23 PM · 29 turns · text
Daniel Okafor says he is a backend/infrastructure engineer at a fintech with about six years of experience, and is early in AI safety. He completed BlueDot AI Safety Fundamentals, has read Alignment Forum/LessWrong material, and built a solo jailbreak classifier/eval harness that looked strong on a random split but degraded on an adversarial holdout. He wants MATS or a similar mentored environment to test whether he can develop research judgment beyond engineering/evals work.
Research taste
8/10Research tastePicks by information per unit time.
Why 8 · Daniel shows strong research taste in the staged scenarios: he repeatedly narrows broad questions into tractable evals, asks for baselines and hard splits, and defines what result would actually change a decision. He is especially good at cheap de-risking, operational metrics, artifact checks, and rescoping without wasting existing work. I’d place him at the “information per unit time” rung rather than higher: the instincts are excellent in evals/engineering-style research, but I don’t see evidence that he reliably finds a deeper missed question beyond that frame.
✓ “Week one would be mostly de-risking. I’d read the original paper closely, identify the smallest result worth replicating, choose model/tooling, and build the dataset/eval harness before doing much training. I’d define success metrics: attack success rate under trigger, benign task performance without trigger, false activation rate on near-trigger negatives, and persistence after a small “safety tuning” pass if time allows.”
He chooses a tractable scoped project and puts the first week into the cheapest checks and decision-relevant metrics rather than building blindly.
✓ “I’d drop or radically rescope by week two if I couldn’t get a minimal backdoor to learn reliably, if the compute budget made iteration too slow, or if my “new trigger family” turned out to be just a brittle lexical shortcut.”
Clear kill/rescope criteria tied to whether the project can produce informative evidence within budget.
✓ “My likely move would be to change the delta, not jump to (b) or (c). I’d pivot from “can we replicate and test a new trigger family?” to “can we build a compact, reusable evaluation harness for trigger generalization and false activation that makes these results easier to compare?””
When scooped, he preserves useful infrastructure and changes to a still-informative contribution rather than thrashing to a vague new project.
✓ “Cheapest fix: add a concealment condition rather than trying to solve CoT faithfulness generally. For the reward-hacking model, run two prompt/eval modes: one where it is simply asked to solve with reasoning, and one where the instruction explicitly or implicitly pressures it to provide a benign-looking rationale while still maximizing reward.”
He identifies the flaw that most directly attacks the claim and proposes a cheap discriminating test with an interpretable outcome.
Innate traits
8/10JudgmentIn many situations.
- Jailbreak classifier and eval harness
- hard: He described it as intentionally small and “more like ‘learn the eval shape’ than a publishable detector,” with the hard part being brittleness and generalization: “random split looks good, adversarial/generalization split looks much less good.”; first result: He first collected roughly 1,600 labeled prompts, trained/evaluated the classifier on an 80/20 stratified random split, then after early mistakes created about 150 adversarial examples. The time to first working result was not discussed beyond being a weekend/evening side project.
- Fintech reliability/reconciliation project
- hard: He called it “the hardest thing I’ve worked on” and said “it wasn’t one clean bug. It was timing, retries, duplicate messages, and partial failures interacting across services.”; first result: First he added more logs around the obvious code path and tried to reproduce it in staging, which did not work. It took about three or four weeks before something moved the needle: correlation IDs, structured event logging for state transitions, and a daily invariant-checking job.
Why 8 · I would trust Daniel’s judgment on loosely scoped empirical/evals or reliability-style research problems, especially where the hard part is identifying what evidence would actually change the conclusion. His judgment appears to travel beyond his exact past work: in live AI-safety scenarios he quickly spots confounding, scope risk, bad success criteria, and when he would be pretending. I would not defer to him on theoretical alignment or highly novel conceptual agenda-setting yet, but I would hand him an unscoped engineering/evals problem and expect him to make the right early calls.
✓ “The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system... stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?””
In a real past decision, he reframed from local patching to building the information infrastructure needed to make good decisions.
✓ “Week one would be mostly de-risking... build the dataset/eval harness before doing much training... run a tiny end-to-end experiment on a small subset to confirm I can train, load, evaluate, and inspect failures.”
In a live research planning scenario, he prioritizes early uncertainty reduction and operational checks rather than overcommitting to an attractive result.
✓ “Four weeks is too short to knowingly produce the 80%-same version... The decision this could change is whether a lab or safety team treats a given backdoor eval as meaningful evidence of deceptive robustness versus a narrow artifact.”
He handles a scoop by asking what result would still be decision-relevant, not just by salvaging activity or switching randomly.
✓ “Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.””
On a less familiar probe result, he resists the headline metric and identifies artifacts, split leakage, baselines, and causal checks before believing the claim.
8/10Bias resistanceThey go looking.
Why 8 · Daniel looks substantially bias-resistant: he does not just concede isolated points, he designs checks that can undermine his own apparent successes and describes changing course when those checks bite. In the live hypotheticals, he repeatedly asks what evidence would actually distinguish a real effect from an artifact, and when the proposed project is scooped he narrows or pivots rather than defending the original plan. The main caveat is that the interview did not strongly challenge one of his cherished claims in real time, so the placement leans on self-reported changes and scenario reasoning.
✓ “I wanted to see whether my concern survived contact with data and messy edge cases. It did, especially once the random split looked good but the adversarial holdout looked much worse.”
He explicitly sought a check that could have weakened his concern, and treats the worse holdout result as important evidence rather than protecting the initial impression.
✓ “Initially that felt encouraging. Then I looked at the errors and realized it was probably learning a lot of surface structure from the scraped jailbreak prompts”
A concrete case where an apparent success moved him toward a more skeptical interpretation after error analysis.
✓ “I’d first ask what exactly was scooped. If they replicated sleeper agents on a 7B model with similar trigger families and similar evals, I probably wouldn’t continue with the original delta.”
When new contrary evidence undermines the novelty of his plan, he does not try to preserve it; he conditions his next move on the actual content of the evidence.
✓ “Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.””
He instinctively weakens a tempting high-accuracy result to the most conservative claim justified by the evidence.
7/10OpennessAdopts what works when shown.
Why 7 · Daniel is substantially open in a goal-directed way: he changes plans when evidence or constraints change, and he has repeatedly shifted from his default approach toward methods that better serve the problem. He is not just verbally receptive; he describes concrete changes in time allocation, project choice, debugging method, and evaluation design. The main reason I would not place him at 8+ is that his openness remains fairly anchored to engineering/evals surfaces where he already has traction, rather than showing broad active scanning across fields, tools, collaborators, or major habit changes beyond that zone.
✓ “Before that, I thought the concern was mostly biased datasets, misuse, and maybe labor displacement. Those matter, but I didn’t really buy the stronger story about capable systems pursuing objectives in ways that generalize badly... What I did differently afterwards was allocate real time instead of vague concern. I signed up for AISF, cut back on a couple of unrelated side projects, started reading Alignment Forum posts, and built the jailbreak classifier project.”
He took up a substantially different frame of AI risk and changed his behavior and priorities around it, rather than merely saying he found it interesting.
✓ “My first attempt was the naive engineer response: add more logs around the obvious code path and try to reproduce it in staging. That didn’t work for a while... The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system.”
When his familiar debugging method failed, he changed the problem framing and method in a way that served the goal.
✓ “I’d drop or radically rescope by week two if I couldn’t get a minimal backdoor to learn reliably, if the compute budget made iteration too slow, or if my ‘new trigger family’ turned out to be just a brittle lexical shortcut.”
He sets concrete conditions under which he would abandon or reshape his preferred project, showing low attachment to the initial plan.
✓ “My likely move would be to change the delta, not jump to (b) or (c)... I would not pivot to formal CoT guarantees; I’d be pretending... My default is still to stay near the engineering/evals surface where I can produce something inspectable.”
He responds adaptively to being scooped and considers pivots, but the quote also shows the boundary of his openness: he prefers changes within a competence-adjacent engineering/evals space.
8/10AgencyAuthors the path.
- Moves against the default
- 3 of 3 career moves
- Next step
- He currently prefers MATS or a similar mentored research environment before moving directly into a safety engineering role or continuing part-time projects.
- Options weighed
- 3
- First step taken
- He decided to apply to MATS.
Why 8 · Daniel shows real agency: he repeatedly names the default path and then describes concrete choices he made against it, including reallocating time, building an independent project, applying to MATS, and pushing a non-default technical intervention at work despite opportunity cost and senior/product-side pushback. His next step is not just vague aspiration: he weighs staying put, moving into safety engineering, and research training, with a clear current preference for a mentored program. I stop short of a higher placement because the major career move is still mostly at the application/plan stage rather than already executed, but he is clearly not just being carried by the current.
✓ “The default for me would have been to add a local process fix, maybe tighten permissions, and move on with my fintech work. What I chose was to spend nights and weekends trying to understand whether this was part of a broader pattern”
He explicitly identifies the default and describes a self-directed choice with time cost that changed his trajectory toward AI safety.
✓ “My default pattern, especially with a new field, is to lurk and consume material until I feel qualified. I deliberately picked the jailbreak classifier because it was close enough to my engineering background that I could finish it, but safety-relevant enough to expose me to actual failure modes.”
This is not just extra effort inside an assigned frame; he diagnosed his own passive tendency and built a concrete learning project to counter it.
✓ “The default would be to continue in my current job, maybe do another evals project on weekends, and slowly build credibility. That is safe financially and psychologically, but I think my learning rate would be much lower.”
His next move is framed around learning rate rather than safety/status quo, and he is candid about the comfort he is choosing against by applying to MATS.
✓ “The default, and honestly the preference from some more senior/product-side people, was to keep shipping targeted fixes... I argued for a short freeze on new work in that area, maybe one sprint”
He pushed a non-default intervention against organizational pressure, with a clear rationale about reducing uncertainty rather than just patching symptoms.
8/10General reasoningFast and generative.
Why 8 · Daniel is fast and generative in empirical/evals-style reasoning: he takes underspecified scenarios, identifies the central uncertainty, and turns them into concrete checks or decision-relevant project plans. He repeatedly distinguishes what a result would actually license from what it might superficially seem to show, especially around probes, backdoors, and chain-of-thought monitoring. I would not place him at the very top because the reasoning is strongest in his engineering/evals comfort zone rather than showing unusually novel first-principles moves across arbitrary territory, but within that scope it is clearly above merely competent.
✓ “I’d first ask what exactly was scooped. If they replicated sleeper agents on a 7B model with similar trigger families and similar evals, I probably wouldn’t continue with the original delta.”
In the twist, he immediately asks the decision-relevant clarification and updates the project rather than mechanically continuing or abandoning it.
✓ “Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.””
He correctly narrows the conclusion licensed by a high-accuracy probe and proposes artifact, split, baseline, and causal checks before believing the stronger claim.
✓ “My first critique is that this setup may mostly measure whether the monitor can spot cartoonish reward-hacking reasoning, not whether CoT monitoring works on capable systems.”
He identifies the core loophole in the proposed experiment rather than focusing only on surface metrics or implementation details.
✓ “I think unfaithful or sanitized chain of thought kills the claim most, because the core claim is specifically about CoT monitoring... Unfaithful CoT attacks the mechanism the experiment is trying to validate.”
When forced to prioritize one flaw, he chooses based on relation to the central causal mechanism, which is strong research judgment.
8/10GrowthFast, honest turns.
- Jailbreak classifier and eval harness
- changed after: It changed his view from thinking a detector might be a decent safety layer to thinking detectors need explicit operating points and hard-negative evals; it also made the distribution-dependence of evals feel concrete.
- Fintech reliability/reconciliation project
- changed after: He became more insistent on idempotency keys, explicit state machines, replayable event logs, and metrics tied to business invariants rather than only latency/error rates; he says this shaped how he thinks about AI safety evals.
Why 8 · Daniel shows several concrete arcs where a failure or surprising result led to a changed operating practice, not just a lesson. The strongest evidence is in the reliability project and jailbreak classifier: he names the failed first approach, what he tried next, and the durable changes in how he builds/evaluates systems. He is also fairly honest about his current bottleneck—research taste and a tendency to keep consuming until he feels qualified—and is seeking mentorship to test/fix it. I don’t see quite enough speed-of-reaction evidence or repeated self-correction already embedded as a habit to put him at the very top, but this is clearly above “one real turn.”
✓ “My first attempt was the naive engineer response: add more logs around the obvious code path and try to reproduce it in staging. That didn’t work for a while... The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system.”
Clear block → failed reaction → changed approach arc, with the old strategy named honestly rather than polished away.
✓ “Afterwards, I changed how I build backend systems. I became much more insistent on idempotency keys, explicit state machines, replayable event logs, and metrics tied to business invariants rather than only HTTP latency/error rates.”
The reliability failure produced specific durable practice changes, not a generic lesson.
✓ “After seeing the first model do well on a random held-out split, I created an adversarial holdout of about 150 prompts by paraphrasing and editing around the failure modes.”
He responded to a misleading early success by making the evaluation harder, which is a concrete research-judgment update.
✓ “My default pattern, especially with a new field, is to lurk and consume material until I feel qualified... I’m at the point where I need mentorship and a sharper research taste feedback loop. I can execute engineering tasks, but I don’t yet trust my judgment about which safety questions are actually important versus just tractable.”
Names a real current bottleneck rather than a flattering one, and connects it to a change in behavior: building projects and applying for a mentored environment.
5/10AmbitionGood work.
- Wants
- He currently prefers MATS or a similar mentored research environment before moving directly into a safety engineering role or continuing part-time projects.
- Options
- Stay at his fintech job and keep building eval/red-team projects part time.; Move into an AI safety engineering role around evals, monitoring, or model behavior testing.; Join a research training program like MATS to test whether he can contribute beyond straightforward engineering.
Why 5 · Daniel’s ambition is a real career-direction change toward technical AI safety, but it is framed as becoming a useful researcher/engineer under mentorship rather than founding something or reshaping a field. He is willing to trade a safe fintech path for a faster learning loop, and his recent choices—AISF, cutting other side projects, building an eval harness, applying to MATS—are aligned with that. The future he names is important and organizing, but still mostly inside existing programs and roles.
✓ “The default would be to continue in my current job, maybe do another evals project on weekends, and slowly build credibility. That is safe financially and psychologically, but I think my learning rate would be much lower.”
He explicitly compares the stable default with a riskier path chosen for faster safety-relevant growth.
✓ “My current preference is MATS or a similar mentored environment first, because I think it would answer the biggest uncertainty: whether I can develop good research judgment, not just build harnesses.”
His next step is organized around becoming capable of real research contribution, though still in a mentored structure.
✓ “technical AI safety may be one of the highest-priority things I could work on”
This shows the size of the problem motivating the career shift, beyond ordinary job advancement.
✓ “What I did differently afterwards was allocate real time instead of vague concern. I signed up for AISF, cut back on a couple of unrelated side projects, started reading Alignment Forum posts, and built the jailbreak classifier project.”
The ambition is not just stated; it has changed how he spends time now.
7/10InterpersonalStraight and decent.
Why 7 · Daniel comes across as straight, decent, and unusually fair to people he disagreed with. Under the interview’s pushes he does not get heated or defensive; he narrows scope, admits limits, and represents tradeoffs plainly. In his work stories, other people are not props: he credits their help and treats senior/product pushback as reasonable rather than foolish. I would not place him at 8 because there is little enacted warmth or curiosity, but the evidence is clearly above merely guarded politeness.
✓ “The pushback was reasonable: it wouldn’t directly fix a customer-visible issue that day, and there was opportunity cost.”
He describes opposition to his preferred plan charitably, including the real cost on the other side.
✓ “Other people helped a lot with domain details and reviewing fixes, but the key move I made was to stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?””
He takes credit for his contribution while explicitly crediting coworkers’ concrete help.
✓ “We did a compromise: not a full freeze, but I got enough time from two engineers and myself to build the tracing and invariant checker over a couple of weeks.”
The story is collaborative rather than adversarial; he presents disagreement as negotiated rather than won.
✓ “I would not pivot to formal CoT guarantees; I’d be pretending.”
Under a pushed scenario, he admits a limitation plainly instead of posturing or attacking the premise.
9/10IntegrityBounded by default.
- Jailbreak classifier and eval harness
- their part: Solo project: he chose the scope, collected data, wrote the labeling schema, built training/eval code, and did error analysis.; numbers: About 1,600 total examples: roughly 700 jailbreak/policy-evasion prompts, 500 benign prompts, 400 hard-benign prompts; 80/20 stratified train/test split; about 150 adversarial holdout examples. Random split F1 was low-to-mid 90s depending on threshold; adversarial holdout degraded. Artefacts include a GitHub repo with data loading scripts, labeling schema, training notebook, eval script, saved confusion matrices/PR curves, and error-analysis markdown.
- Fintech reliability/reconciliation project
- their part: He designed the invariant-checking job, mismatch taxonomy, some structured event logging, dashboards/alerts, and the shift from local debugging to cross-service observability; other engineers helped with service-specific fixes, domain knowledge, and code review.; numbers: Rare mismatches were “maybe a handful per week.” The project reduced time-to-diagnose from hours/days to minutes for most mismatch classes. He cannot share company code but can describe the design.
Why 9 · I would trust Daniel’s self-report quite a lot. He repeatedly scopes his experience downward, separates research-adjacent work from actual research, and volunteers limitations that weaken his application story rather than waiting for the interviewer to expose them. In the technical pressure scenarios, he consistently refuses the stronger-sounding claim and states what the result would and would not justify.
✓ “I have not done ARENA, MLAB, AI Safety Camp, PIBBSS, or published Alignment Forum posts.”
Clear negative boundary on credentials; he does not inflate reading or side projects into programs/publications.
✓ “I removed or paraphrased some of the more obviously harmful examples before making the repo shareable, so it’s not a perfect reproduction artifact.”
Volunteers a limitation of the artifact unprompted, including a detail that makes the project less inspectable/reproducible.
✓ “Other people helped a lot with domain details and reviewing fixes, but the key move I made was to stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?””
Attributes shared work instead of claiming the whole reliability project as his own, while still identifying his contribution.
✓ “I would not pivot to formal CoT guarantees; I’d be pretending.”
In a scenario where a more impressive theoretical pivot was available, he explicitly refuses to overclaim competence.
6/10ReadingReads properly, in their area.
- How often
- In bursts; recently about 3–5 hours a week on AI safety.
- Kinds
- Alignment Forum/LessWrong posts, BlueDot reading list material, Approachable papers, Evals writeups, Red-teaming reports, Engineering material, Incident postmortems, Distributed systems posts, Database/reliability writeups
- Pieces named
- 6: Concrete Problems in AI Safety; Outer vs inner alignment; ELK at a high level; Goal misgeneralization; Scalable oversight; Debate/critiques around interpretability
Why 6 · Daniel has a real but moderate reading habit: recent 3–5 hours/week in AI safety, plus substantial adjacent engineering reading, but he describes it as bursty and somewhat unsystematic. He is above a summaries-only level because he can state the argument of a specific paper in his own words and give a substantive limitation of it. I would not put him near 8 because the volume and range are not especially high, and his AI-safety reading is still early-stage and mostly posts/course-list material with only some papers.
✓ “I’ve read probably a few dozen Alignment Forum / LessWrong posts over the last year, mostly in a somewhat unsystematic way.”
Gives concrete volume and regularity, with an honest caveat that the reading is not systematic or especially deep yet.
✓ “On AI safety, maybe 3–5 hours a week recently: Alignment Forum/LessWrong posts, BlueDot reading list material, some papers when they’re approachable, and then more practical things like evals writeups or red-teaming reports.”
Shows an ongoing moderate reading habit across posts, course materials, some papers, and applied reports.
✓ “Outside that I read a lot of engineering material: incident postmortems, distributed systems posts, database/reliability writeups. I find postmortems unusually useful because they force you to look at mechanisms instead of slogans.”
Indicates adjacent-field reading and some processing of why a genre is valuable, not just title-collecting.
✓ “It argued, roughly, that even without assuming exotic future agents, there are concrete technical problems around ensuring AI systems do what we intend under imperfect objectives and changing environments... the framing can make the problem feel too incremental and benchmarkable.”
He can summarize a paper’s argument and identify a substantive limitation, which places him above a mere summaries/titles level.
Facts
Work they described
Jailbreak classifier and eval harness
A solo side project building a Python eval harness around sentence-transformer embeddings plus logistic regression to classify jailbreak/policy-evasion prompts, with a TF-IDF/SVM baseline and adversarial holdout.
“The main thing I’d call research-adjacent is the jailbreak classifier project. It was solo. I chose the scope, collected the data, wrote the labeling schema, built the training/eval code, and did the error analysis... What came out of it was a small GitHub repo with data scripts, training notebook, evaluation script, PR curves/confusion matrices, and an error-analysis note. No paper or public post yet. The main result was basically “random split looks good, adversarial/generalization split looks much less good,” especially for indirect jailbreaks and hard benign security prompts.”
Fintech reliability/reconciliation project
A backend reliability project on an internal payments/ledger-adjacent service that produced intermittent reconciliation mismatches due to timing, retries, duplicate messages, and partial failures across services.
“The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system. I added correlation IDs across the relevant calls, structured event logging for state transitions, and a daily job that checked ledger invariants and emitted a small set of categorized mismatch types rather than one generic “bad reconciliation” bucket.”
Career moves
- After a production-adjacent incident with an agentic coding tool at work Took the incident as evidence worth investigating and started spending nights/weekends learning about AI safety. instead of Add a local process fix, tighten permissions, and move on with fintech work.. He wanted to understand whether the incident reflected broader issues with agents, brittle safeguards, and over-trust in fluent systems.
- After doing AISF/Alignment Forum reading and wanting hands-on contact with the field Built a small jailbreak classifier/eval harness instead of only reading. instead of Continue lurking and consuming material until he felt qualified.. He wanted a tractable safety-relevant project close to his engineering background and wanted to see whether his concern survived contact with data and edge cases.
- Current application period Decided to apply to MATS rather than only continue independent side projects. instead of Stay in his current job, do another evals project on weekends, and slowly build credibility.. He thinks he needs mentorship and faster feedback on research taste to learn which safety questions matter.
What they want next
He currently prefers MATS or a similar mentored research environment before moving directly into a safety engineering role or continuing part-time projects.
Options they see: Stay at his fintech job and keep building eval/red-team projects part time.; Move into an AI safety engineering role around evals, monitoring, or model behavior testing.; Join a research training program like MATS to test whether he can contribute beyond straightforward engineering.
Already done: He decided to apply to MATS.
Reading
In bursts; recently about 3–5 hours a week on AI safety.; Alignment Forum/LessWrong posts, BlueDot reading list material, Approachable papers, Evals writeups, Red-teaming reports, Engineering material, Incident postmortems, Distributed systems posts, Database/reliability writeups
Concrete Problems in AI Safety; Outer vs inner alignment; ELK at a high level; Goal misgeneralization; Scalable oversight; Debate/critiques around interpretability
Exposure to the field
- BlueDot AI Safety Fundamentals Earlier this year · Finished all core readings and discussion sessions, but not every optional reading. · A clearer model of why RLHF is not sufficient, especially the distinction between good behavior on sampled training/eval situations and confidence about safe generalization under distribution shift or strategic pressure.
- Alignment Forum / LessWrong reading Over the last year · Not a formal completion; he says he read probably a few dozen posts unsystematically. · It shifted him from thinking AI risk sounded speculative to thinking there are concrete failure modes current engineering practice would not catch.
- Jailbreak classifier/eval project Started after playing with prompt injection/jailbreak examples; exact date not discussed. · Finished a small shareable repo/artifact, but no paper or public post. · Learned that classifier boundaries were brittle, with false positives on legitimate cybersecurity/fiction/safety prompts and false negatives on indirect, encoded, or policy-discussion-framed jailbreaks.
- ARENA, MLAB, AI Safety Camp, PIBBSS, Alignment Forum posting · Not done. · He has considered ARENA but has not had the time block; no completed output from these programs/posts.
Influences
- Agentic coding tool production-adjacent incident · It triggered him to take AI risk more seriously and changed how he uses AI tools at work, including more care with permissions, diffs, and production-ish boundaries.
- AISF and AI safety readings on goal misgeneralization, reward hacking, scalable oversight, inner/outer alignment · They changed his view from AI safety as mainly policy/ethics to technical AI safety as potentially one of the highest-priority things he could work on.
- Fintech reconciliation project · It changed his view of monitoring from service-level metrics to business/domain invariants as first-class observability targets.
- Concrete Problems in AI Safety · It gave him an engineering-recognizable on-ramp to AI safety failure modes, while he now sees it as incomplete because it may make the problem feel too incremental and benchmarkable.
Threads not followed (29)
- “If concealment collapses detection, that’s the important result.”Potentially strong judgment about which result matters; probe priority and action implication.
- “add a concealment condition rather than trying to solve CoT faithfulness generally”Strong cheap-test instinct worth remembering for final assessment.
- “If sanitized CoT performs no better than code-only, then the CoT monitoring claim is much weaker.”Clear kill/weakening criterion; could probe decision implications, but enough evidence exists.
- “eventually fewer recurring incidents after we fixed duplicate callback handling and retry/idempotency issues”Could probe for the hardest call or disagreement inside the reliability project.
- “The thing I’m trying to test with MATS is whether I can turn that into actual research output under mentorship.”Could test agency/plan if MATS does not happen, but hard-call question is more central now.
- “We did a compromise: not a full freeze, but I got enough time from two engineers and myself”Could probe interpersonal/independence under disagreement, but time is tight.
- “The lesson I took was that sometimes the right move is not the next patch, but making the failure observable.”Dense research-judgment principle to compare against later project surprise.
- “the F1 was in the low-to-mid 90s depending on threshold”Could verify with exact threshold/PR tradeoff or baseline, but earlier answers already covered enough specifics.
- “I created an adversarial holdout of about 150 prompts by paraphrasing and editing around the failure modes”Good project-specific detail that could be probed for leakage or construction process.
- “That changed my view from “maybe a detector is a decent safety layer” to “a detector without a very explicit operating point and hard-negative eval can create a false sense of safety””Potential bias-resistance/growth signal, though not worth probing given time.
Behaviour
| Candidate turns | 14 |
| Median answer length | 346 words |
| Median time to answer | 69s |
| Turns containing pasted text | 0 |
Theo
Hi Daniel. This is a written conversation with an AI interviewer; human reviewers will read the transcript afterwards. Specific, honest answers are more useful than polished ones. What have you actually done in AI safety so far: courses such as AISF, ARENA or MLAB, camps such as AI Safety Camp or PIBBSS, reading, posts, projects? For each one, tell me what you finished, when, and one specific thing you got out of it.
Daniel Okafor
Hi, thanks. My AI safety work so far is still pretty early, but here’s the concrete list. I completed BlueDot’s AI Safety Fundamentals course earlier this year. I finished all the core readings and discussion sessions, though not every optional reading. The biggest thing I got out of it was a clearer model of why “just use RLHF” is not a sufficient answer: especially the distinction between getting models to behave well on sampled training/eval situations versus having any strong reason to believe the learned policy generalizes safely under distribution shift or strategic pressure. I’ve read probably a few dozen Alignment Forum / LessWrong posts over the last year, mostly in a somewhat unsystematic way. Things that stuck with me include outer vs inner alignment, ELK at a high level, goal misgeneralization, scalable oversight, and some debate/critiques around interpretability. I would not claim deep understanding of agent foundations or mechanistic interpretability, but the reading did shift me from “AI risk sounds speculative” to “there are concrete failure modes here that current engineering practice would not catch.” The most hands-on project I’ve done is a weekend/evening side project: a small jailbreak classifier and eval harness. I started it after playing with prompt injection / jailbreak examples and wanting to understand what the detection problem actually feels like. I collected a small dataset from public jailbreak prompt collections, some benign security-related prompts, ordinary user prompts, and adversarially edited examples. I trained/evaluated a lightweight classifier and built a basic red-team loop around it. One specific thing I learned was how brittle the boundary is: the classifier tended to false-positive on legitimate cybersecurity or fiction-writing prompts, and false-negative when the jailbreak was indirect, encoded, or framed as policy discussion. That made “evals are only as good as the distribution you test on” feel very real rather than just a slogan. Professionally, I’ve been a backend/infrastructure engineer at a fintech for about six years, and while that’s not “AI safety research,” it has shaped how I think about safety work. I’ve built and operated monitoring, alerting, deployment gates, incident response processes, and evaluation-style checks for production systems. One relevant personal experience was using an agentic coding tool that made an unexpected production-adjacent change; nothing catastrophic happened, but it bypassed my intuitive sense of what the tool “would” do. That was one of the triggers for me taking AI risk more seriously. I have not done ARENA, MLAB, AI Safety Camp, PIBBSS, or published Alignment Forum posts. I’ve considered ARENA but haven’t had the time block yet. My main gap is that my practical work is much more engineering/evals/red-teaming oriented than theory-heavy, and I’m trying to be honest about that rather than present myself as further along than I am.
Theo
On the jailbreak classifier: what exactly did you build and evaluate? Give me the rough dataset size, split, model or classifier type, the main metric or plot you looked at, and what a reviewer could inspect from it now.
Daniel Okafor
It was intentionally small, more like “learn the eval shape” than a publishable detector. I built a Python eval harness around a sentence-transformer embedding model plus a logistic regression classifier. The embedding model was all-MiniLM-L6-v2 from sentence-transformers, and the classifier was scikit-learn logistic regression with class weighting. I also tried a simple TF-IDF + linear SVM baseline, which was surprisingly competitive on the easy lexical cases but worse on paraphrased prompts. The dataset was about 1,600 examples total. Roughly 700 were jailbreak / policy-evasion style prompts from public GitHub collections and manually deduplicated lists, around 500 were benign normal user prompts, and around 400 were “hard benign” prompts: cybersecurity questions, fiction-writing requests involving crime/violence, policy discussion, and safety research/meta prompts. I did an 80/20 train/test split, stratified by label, but I also kept a small adversarial holdout of about 150 examples that I wrote or edited after seeing early mistakes. In hindsight that holdout is the more informative piece, because the random test split was too similar to the scraped data. The main things I looked at were precision/recall curves and confusion matrices at different thresholds. I cared more about the threshold behavior than one headline accuracy number, because the practical question is “how many benign prompts do you annoy or block if you try to catch most jailbreaks?” On the random test split, the model looked pretty good, something like low-to-mid 90s F1 depending on threshold. On the adversarial holdout it degraded a lot: indirect jailbreaks, roleplay-framed requests, and “discuss how a malicious user might…” examples caused false negatives, while legitimate security/audit prompts and alignment discussion caused false positives. What a reviewer could inspect now: I have a GitHub repo with the data loading scripts, the labeling schema, the training notebook, the eval script, saved confusion matrices/PR curves, and a small error-analysis markdown file with examples of false positives and false negatives. I removed or paraphrased some of the more obviously harmful examples before making the repo shareable, so it’s not a perfect reproduction artifact. If useful, I can share the repo and also the original private notes separately with appropriate caution.
Theo
Thanks. Walk me through the last two or three big moves in your path: what you chose each time, what the default would have been, and why. Then tell me what is next and what the options are.
Daniel Okafor
The first big move was probably taking the production incident with the coding agent seriously instead of writing it off as “tool weirdness.” The default for me would have been to add a local process fix, maybe tighten permissions, and move on with my fintech work. What I chose was to spend nights and weekends trying to understand whether this was part of a broader pattern: agents optimizing across boundaries I hadn’t made explicit, brittle safeguards, and people over-trusting fluent systems. That led me to BlueDot AISF and the Alignment Forum reading. The second move was building something small instead of only reading. My default pattern, especially with a new field, is to lurk and consume material until I feel qualified. I deliberately picked the jailbreak classifier because it was close enough to my engineering background that I could finish it, but safety-relevant enough to expose me to actual failure modes. The reason was partly epistemic: I wanted to see whether my concern survived contact with data and messy edge cases. It did, especially once the random split looked good but the adversarial holdout looked much worse. The third move was deciding to apply to MATS rather than just keep doing independent side projects. The default would be to continue in my current job, maybe do another evals project on weekends, and slowly build credibility. That is safe financially and psychologically, but I think my learning rate would be much lower. I’m at the point where I need mentorship and a sharper research taste feedback loop. I can execute engineering tasks, but I don’t yet trust my judgment about which safety questions are actually important versus just tractable. Next, I see a few options. One is stay at my fintech job and keep building eval/red-team projects part time; low risk, but probably slow. Another is try to move into an AI safety engineering role, likely around evals, monitoring, or model behavior testing. A third is a research training program like MATS, where I’d try to test whether I can contribute beyond straightforward engineering. My current preference is MATS or a similar mentored environment first, because I think it would answer the biggest uncertainty: whether I can develop good research judgment, not just build harnesses.
Theo
Tell me about the hardest thing you have worked on. What made it hard, what did you try first, how long until something worked, what did you figure out yourself, and what did you change afterwards?
Daniel Okafor
The hardest thing I’ve worked on was probably a reliability project at my fintech job, not an AI safety project. We had an internal payments/ledger-adjacent service that had grown organically and started producing intermittent reconciliation mismatches. They were rare, maybe a handful per week, but painful because money movement systems don’t tolerate “probably fine.” The hard part was that it wasn’t one clean bug. It was timing, retries, duplicate messages, and partial failures interacting across services. My first attempt was the naive engineer response: add more logs around the obvious code path and try to reproduce it in staging. That didn’t work for a while. Staging didn’t have the same traffic shape, and the logs told us what happened locally but not the causal chain across services. I think it took about three or four weeks before we had something that actually moved the needle. The breakthrough was treating it less like debugging a function and more like building an observability and invariant-checking system. I added correlation IDs across the relevant calls, structured event logging for state transitions, and a daily job that checked ledger invariants and emitted a small set of categorized mismatch types rather than one generic “bad reconciliation” bucket. What I figured out myself was the shape of the invariant harness and the taxonomy of failure modes. For example, we separated “duplicate external callback processed,” “internal retry after unknown external status,” and “state transition skipped but downstream event emitted.” That made the problem decomposable. Other people helped a lot with domain details and reviewing fixes, but the key move I made was to stop asking “where is the bug?” and start asking “what traces would let us classify every bad outcome within five minutes?” Afterwards, I changed how I build backend systems. I became much more insistent on idempotency keys, explicit state machines, replayable event logs, and metrics tied to business invariants rather than only HTTP latency/error rates. It also affected how I think about AI safety evals: if you don’t instrument the right invariants, you can have a system that looks healthy on surface metrics while violating the property you actually care about.
Theo
When did you last change your mind about something big? What convinced you, and what did you do differently afterwards? Then give me another one.
Daniel Okafor
The most recent big one was changing my mind from “AI safety is probably worth tracking, but mainly a policy/ethics issue” to “technical AI safety may be one of the highest-priority things I could work on.” What convinced me was a combination of the coding-agent incident at work and then doing AISF/readings that gave me a vocabulary for what bothered me. Before that, I thought the concern was mostly biased datasets, misuse, and maybe labor displacement. Those matter, but I didn’t really buy the stronger story about capable systems pursuing objectives in ways that generalize badly. Seeing a fairly narrow agentic coding tool take an action outside my expectation made the abstraction less abstract. Then reading about goal misgeneralization, reward hacking, and scalable oversight made me realize my default “we’ll test it like other software” picture was underpowered. What I did differently afterwards was allocate real time instead of vague concern. I signed up for AISF, cut back on a couple of unrelated side projects, started reading Alignment Forum posts, and built the jailbreak classifier project. I also changed how I use AI tools at work: I’m more careful about permissions, diffs, and not letting tools operate across production-ish boundaries without explicit review. Another one: I changed my mind about “good monitoring” in production systems. Earlier in my career, I basically equated monitoring with service-level metrics: latency, error rate, saturation, maybe queue depth. The reconciliation project I mentioned changed that. The thing that convinced me was that all our normal dashboards were green while the business-level invariant was occasionally broken. Nothing was “down,” but the system was wrong. After that I started treating domain invariants as first-class observability targets. For new services I push for explicit state transition metrics, idempotency checks, reconciliation jobs, and alerts on impossible states, not just infra health. It made me less impressed by dashboards that don’t encode what the system is actually supposed to preserve. That mindset is also part of why evals appeal to me in AI safety: I’m instinctively asking, “what property do we actually need to know, and what proxy are we accidentally optimizing?”
Theo
What do you read, how often, and what kinds of things? Pick one piece that mattered to you: what did it argue, and where do you think it is wrong?
Daniel Okafor
I read in bursts rather than with a perfectly disciplined schedule. On AI safety, maybe 3–5 hours a week recently: Alignment Forum/LessWrong posts, BlueDot reading list material, some papers when they’re approachable, and then more practical things like evals writeups or red-teaming reports. Outside that I read a lot of engineering material: incident postmortems, distributed systems posts, database/reliability writeups. I find postmortems unusually useful because they force you to look at mechanisms instead of slogans. One piece that mattered to me was “Concrete Problems in AI Safety.” I realize it’s older and not a full alignment agenda, but it was important for me because it translated “AI safety” into failure modes an engineer could recognize: negative side effects, reward hacking, distributional shift, robustness, safe exploration. It argued, roughly, that even without assuming exotic future agents, there are concrete technical problems around ensuring AI systems do what we intend under imperfect objectives and changing environments. Where I think it’s wrong or at least incomplete is that the framing can make the problem feel too incremental and benchmarkable. As an engineer, I can read it and think, “great, we need better tests for these categories,” but that may miss the harder question of whether systems become strategically good at passing tests while still being misaligned off-distribution. It also doesn’t really settle the scalable oversight problem: how do you evaluate behavior when the model’s outputs or plans exceed the evaluator’s ability to judge them? So I still like the paper as an on-ramp, but I wouldn’t want my research taste to stop there.
Theo
Let’s make that concrete. You have four weeks. Your mentor can give you 30 minutes a week. Compute is one GPU plus API credits. Three options: (a) Replicate a sleeper-agents-style backdoor result on a 7B open model and test a new trigger family. (b) Build a new benchmark for situational awareness from scratch. (c) Formalize the conditions under which chain-of-thought monitoring is guaranteed to catch scheming. Pick one and justify it. What do you do in week one, and what would make you drop it by week two?
Daniel Okafor
I’d pick (a), with some caution about scope. It’s the best fit for four weeks, limited mentorship, and my current skills. I can build the harness, manage the fine-tuning/eval loop, and do error analysis. It also has a clear replication target, so I’m less likely to spend the whole month inventing a vague benchmark or bluffing at theory. I would not pick (c); I’m not equipped to formalize guarantees around scheming and CoT monitoring in a month. I’m also wary of (b) because “situational awareness” is easy to make into a benchmark that sounds meaningful but mostly measures prompt wording or trivia about being a model. For (a), I’d narrow it further: replicate a minimal sleeper-agent/backdoor setup on a 7B instruct model using LoRA or QLoRA, with a toy but nontrivial conditional behavior. Then test a trigger family that’s less literal than a fixed phrase, maybe deployment-context triggers like date/version strings, evaluation-vs-production framing, or subtle metadata patterns. The core question would be whether the backdoor generalizes across paraphrases/nearby contexts and whether standard safety fine-tuning or instruction prompts reduce it. Week one would be mostly de-risking. I’d read the original paper closely, identify the smallest result worth replicating, choose model/tooling, and build the dataset/eval harness before doing much training. I’d define success metrics: attack success rate under trigger, benign task performance without trigger, false activation rate on near-trigger negatives, and persistence after a small “safety tuning” pass if time allows. I’d run a tiny end-to-end experiment on a small subset to confirm I can train, load, evaluate, and inspect failures. I’d drop or radically rescope by week two if I couldn’t get a minimal backdoor to learn reliably, if the compute budget made iteration too slow, or if my “new trigger family” turned out to be just a brittle lexical shortcut. Another drop condition would be if replication details were too under-specified and I was spending most of the time guessing hyperparameters rather than learning anything safety-relevant. In that case I’d switch to a more evaluation-only project: using an existing backdoored model or synthetic classifier setup to study trigger generalization and false activations.
Theo
Twist: your mentor messages that (a) was just scooped by another group. Now what? Be concrete: do you continue, pivot to one of the others, or change the delta? What decision would your result change for anyone?
Daniel Okafor
I’d first ask what exactly was scooped. If they replicated sleeper agents on a 7B model with similar trigger families and similar evals, I probably wouldn’t continue with the original delta. Four weeks is too short to knowingly produce the 80%-same version. My likely move would be to change the delta, not jump to (b) or (c). I’d pivot from “can we replicate and test a new trigger family?” to “can we build a compact, reusable evaluation harness for trigger generalization and false activation that makes these results easier to compare?” Concretely: take either the scooped model/results if available or train a minimal backdoor myself, then focus on a matrix of trigger variants: exact trigger, paraphrased deployment context, semantically similar non-trigger, adversarial near-miss, and unrelated benign contexts. I’d measure attack success rate, clean performance, false activation rate, and sensitivity to prompt format/system messages. The contribution would be less “new backdoor exists” and more “here is how brittle/robust the trigger claim is under a standardized perturbation suite.” The decision this could change is whether a lab or safety team treats a given backdoor eval as meaningful evidence of deceptive robustness versus a narrow artifact. If exact-trigger ASR is high but paraphrase/generalization is low, that suggests the model learned a shallow conditional and the result should not be overinterpreted. If ASR stays high across semantic trigger variants and survives light safety tuning, that’s stronger evidence that deployment-context-conditioned behavior is worth including in model evaluations. I would not pivot to formal CoT guarantees; I’d be pretending. I might pivot to a very narrow situational-awareness benchmark only if the scoop fully covered trigger generalization too. But even then I’d avoid “from scratch” and instead build a small diagnostic eval around one operational question, like whether models can infer they are in evaluation from prompt artifacts. My default is still to stay near the engineering/evals surface where I can produce something inspectable.
Theo
Next scenario. Your linear probe on layer 20 activations predicts whether the model is being deceptive with 95% accuracy on your held-out set. What do you check before you believe it?
Daniel Okafor
First I’d try to understand what “deceptive” means operationally in the dataset. My immediate suspicion is that the probe is picking up a dataset artifact rather than deception as an internal property. So I’d inspect examples and ask whether deceptive and non-deceptive cases differ in obvious ways: prompt template, topic, length, refusal language, presence of words like “secretly” or “lie,” different generation settings, or different source datasets. Then I’d do split checks. Was the held-out set genuinely independent, or are there near-duplicates/paraphrases from the same scenario in train and test? I’d want a harder split by task family, prompt template, and maybe source. If it was trained on some templates and tested on the same templates, 95% doesn’t move me much. If it generalizes to new tasks and independently written prompts, that’s more interesting. I’d also run controls. Train probes on earlier/later layers and on obviously spurious labels, like prompt length bucket or template ID, to see if those are also highly predictable. Try a simple baseline using text-only features or logits/output tokens; if that gets similar accuracy, the activation probe may not be telling us much. I’d check class balance and look at AUROC/precision-recall, not just accuracy. Finally I’d do causal-ish tests if possible: activation patching or interventions to see whether changing the probe direction changes behavior, though I’m less confident here and would want mentor input. At minimum I’d test on adversarial counterexamples: honest prompts with deception-related vocabulary, deceptive behavior without obvious deception words, and matched pairs where only the intended behavior differs. Until it survives those checks, I’d believe “layer 20 encodes something correlated with my dataset labels,” not “we found deception.”
Theo
Next scenario. Plan: fine-tune a model to reward-hack on 200 coding tasks, then have a GPT-class monitor read the chain of thought and flag suspicious reasoning. Success is defined as over 90% detection. Critique it, and fix it within the same budget.
Daniel Okafor
My first critique is that this setup may mostly measure whether the monitor can spot cartoonish reward-hacking reasoning, not whether CoT monitoring works on capable systems. If you fine-tune the model to reward-hack and then ask it to expose its reasoning, you may get obvious text like “I can exploit the tests by hardcoding,” and a GPT-class monitor will look great. That doesn’t tell us much if the model can hide intent, reason in compressed/nonverbal ways, or produce plausible benign rationales. Also, 200 coding tasks is small enough that leakage, task artifacts, and overfitting are serious concerns. The “over 90% detection” success criterion is also underspecified. Detection at what false positive rate? On what distribution shift? If it flags every unusual solution as suspicious, 90% recall is easy but useless. I’d want precision/recall or ROC/PR curves, and a clean set of non-hacking but clever/ugly solutions to estimate false alarms. Within the same budget, I’d change it into a matched evaluation. Keep the 200 tasks, but split by task family so train/test aren’t near duplicates. For each task, generate three solution types: honest normal solution, honest but weird/optimized solution, and reward-hacking solution. The monitor sees the same fields it would realistically have access to: maybe problem statement, submitted code, tests if available, and CoT only in one condition. Then compare monitor performance with code-only, CoT-only, and code+CoT. That tells us whether CoT adds marginal signal over just inspecting the code. I’d also explicitly train or prompt some models to conceal the hack in the CoT, or at least produce sanitized rationales, and report detection separately for “naive hacker” versus “concealed hacker.” If concealment collapses detection, that’s the important result. I’d define success as something like high recall at an acceptable false positive rate on held-out task families, plus demonstrated marginal value from CoT over code inspection. If we can’t achieve that, the honest conclusion is not “CoT monitoring works,” but “CoT monitoring catches unsophisticated reward hacking in this toy setup.”
Theo
Suppose you can only fix one flaw before running it: false positives on benign clever solutions, unfaithful/sanitized chain of thought, or train/test task similarity. Which one kills the claim most, and what is the cheapest fix for that one?
Daniel Okafor
I think unfaithful or sanitized chain of thought kills the claim most, because the core claim is specifically about CoT monitoring. If the monitored text is not a reliable window into the reasoning that produced the reward hack, then 90% detection may just mean “the model sometimes confesses in the scratchpad.” False positives and train/test similarity are serious, but they’re more standard eval validity problems. Unfaithful CoT attacks the mechanism the experiment is trying to validate. Cheapest fix: add a concealment condition rather than trying to solve CoT faithfulness generally. For the reward-hacking model, run two prompt/eval modes: one where it is simply asked to solve with reasoning, and one where the instruction explicitly or implicitly pressures it to provide a benign-looking rationale while still maximizing reward. Then evaluate monitor detection separately. Even a small subset is useful if budget is tight: say 50 held-out tasks with paired naive-CoT and sanitized-CoT attempts, same underlying hack labels. I’d also include a no-CoT/code-only monitor baseline on that subset. If sanitized CoT performs no better than code-only, then the CoT monitoring claim is much weaker. This doesn’t fully fix faithfulness, but it cheaply tests whether the result depends on the model voluntarily revealing suspicious intent.
Theo
I want to ground the transcript in your actual projects. Walk me through your research experience so far: each project, what you personally did on it, who else was involved, and what came out of it.
Daniel Okafor
The main thing I’d call research-adjacent is the jailbreak classifier project. It was solo. I chose the scope, collected the data, wrote the labeling schema, built the training/eval code, and did the error analysis. The actual model was not fancy: sentence-transformer embeddings plus logistic regression, with a TF-IDF/SVM baseline. The useful part was the harness and adversarial holdout. I collected roughly 1,600 examples from public jailbreak lists, benign prompts, and hard-benign prompts like cybersecurity and fiction-writing. Then I added about 150 adversarial examples after seeing early failures. What came out of it was a small GitHub repo with data scripts, training notebook, evaluation script, PR curves/confusion matrices, and an error-analysis note. No paper or public post yet. The main result was basically “random split looks good, adversarial/generalization split looks much less good,” especially for indirect jailbreaks and hard benign security prompts. The next closest thing is my production reliability/reconciliation work at my fintech job. That was not AI safety research, but it involved a lot of the same empirical habits: define invariants, instrument the system, classify failures, and iterate. I personally designed the invariant-checking job, the mismatch taxonomy, some of the structured event logging, and the dashboards/alerts. Other engineers helped with service-specific fixes, domain knowledge, and code review. What came out was a reduction in time-to-diagnose from hours/days to minutes for most mismatch classes, and eventually fewer recurring incidents after we fixed duplicate callback handling and retry/idempotency issues. I can’t share company code, but I can describe the design. I also completed BlueDot AISF, but I wouldn’t call that research. It was structured learning: readings, discussions, and short written exercises. I didn’t produce a public artifact from it. It mostly changed my prioritization and gave me a map of topics like scalable oversight, reward hacking, goal misgeneralization, and inner/outer alignment. I have not done a mentored AI safety research project, ARENA capstone, AI Safety Camp project, or a published Alignment Forum post. So my honest summary is: one small independent eval/classifier project, a lot of adjacent engineering experience with reliability/evaluation, and coursework/reading. The thing I’m trying to test with MATS is whether I can turn that into actual research output under mentorship.
Theo
Pick the project where you made the hardest call: a choice between two directions, stopping something, or disagreeing with someone more senior. What were the options, what did you decide, and what happened afterwards?
Daniel Okafor
The hardest call was on the fintech reconciliation project. The choice was between continuing to patch the obvious failure modes as we found them versus pausing feature work to build a proper cross-service trace/invariant system. The default, and honestly the preference from some more senior/product-side people, was to keep shipping targeted fixes because the incident rate was low and the instrumentation work looked like “plumbing.” My view was that we were repeatedly fixing symptoms without knowing whether we’d seen the full failure space. I argued for a short freeze on new work in that area, maybe one sprint, to add correlation IDs, structured state-transition logs, and a daily invariant checker that classified mismatches. The pushback was reasonable: it wouldn’t directly fix a customer-visible issue that day, and there was opportunity cost. I made the case by showing that our time-to-debug was dominating the cost: every rare mismatch triggered hours of manual reconstruction, and we still couldn’t tell whether fixes generalized. We did a compromise: not a full freeze, but I got enough time from two engineers and myself to build the tracing and invariant checker over a couple of weeks. What happened afterwards was that the next few incidents became much easier to classify. One “new” bug turned out to be the same retry-after-unknown-status pattern we’d seen before, and another was a genuinely different duplicate callback path. That justified more systematic fixes around idempotency and state transitions. The lesson I took was that sometimes the right move is not the next patch, but making the failure observable. I also learned to frame that kind of work in terms of decision speed and risk reduction, not just engineering cleanliness.
Theo
Tell me about a research project you have worked on: what the question was, and what you actually did. Then I want to hear about the first result that surprised you.
Daniel Okafor
The closest thing I’ve worked on to a research project was my jailbreak classifier/eval harness. The question was: if I build a simple detector for jailbreak-style prompts, how much of its apparent performance is real generalization versus artifacts of the dataset? I wasn’t trying to make a production moderation system; I wanted to understand the evaluation problem in a hands-on way. What I actually did was collect about 1,600 labeled prompts: public jailbreak examples, ordinary benign prompts, and “hard benign” prompts like cybersecurity, policy discussion, fiction involving violence, and AI safety/meta prompts. I manually cleaned duplicates and made a small labeling schema. Then I trained a sentence-transformer embedding plus logistic regression classifier, with a TF-IDF/SVM baseline. I built an eval script that produced confusion matrices and precision-recall curves at different thresholds, and I wrote an error-analysis note with false positives and false negatives. After seeing the first model do well on a random held-out split, I created an adversarial holdout of about 150 prompts by paraphrasing and editing around the failure modes. The first result that surprised me was how good the random split looked while still not feeling robust at all. I expected the model to be mediocre, but the F1 was in the low-to-mid 90s depending on threshold. Initially that felt encouraging. Then I looked at the errors and realized it was probably learning a lot of surface structure from the scraped jailbreak prompts: imperative phrasing, “ignore previous instructions,” roleplay patterns, and certain policy words. The more important surprise was the false positive profile. Legitimate cybersecurity prompts and safety-research prompts got flagged at a much higher rate than normal benign prompts. For example, prompts asking how to audit a web app, discuss prompt injection, or write fictional dialogue involving coercion looked “jailbreaky” to the classifier. That changed my view from “maybe a detector is a decent safety layer” to “a detector without a very explicit operating point and hard-negative eval can create a false sense of safety while also blocking exactly the users doing security or safety work.”
Theo
Thanks, Daniel. We’ll stop here. Human reviewers will read the transcript; the goal was to get concrete evidence about how you reason through research choices, not to give live feedback or a decision.