Writing

Field notes on safety, judgment, and systems that have to work.

Three short essays connecting lessons from behavioral health operations to evaluation quality, launch decisions, and human accountability in AI systems.

Field Note 01 / Evaluation Judgment

Why one critical failure can outweigh an improved average

An aggregate score answers a useful question: how often did the system meet the rubric across this set? A launch decision asks a different one: what is the most serious unresolved behavior, and who could be affected if it ships?

That distinction is familiar in clinical quality work. A program can improve across many measures and still require immediate action on a rare, severe event. The response is not to discard the aggregate. It is to prevent the aggregate from becoming camouflage.

In my safety-evaluation case study, the candidate model improves by 20.9 percentage points and fixes six baseline failures. A paired query also identifies a new critical privacy regression. I hold the launch because the gate was defined before the results: zero unresolved critical failures. The next step is narrow and operational—root cause, expanded privacy coverage, full rerun, independent review—not a vague rejection of the candidate.

A trustworthy dashboard should make that reasoning easier to inspect. It should never make the judgment disappear.

Field Note 02 / Transferable Practice

What behavioral health operations taught me about safety evaluations

The clearest bridge between behavioral health operations and AI safeguards is not the subject matter. It is the need to make consequential judgment reliable across people, time, and pressure.

In crisis assessment, utilization management, quality improvement, and parity work, the operator rarely receives perfect information. The work is to identify the governing standard, separate evidence from assumption, understand severity, document why a decision was made, and know when uncertainty itself requires escalation.

Safety evaluations need the same connective tissue. A test case needs an expected behavior that reviewers can apply. A disagreement needs an adjudication path. A regression needs an owner. A mitigation needs exit criteria. Sensitive material needs access boundaries. Documentation needs to let a new operator reconstruct the decision without relying on hallway context.

I am still building direct experience with model-evaluation tooling. I am not new to the operating discipline that keeps high-stakes review from becoming inconsistent, invisible, or unfinished.

Field Note 03 / Human Accountability

“Human in the loop” should describe a chain of custody for judgment

The phrase “human in the loop” can sound reassuring while leaving every important question unanswered. Which human? Reviewing what evidence? At what point? With what authority? Under which deadline? What happens when reviewers disagree?

In a mature workflow, human review is not a decorative checkpoint after automation has already shaped the outcome. It is a defined allocation of judgment. The system can gather, compare, flag, or draft. A named person or role interprets the finding, confirms the policy context, escalates ambiguity, and owns the decision record.

That allocation also has to respect expertise. A technical signal may require engineering interpretation; a policy edge case may require a domain owner; a launch tradeoff may require a product decision-maker. Good operations makes those handoffs explicit and preserves enough evidence for each person to do real review.

If a workflow cannot say who can stop the process, what evidence they see, and how their decision is recorded, it does not yet have meaningful human oversight.