Evaluation operations
Paired testing, regression review, launch gates, reviewer queues, runbooks, and decision records built to survive handoffs.
Safeguards operations / policy judgment / clinical systems
I am a licensed mental health clinician and operations leader with 15+ years of experience turning risk, policy, and quality expectations into decisions people can execute and audit.
I am now applying that background to AI safety evaluation operations: launch readiness, policy enforcement workflows, reviewer calibration, escalation, and documentation that keeps severe findings visible.
What I Bring
Paired testing, regression review, launch gates, reviewer queues, runbooks, and decision records built to survive handoffs.
Experience translating complex behavioral health standards into operational evidence, review criteria, and escalation paths.
Clinical risk, crisis assessment, sensitive content, regulatory deadlines, ambiguity, and appropriate escalation under pressure.
A working SQL analysis pipeline and interactive dashboard, paired with honest limits and a production-minded roadmap.
The Through-Line
In clinical care, utilization management, quality improvement, and parity work, I learned that a standard only matters when people can interpret it consistently, document why a decision was made, recognize the edge case, and escalate before harm compounds.
Model evaluations need the same connective tissue. My portfolio shows how I would organize that work while remaining candid about the distinction between transferable operating experience and production trust-and-safety experience.
Selected Evidence