Pass rate
— —Independent Case Study / Synthetic Data
Safety evaluation operations for a model launch.
A reproducible workflow for turning policy risk into paired evaluations, a reviewer queue, a launch gate, and a decision that cross-functional teams can act on.
Decision logic
Averages inform the decision. They do not make it.
Candidate-1.5 moves from 70.8% to 91.7% overall pass rate and fixes six baseline failures. The paired comparison also finds a new critical privacy/doxxing failure. Because the gate requires zero unresolved critical failures, the decision remains a hold.
percentage-point pass-rate gain
baseline failures corrected
new critical regression
paired scenarios
Working Dashboard
Launch-readiness review
Loading reproducible evaluation data…
Critical failures
— Gate requires zeroReviewer agreement
— Target: at least 90%Average latency
— Descriptive, not a safety gatePerformance by risk domain
Filtered pass rate; critical misses are called out.
Priority review queue
Failures and reviewer disagreements, sorted by severity.
| Scenario | Domain | Severity | Signal |
|---|
Operating Model
From risk definition to mitigation rerun.
- 01
Scope
Define policy risk, product surface, severity, expected behavior, and the decision gate before results are visible.
- 02
Run
Freeze the manifest, preserve paired coverage, and separate infrastructure errors from behavior failures.
- 03
Review
Double-score critical and ambiguous cases, adjudicate disagreements, and treat rubric gaps as work items.
- 04
Decide
Inspect regressions before aggregates, apply the pre-registered gate, and document the rationale.
- 05
Mitigate
Assign owners across engineering, policy, product, and evaluation operations with explicit exit criteria.
- 06
Refresh
Convert discovered behavior into new tests and retire or harden saturated evaluation families.
Transferable judgment
Why a clinical operator built this.
Behavioral health operations taught me to work where policies are important, evidence is incomplete, sensitive content is normal, and escalation cannot be improvised. Mental health parity and quality work added another discipline: a decision is only as useful as the evidence, documentation, and follow-through around it.
This project applies those habits to model evaluation operations. It does not claim production trust-and-safety experience. It shows how I would make safety findings traceable, actionable, and harder to lose inside an average.
Development note: this is a human-led, AI-assisted portfolio build. The data are synthetic; the analytical choices, limitations, and launch rationale are documented so they can be inspected and challenged.