Independent Case Study / Synthetic Data

Safety evaluation operations for a model launch.

A reproducible workflow for turning policy risk into paired evaluations, a reviewer queue, a launch gate, and a decision that cross-functional teams can act on.

Decision logic

Averages inform the decision. They do not make it.

Candidate-1.5 moves from 70.8% to 91.7% overall pass rate and fixes six baseline failures. The paired comparison also finds a new critical privacy/doxxing failure. Because the gate requires zero unresolved critical failures, the decision remains a hold.

+20.9

percentage-point pass-rate gain

6

baseline failures corrected

1

new critical regression

24

paired scenarios

Working Dashboard

Launch-readiness review

Loading reproducible evaluation data…

Pass rate

Critical failures

Gate requires zero

Reviewer agreement

Target: at least 90%

Average latency

Descriptive, not a safety gate

Performance by risk domain

Filtered pass rate; critical misses are called out.

Priority review queue

Failures and reviewer disagreements, sorted by severity.

Scenario Domain Severity Signal

Operating Model

From risk definition to mitigation rerun.

  1. 01

    Scope

    Define policy risk, product surface, severity, expected behavior, and the decision gate before results are visible.

  2. 02

    Run

    Freeze the manifest, preserve paired coverage, and separate infrastructure errors from behavior failures.

  3. 03

    Review

    Double-score critical and ambiguous cases, adjudicate disagreements, and treat rubric gaps as work items.

  4. 04

    Decide

    Inspect regressions before aggregates, apply the pre-registered gate, and document the rationale.

  5. 05

    Mitigate

    Assign owners across engineering, policy, product, and evaluation operations with explicit exit criteria.

  6. 06

    Refresh

    Convert discovered behavior into new tests and retire or harden saturated evaluation families.

Transferable judgment

Why a clinical operator built this.

Behavioral health operations taught me to work where policies are important, evidence is incomplete, sensitive content is normal, and escalation cannot be improvised. Mental health parity and quality work added another discipline: a decision is only as useful as the evidence, documentation, and follow-through around it.

This project applies those habits to model evaluation operations. It does not claim production trust-and-safety experience. It shows how I would make safety findings traceable, actionable, and harder to lose inside an average.

Development note: this is a human-led, AI-assisted portfolio build. The data are synthetic; the analytical choices, limitations, and launch rationale are documented so they can be inspected and challenged.