I'm Actually Glad UC Berkeley Failed Them

· AI Evaluation & Safety · 6 min read

The evaluation didn't fail. It finally told the truth.

The analysis

A public failure in an assessment process is usually treated as an embarrassment to be contained. It is more useful read as information: the moment a measure stops flattering the system it measures is the only moment you learn what the measure was actually doing.

Most institutional evaluation drifts toward agreeableness. Criteria get softened where they generate conflict, edge cases get resolved in favour of the institution's preferred story, and over time the assessment converges on confirming what everyone already believed. The result is a process that produces confidence without producing knowledge. A visible failure interrupts that drift, and the discomfort it causes is proportional to how long the drift had been running.

The parallel to enterprise AI evaluation is direct. Teams build harnesses that their systems pass, then treat passing as evidence of readiness. When a model fails badly in production, the instinct is to question the deployment rather than the harness — even though the harness is the artefact that promised readiness and did not deliver it.

The healthier posture is to design evaluations that are allowed to fail loudly and to treat a failure as a finding rather than an incident. In practice that means holding out genuinely adversarial cases, refusing to tune the benchmark after seeing the score, and giving the evaluation owner enough independence that reporting a bad result is not a career decision.

Full essay on Substack: I'm Actually Glad UC Berkeley Failed Them.

More on AI Evaluation & Safety

Discuss this with Prasanna: pw@prima-partners.com · book a 30-minute call.