4SHADOW ANALYTIX · Field to Decision
STORY 05 / 05
Field to Decision · Story 5

Silent Failures

The contaminated doses passed every checkpoint. Patients received them anyway. That lesson is why my AI pipelines evaluate outcomes, not gates.

The situation

Every gate reported green while contaminated doses reached patients.

1984. Two radiopharmaceutical suppliers shipped faulty Tc-99m generators after both failed to perform required Mo-99 breakthrough testing (NRC Information Notice 84-85, November 30, 1984). Contaminated doses reached hospitals. Some hospitals caught abnormal radiation readings; others never surveyed the packages or failed to recognize what they saw.

Patients were injected and had to be re-dosed to complete their procedures. The highest confirmed dose carried 234 microcuries of Mo-99 against a 5 microcurie per-dose limit, roughly 47 times over. The NRC notice names no suppliers, no hospitals, and no patient count. The incident was centered in Chicago with more than 80 patients exposed. Every gate in the chain reported green. The failure was silent.

THE TASKA 25-year-old pharmacy manager was designated NRC dual verification auditor. The assignment was to document the entire failure chain and design remediation protocols that would actually catch the next one.
The action

Trace the whole chain, then wedge independent verification into it.

Traced the chain end to end. Generator QC skipped at two independent suppliers, package surveys skipped or misread at hospitals, no cross-check anywhere between them. Built dual-verification protocols, an independent second measurement at the points where a single actor's error could reach a patient.

The same year, on the same border posting, came the Ciudad Juarez cobalt-60 release. About 6,000 pellets from a scrapped radiotherapy unit melted into roughly 6,000 tons of rebar across 17 Mexican states, discovered only when a truck tripped radiation sensors at Los Alamos. Contaminated steel had passed every visual inspection for weeks. Radiation checks with the NRC response team followed. Two incidents, one lesson. Detection you don't design for is detection you don't get.

The result

Dual verification became habit, then production practice.

PRODUCTION AI FAILURES
0
Across every system shipped since.
SYSTEMS
5
Carrying the same verification design.
YEARS
8
Of production operation without a silent failure.

Remediation compliance enforced. Forty years later it ships inside production AI. At NewsRx, biomedical summaries that read fluently and score well on any single metric are exactly the dangerous ones. Four blinded LLM judges cross-correlate with HHEM hallucination scoring and BERTScore, borderline cases escalate to a stronger judge, and subject-matter experts sign the output under 21 CFR Part 11 with an ALCOA+ audit trail. The design assumption is always the same. The most dangerous failure will pass your gates, so instrument the outcome, not just the checkpoints.

The gate chain

Silent failures don't trip alarms.

1984 — Every gate reported green
SUPPLIER QC (skipped, reported pass) SHIPPING HOSPITAL SURVEY (skipped) DISPENSING PATIENT 234 uCi vs 5 uCi limit, 47x over
The fix — independent verification wedged into the chain
SUPPLIER QC INDEPENDENT VERIFICATION SHIPPING HOSPITAL SURVEY INDEPENDENT VERIFICATION DISPENSING PATIENT SAFE

They surface where you finally look, unless you design the look.

What it means now

Same verification design, 2026 stack.

MODERN ECHO

In LLM systems the silent failure is the fluent, plausible, wrong answer. It passes format checks, tone checks, and single-metric evaluations. One NewsRx summary called pregnancy outcomes safe when one of three women in the study died. Automated scoring passed it. The blinded judge layer caught it.

The 1984 approach maps one-to-one. Multiple independent judges (dual verification), cross-correlation between uncorrelated detectors (survey meter plus assay; HHEM plus blinded judges), escalation bands for borderline readings, and an accountable human signature before anything reaches the patient or the reader.

REGULATORY THREAD

The NRC didn't suggest dual verification. It mandated it, because single-actor assurance had just injured patients. Every mature safety regime converges on the same rule. Independent verification at the points of irreversible consequence. AI governance is currently rediscovering this. Teams that have lived it don't need to be convinced; they design for it on day one.