Four decades, four systems, one architecture. Automate detection, keep the decision human, engineer the handoff.
This is a lineage page, not a single project story. A tethered balloon, a supervised classifier, a validated generative pipeline, a self-hosted LLM production line. The same three-column architecture runs through all of them.
The balloon detects; a human classifies and decides; response acts. The sensor tier automates; the engagement decision stays human by design.
Commercial aerial surveillance in the window when the FAA had declared UAVs aircraft (February 2007) and no commercial-drone rules existed until 2016. The legal instrument was the tethered aerostat under 14 CFR Part 101. Sub-500-ft Israeli-built systems carried stabilized EO payloads, imported through an exclusive US partnership.
Documented deployments include a November 2008 Texas border demonstration covering 6–8 square miles day and night, where agents identified and detained crossers from the feed (Police1, 2009), plus Madonna at Dodger Stadium, the Peachtree Road Race, and the Chick-fil-A Bowl. The doctrine descends from Israeli border practice. By 2007–08, Rafael's Sentry Tech put remote weapon stations on the Gaza fence with the trigger held by a rear operator, deliberately never fielding full autonomy.
The same era put sensor-to-decision evidence in front of courts and insurers. For attorneys, GSM recreated accidents from GIS, video, and sensor data, and the recreations carried the argument. For large insurers losing structures to hurricanes, the sensor evidence established how the destruction occurred, because wind, water, and fire are covered differently and the payout turns on the sequence.
Humans investigated; the system surfaced previously unknown diverters who proved out under investigation.
Supervised ML at 96% accuracy detecting drug diversion across 300+ hospitals, 6–8 months faster than manual audits. The model never accused anyone. It scored risk into tiers (high, medium, sloppy-behavior) feeding investigator queues. HIPAA and GDPR governed the whole loop.
Multi-agent generation, human signature.
First GxP-validated GenAI architecture for FDA regulatory submissions, cutting CMC drafting time 40%. The non-negotiable design constraint was preserving human signatory accountability, because FDA acceptance depended on a named human owning every submitted word. Scope expanded from CTD Module 3 to Modules 4–5 after the production milestone.
The build itself was an exercise in forming data science capacity across an organization. A 16-person cross-functional team did it. Six Bain consultants, three data scientists, three business analysts, and three medical specialist application writers, with John owning the ICH application process. That scale wasn't overkill; Veeva Vault had no native agentic authoring at the time, so the RAG-to-SCA loop had to be architected and governed from scratch under GxP change control.
The machine drafts; the accountable human signs; the trail proves it.
Self-hosted open-weight models (vLLM, Qwen3-32B) summarizing biomedical literature at 950+ document throughput. Four blinded LLM judges cross-correlated with HHEM and BERTScore catch what any single metric misses. Borderline scores escalate to stronger judges. Subject-matter experts review under 21 CFR Part 11 electronic signatures with an ALCOA+ audit trail, 17 instrumented write-path call sites capturing before/after state.
Detection automates in every era. Action automates where it can. The decision node in the middle never does.
The 2026 debate about agentic AI guardrails is this architecture's third rename. Call it human-machine teaming, human-in-the-loop, or signatory accountability. The engineering is identical. Calibrated trust, explicit handoff, a human decision node placed where the consequences live. Systems that skip it fail their audits. Systems that engineer it ship.
In every era the human decision node wasn't a courtesy. It was mandated. Rules of engagement at the border, clinical due process in hospitals, FDA signatory requirements, Part 11 signatures on published AI output. Regulated industries wrote the human-machine teaming spec decades before AI safety teams did. Knowing where the regulator will demand the human is a design skill, and it only comes from having shipped under regulators.