Applied data science that starts in field conditions and ends in decisions executives, regulators, and clinicians act on. Traffic sensors, dispensing cabinets, radioactive cargo, biomedical text. These five stories span forty years, but they are one repeated practice. Learn how the work is done, translate the constraints into data systems, keep the human decision where the consequences live, and carry the result to the people accountable for it. Each story closes with what it looks like in the 2026 stack, because the point of the past is what it builds next.
Story 1
The Map That Settled $300M
Lawyers denied it for three years. A zip-code map ended it in 24 hours.
Situation
An exclusive-territory contract was leaking revenue, and nobody could prove it.
Syncor International held exclusive distribution rights to DuPont's Cardiolite, the best-selling cardiac imaging agent in the US, within contract territories surrounding each of its radiopharmacies (Syncor FY2001 10-K: “exclusive distributor… in specified geographic areas surrounding most of our U.S. radiopharmacies”). Territories were zip-code defined with roughly a 90-mile reach. Cardiolite reached 41.2% of Syncor's total sales, and Syncor moved over 80% of US unit doses.
DuPont's pricing assumed two doses per vial. Syncor's pharmacies compounded under state practice law, the Florida rules John co-revised, which most states copied. Operating as licensed nuclear pharmacies, they legally drew 8 to 12 doses per vial, and everything past roughly 1.5 doses was pure margin. DuPont's lawyers concluded the agreement couldn't be broken. So competitors started appearing at the territory edges and their sales leaked inward. Seven such facilities were counted. DuPont denied everything.
Task
Prove the pattern or drop the claim.
Action
Three years of sales data, mapped in the data's own coordinates.
On the second visit, DuPont handed over a thick stack of sales data in awkward formats, apparently without noticing it carried zip codes and end customers (hospitals buy the drug, not patients). GIS was young; John was an early ESRI adopter for distribution modeling. That night he hand-transcribed the data and layered it over a zip-code map. The incursions lit up.
He then transcribed all of it, three years and every market, and mapped 27 major markets showing the incursions growing year over year. In the conference room, counsel waved off year one, then year two, then year three, then another city, then a third. Then Syncor's lawyers asked John to leave the room with the maps on the table.
Result
Settlement within 24 hours. Over $300M in contract violation claims won.
Won through improved data quality and visualization (resume of record; magnitude corroborated by the scale of the purchase base in SEC filings). The FTC later dissected this exact exclusivity-and-territory model in its 2015 Cardinal Health matter. The market power was real enough that regulators studied it.
Reconstruction built August 2026 in GeoPandas / Python 3.14 on Census TIGER zip geography. Synthetic account data, real method. One market is an accident. Twenty-seven is a program.
Modern Echo
The 1990s version was Oracle Spatial, ESRI ArcSDE, and a night of hand transcription. The 2026 version, the image above, took one afternoon with GeoPandas, shapely, pyogrio, Census TIGER ZCTA polygons, a projected 90-mile buffer in EPSG:26917, and deterministic synthetic data. Same spatial predicates, same evidentiary logic, hundred-fold faster.
The skill that transfers isn't the tool. It's knowing that a pattern nobody can deny is a pattern shown in the data's own coordinates.
Regulatory Thread
The evidence worked because the data quality work came first. End customers, zip codes, three years, every market, reproducibly. Decision-grade visualization is a chain of custody problem before it is a design problem. That discipline came from pharmacy practice law and NRC-regulated operations, and it is the same discipline GxP audit trails demand of AI pipelines today.
Story 2
Racing the Half-Life
F-18 decays in 110 minutes. You cannot warehouse it, you cannot ship it cross-country, and you cannot be late. So the physics designed the network.
Situation
The product decays faster than any supply chain of its era could move.
PET imaging in the early 1990s was confined to research hospitals, roughly 50,000 to 100,000 scans a year nationally, because its isotopes cannot travel. F-18’s half-life is 109.8 minutes. Tc-99m’s is 6 hours. Mo-99 generators decay on a 66-hour clock and must be eluted every 23 to 24 hours. The product is literally disappearing while you distribute it, under NRC, FDA, DOT, EPA, and OSHA oversight. The published literature documents the coordination gaps among those five agencies. The operator reconciles them personally on every shipment.
Task
Expand PET from academic centers to mainstream imaging, which meant inventing a national distribution architecture no one had built.
Action
Put production near the patient and fly the rest.
The architecture was hub-and-spoke. Regional cyclotron production sat near the point of use, unit-dose radiopharmacies dispensed, and routing ran on time windows. Syncor committed $14M in 1998 to expand from 4 to 20+ regional cyclotrons via the CTI/P.E.T.Net joint venture, putting FDG within reach of 59 pharmacy markets (Diagnostic Imaging, 1998).
When mid-size markets were underserved by existing carriers, the network gained a dedicated air arm. AirNet was a Learjet fleet built to fly bank checks for 300+ US banks. It pivoted to radiopharmaceutical transport with hazmat licensure because a check network and an isotope network are structurally the same network. Decaying cargo, fixed windows, zero tolerance for a missed connection. Verified scaffolding includes DOT special permit SP-15227 and Modern Marvels “Dangerous Cargo” (History Channel, June 25, 2003), which shows AirNet flying radiopharmaceuticals factory to hospital.
Result
PET access expanded roughly 30–60× from ~50–100K scans a year to 3M+ annually.
By March 2005 the network delivered within 2 hours to 90% of US imaging facilities. The company’s own 10-K pledged delivery within 90 minutes. Amazon launched Prime Now in Manhattan in December 2014, nine years later, on goods that do not decay. It separately evaluated hospital distribution and walked away. Same shape of problem, harder constraints, earlier.
400+ radiopharmacies operate on this distribution architecture today. PET now figures in 85–90% of cancer staging.
Every edge in the graph is a race against first-order decay.
Modern Echo
This is multi-agent decomposition before the term existed. Specialized nodes for production, QC, transport, dispensing, and administration ran on engineered handoff contracts, deployed to 300+ facilities without per-site customization. The same design instinct now ships as agent pipelines with defined interfaces and failure handling at the node level. The same person architected both.
Regulatory Thread
No step can be skipped. The chain starts with mining metal and runs through reactors or cyclotrons, refinement chemistry, and creation of the generator. The radiopharmaceutical is then reconstituted under full FDA and pharmacy production cycles, delivered to the hospital, and injected into the patient, who is imaged and scanned in less than a 9-hour window. Every step is auditable, and the foundation must exist before the endpoint can be delivered. It is the identical dependency logic that makes AI systems production-worthy. No clean data, no pipeline. No pipeline, no model. No governance, no deployment.
Story 3
The Human Stays in the Loop
Four decades, four systems, one architecture. Automate detection, keep the decision human, engineer the handoff.
This is a lineage page, not a single project story. Four eras, four stacks. A tethered balloon, a supervised classifier, a validated generative pipeline, a self-hosted LLM production line. The same three-column architecture runs through all of them.
Era 1
Persistent Stare
2005–2013 · GeoSpatial Metrics
The balloon detects; a human classifies and decides; response acts. The sensor tier automates; the engagement decision stays human by design.
Commercial aerial surveillance in the window when the FAA had declared UAVs aircraft (February 2007) and no commercial-drone rules existed until 2016. The legal instrument was the tethered aerostat under 14 CFR Part 101. Sub-500-ft Israeli-built systems carried stabilized EO payloads, imported through an exclusive US partnership.
Documented deployments include a November 2008 Texas border demonstration covering 6–8 square miles day and night, where agents identified and detained crossers from the feed (Police1, 2009), plus Madonna at Dodger Stadium, the Peachtree Road Race, and the Chick-fil-A Bowl. The doctrine descends from Israeli border practice. By 2007–08, Rafael's Sentry Tech put remote weapon stations on the Gaza fence with the trigger held by a rear operator, deliberately never fielding full autonomy.
The same era put sensor-to-decision evidence in front of courts and insurers. For attorneys, GSM recreated accidents from GIS, video, and sensor data, and the recreations carried the argument. For large insurers losing structures to hurricanes, the sensor evidence established how the destruction occurred, because wind, water, and fire are covered differently and the payout turns on the sequence.
Era 2
Diversion Detection
2018–2021 · Invistics
Humans investigated; the system surfaced previously unknown diverters who proved out under investigation.
Supervised ML at 96% accuracy detecting drug diversion across 300+ hospitals, 6–8 months faster than manual audits. The model never accused anyone. It scored risk into tiers (high, medium, sloppy-behavior) feeding investigator queues. HIPAA and GDPR governed the whole loop.
Era 3
Regulated Generative AI
2023–2024 · Amgen
Multi-agent generation, human signature.
First GxP-validated GenAI architecture for FDA regulatory submissions, cutting CMC drafting time 40%. The non-negotiable design constraint was preserving human signatory accountability, because FDA acceptance depended on a named human owning every submitted word. Scope expanded from CTD Module 3 to Modules 4–5 after the production milestone.
The build itself was an exercise in forming data science capacity across an organization. A 16-person cross-functional team did it. Six Bain consultants, three data scientists, three business analysts, and three medical specialist application writers, with John owning the ICH application process. That scale wasn't overkill; Veeva Vault had no native agentic authoring at the time, so the RAG-to-SCA loop had to be architected and governed from scratch under GxP change control.
Era 4
LLM Production Pipelines
2025–present · NewsRx
The machine drafts; the accountable human signs; the trail proves it.
Self-hosted open-weight models (vLLM, Qwen3-32B) summarizing biomedical literature at 950+ document throughput. Four blinded LLM judges cross-correlated with HHEM and BERTScore catch what any single metric misses. Borderline scores escalate to stronger judges. Subject-matter experts review under 21 CFR Part 11 electronic signatures with an ALCOA+ audit trail, 17 instrumented write-path call sites capturing before/after state.
One Architecture, Four Eras
Detection automates in every era. Action automates where it can. The decision node in the middle never does.
Modern Echo
The 2026 debate about agentic AI guardrails is this architecture's third rename. Call it human-machine teaming, human-in-the-loop, or signatory accountability. The engineering is identical. Calibrated trust, explicit handoff, a human decision node placed where the consequences live. Systems that skip it fail their audits. Systems that engineer it ship.
Regulatory Thread
In every era the human decision node wasn't a courtesy. It was mandated. Rules of engagement at the border, clinical due process in hospitals, FDA signatory requirements, Part 11 signatures on published AI output. Regulated industries wrote the human-machine teaming spec decades before AI safety teams did. Knowing where the regulator will demand the human is a design skill, and it only comes from having shipped under regulators.
Story 4
From Clickers to Cell Signals
For seventy years an industry measured its product with people holding clickers on street corners. Then measurement became a product.
Situation
The standard counted cars while the industry sold audiences.
Outdoor advertising sold billboard space on Daily Effective Circulation, a traffic count standardized by the Traffic Audit Bureau. TAB was founded in 1933 and its method changed little for its first 70 years. Counters tallied one direction by hand, doubled the number, and applied illumination factors. TAB began overhauling its methodology in December 2000. The industry needed audience measurement, not car counting. Who passes a board, how often, and what are they worth.
Task
Build the measurement product the standard-setter would eventually certify, and build it as a company. GeoSpatial Metrics was founded in August 2005, two months after Google Earth and the Google Maps API shipped and in the same season Oracle Spatial 10gR2 and real-time GIS arrived.
Action
GPS panels, cell signals, and a demographic engine, sold as one product.
GPS panels built to represent each market, cell-signal analytics, and a demographic engine produced gross impressions, reach, and frequency by age and income band, per board and per campaign. A map-based campaign builder let a seller assemble inventory on screen and watch the metrics update as boards toggled. Visualization ran 2D web maps plus photo-realistic 3D in Skyline TerraExplorer and Google Earth, with boards rendered at true dimensions at street level. Lean Six Sigma discipline drove the panel methodology. Delivery ran through licensed software contracts with ESRI, Skyline, and MapInfo, composed into GSM's own product. Build on platforms you license, sell the composition.
Result
The Traffic Audit Bureau certified the product. The industry's own standard-setter accepted cell-signal measurement in place of clickers.
Billboard clients saw a 5 to 10% revenue lift, worth $300K to $500K annually per $5M of inventory. The geospatial predictive-visualization work with Skyline earned Finalist recognition from the US-Israel BIRD Foundation in 2007.
Today TAB is Geopath. Its current methodology of GPS panels, cellular data, sensor fusion, and demographic attribution is the same stack, industrialized. The measurement approach outlived the company that pioneered it. Two of the era's algorithms sit cited on an NIH biosketch. One is the TAB traffic-pattern recognition engine. The other is a cluster algorithm that linked retail pharmacy sales to CDC outbreak prediction with 3 to 14 days of early warning.
The measurement became the product; the product became the standard.
Modern Echo
This is the Center-of-Excellence commercialization loop executed once already. Find an operating problem an industry tolerates. Productize the data solution on platforms you license rather than build. Get the industry body to certify it. Price it on client value. In 2026 the equivalent motion is decision-intelligence products on governed geospatial estates. Same loop, richer stack.
Regulatory Thread
Certification is a regulator by another name. The product won because its methodology survived audit by the standard-setter on panel design, statistical discipline, and reproducibility. Measurement products live or die on whether a skeptical auditor can retrace them. That bar came from Six Sigma and pharmacy-law habits, and it is the same bar AI-derived analytics face with every enterprise client today.
Story 5
Silent Failures
The contaminated doses passed every checkpoint. Patients received them anyway. That lesson is why my AI pipelines evaluate outcomes, not gates.
Situation
Every gate reported green while contaminated doses reached patients.
1984. Two radiopharmaceutical suppliers shipped faulty Tc-99m generators after both failed to perform required Mo-99 breakthrough testing (NRC Information Notice 84-85, November 30, 1984). Contaminated doses reached hospitals. Some hospitals caught abnormal radiation readings; others never surveyed the packages or failed to recognize what they saw.
Patients were injected and had to be re-dosed to complete their procedures. The highest confirmed dose carried 234 microcuries of Mo-99 against a 5 microcurie per-dose limit, roughly 47 times over. The NRC notice names no suppliers, no hospitals, and no patient count. The incident was centered in Chicago with more than 80 patients exposed. Every gate in the chain reported green. The failure was silent.
Task
A 25-year-old pharmacy manager was designated NRC dual verification auditor. The assignment was to document the entire failure chain and design remediation protocols that would actually catch the next one.
Action
Trace the whole chain, then wedge independent verification into it.
Traced the chain end to end. Generator QC skipped at two independent suppliers, package surveys skipped or misread at hospitals, no cross-check anywhere between them. Built dual-verification protocols, an independent second measurement at the points where a single actor's error could reach a patient.
The same year, on the same border posting, came the Ciudad Juarez cobalt-60 release. About 6,000 pellets from a scrapped radiotherapy unit melted into roughly 6,000 tons of rebar across 17 Mexican states, discovered only when a truck tripped radiation sensors at Los Alamos. Contaminated steel had passed every visual inspection for weeks. Radiation checks with the NRC response team followed. Two incidents, one lesson. Detection you don't design for is detection you don't get.
Result
0
Production AI failures
5
Systems
8
Years
Remediation compliance enforced; dual verification became habit, then career doctrine. Forty years later it ships inside production AI. At NewsRx, biomedical summaries that read fluently and score well on any single metric are exactly the dangerous ones. Four blinded LLM judges cross-correlate with HHEM hallucination scoring and BERTScore, borderline cases escalate to a stronger judge, and subject-matter experts sign the output under 21 CFR Part 11 with an ALCOA+ audit trail. Zero production AI failures across five systems and eight years. The design assumption is always the same. The most dangerous failure will pass your gates, so instrument the outcome, not just the checkpoints.
The Gate Chain
1984 — Every gate reported green
The fix — independent verification wedged into the chain
Silent failures don't trip alarms. They surface where you finally look, unless you design the look.
Modern Echo
In LLM systems the silent failure is the fluent, plausible, wrong answer. It passes format checks, tone checks, and single-metric evaluations. One NewsRx summary called pregnancy outcomes safe when one of three women in the study died. Automated scoring passed it. The blinded judge layer caught it.
The 1984 doctrine maps one-to-one. Multiple independent judges (dual verification), cross-correlation between uncorrelated detectors (survey meter plus assay; HHEM plus blinded judges), escalation bands for borderline readings, and an accountable human signature before anything reaches the patient or the reader.
Regulatory Thread
The NRC didn't suggest dual verification. It mandated it, because single-actor assurance had just injured patients. Every mature safety regime converges on the same rule. Independent verification at the points of irreversible consequence. AI governance is currently rediscovering this. Teams that have lived it don't need to be convinced; they design for it on day one.
Appendix
Data to LLM: The 2026 Stack, Hands-On
Not from a course. This stack runs in production right now, and the open-source side has closed the gap enough that self-hosting is a real default choice today, not a budget compromise.
The Pipeline, Raw Data to LLM-Ready
Open source at every step. This is the working pattern behind the NewsRx production line in Story 3 and Story 5.
1
Parse
Turn PDFs, docx, and HTML into clean text. Apache Tika or Unstructured.io.
2
Chunk
Split text into retrievable pieces. Hierarchical for structured docs (3 to 5x better F1), semantic for narrative text. LangChain / LlamaIndex splitters.
3
Embed
Convert each chunk to a vector. BGE-M3 (dense, sparse, and multi-vector in one model) or Qwen3-Embedding-8B (#1 open MTEB multilingual, 32K context).
4
Serve
Run the embedding model at scale. TEI for auto-batched production serving, or raw vLLM/SGLang DIY.
5
Store + Retrieve
Index vectors for search. Qdrant or Milvus (43K+ stars, hybrid dense+sparse). Fuse with BM25 via Reciprocal Rank Fusion, then rerank with BGE-reranker-v2.
6
Generate
Feed the top reranked chunks to the LLM. Self-hosted: vLLM serving an open model (Qwen3-32B and similar).
Where Process Meets the Model
The approach behind the builds. Grounding lives in a knowledge graph beside Postgres running hybrid RAG and vector search, because ontology and semantics are part of the data, not decoration on it.
A GenAI platform in the Cortex class doesn't just draft text. It does data pulls, and the human in the loop rarely catches a bad one, because reviewers aren't taught critical thinking and a plausible number gets accepted. The bad data flows into analysis. Business, shipping, and manufacturing paths change on it. The changed outputs get written back, and the bad data is reintroduced to the very system it came from. The loop closes, and the estate starts poisoning itself.
The loop nobody instruments closes at the write-back.
The Countermeasure
Ground retrieval in ontology and semantics so a pull carries its meaning and lineage with it. That's the knowledge graph's job beside the vector index, and it's why GraphRAG matters below. Treat "the reviewer will catch it" as an unverified gate, the same gate that reported green in 1984. Instrument the write-back path, because that's the point of irreversible consequence — the 17 instrumented write-path call sites at NewsRx are exactly this, applied to a publishing estate.
What Changed in the Past 12 Months
Currency is a claim you have to re-earn every year. This is the current year's ledger.
Dense-Only Retrieval Lost
Benchmarks (BEIR, MTEB, Anthropic's own contextual retrieval work) settled it. BM25 plus dense embeddings, fused with reciprocal rank fusion, beats either alone, and a cross-encoder reranker adds another 5 to 15 points of MRR on hard queries.
12 months ago dense-only vector search was the default pattern most tutorials taught. Now hybrid search runs in 72% of production RAG systems.
Agentic RAG Went Mainstream
The biggest paradigm shift of the year. Instead of a fixed retrieve-then-generate pipeline, the model now controls retrieval itself, deciding to ask for more evidence, rewrite its own query, or stop early (Self-RAG, FLARE patterns).
12 months ago retrieval was a preprocessing step. Now it is an action the agent takes mid-reasoning — which is exactly why ungoverned data pulls became the failure mode above.
GraphRAG Hit Production Scale
Microsoft's GraphRAG (MIT-licensed) moved knowledge-graph-augmented retrieval from research papers into real, adopted production architecture.
12 months ago GraphRAG was a research pattern. Now it is a named, shipping option for teams that need entity and relationship reasoning — ontology and semantics — not just similarity search.
Open Embeddings Closed the Gap
BGE-M3 and Qwen3-Embedding-8B now rival or beat proprietary embedding APIs on MTEB, trained on massive multilingual corpora, both freely self-hostable.
12 months ago closed APIs were the safe default for embedding quality. Now a self-hosted open-weight model is a legitimate first choice.
The Amgen Build, Then vs. Now
The build-versus-buy line moves as platforms mature. Knowing when to stop building and start configuring is the judgment call.
What It Took in 2023–2024
Amgen's first GxP-validated GenAI system for FDA NDA submissions (ICH-standard CTD content, RAG-grounded, feeding Structured Content Authoring into Veeva Vault) took a 16-person build. Six Bain consultants, three data scientists, three business analysts, three medical specialist application writers, and John owning the ICH application process. Vault had no native agentic authoring then, so the RAG-to-SCA loop was architected and governed from scratch under GxP change control.
The build was custom because the platform hadn't caught up yet.
What Ships Natively, August 2026
Veeva's own Agentic Authoring application integrates natively with Vault RIM and Word, proactively drafts submissible documents, and monitors incoming data to initiate drafting. Clinical, Regulatory, and Medical AI Agents are slated for this same month, and Veeva's roadmap calls RIM+AI the standard for 2026 to 2027.
What took a 16-person custom build in 2023 is a configurable vendor feature by 2026.
The Insight, Not Just the Fact
Built today, the Amgen system likely wouldn't need a bespoke GenAI-to-SCA pipeline at all. It would configure and extend the vendor's native agents. That's not a knock on the original build. The industry converging on the same architecture two years later is validation the design was right. What it shows is that work genuinely custom-necessary one year becomes vendor commodity a couple of years later, and knowing when to stop building and start configuring is the strategist judgment call, not a technical one.
The Long Curve
"I ran plasma flow calculations on a Cray-1A for DARPA in 1980. 160 megaflops, state of the art at the time. Apple's A17 Pro does over 5 teraflops, about 31,000 times faster, in something that fits in a pocket. That's the same curve as a bespoke 2023 GenAI pipeline against a 2026 native platform feature, just compressed into two years instead of forty-five."
Some stories draw on reconstruction and first-person material where the underlying records are private. Figures are cited to SEC filings, NRC notices, and trade press where public records exist.