Case studies
8 case studies · All hypothetical scenarios
Where discipline pays
Hypothetical scenarios modelled on production AI as it is actually run — what evaluation gates, guardrails, and traces return when applied in earnest.
01 — Case studies
01
A Series-B fintech platform
The copilot depended on hosted models that changed underneath it.
Read the case study →
Financial services
02
A clinical documentation platform
The scribe produced fluent, well-structured notes — which was precisely the danger.
Read the case study →
Healthcare technology
03
A European industrial-equipment manufacturer
The maintenance copilot could assemble plausible repair advice that appeared in no approved procedure — tolerable in a chat window, unacceptable beside a torque specification on safety-critical machinery.
Read the case study →
Industrial manufacturing
04
A national general insurer
The triage copilot's usefulness was never in doubt; its governability was.
Read the case study →
Insurance
05
A legal-tech research company
Agent runs failed invisibly in the middle of the chain.
Read the case study →
Legal technology
06
A top-20 pharmaceutical company
First-pass literature screening for adverse-event signals had to scale across thousands of items a week without compromising accountability.
Read the case study →
Pharmaceuticals
07
A central-government benefits agency
Case files run to hundreds of pages of medical evidence, employment history, and prior correspondence, and caseworkers spent much of each day assembling a picture of the case before judgement could begin.
Read the case study →
Public sector
08
A multi-brand online retailer
A shopping and service assistant operating across several brands could apply discounts and initiate refunds in free text — in effect, spending company money on its own judgement.
Read the case study →
Retail and e-commerce