Skip to content

Case studies

8 case studies · All hypothetical scenarios

Where discipline pays

Hypothetical scenarios modelled on production AI as it is actually run — what evaluation gates, guardrails, and traces return when applied in earnest.

01 — Case studies

01

A Series-B fintech platform

The copilot depended on hosted models that changed underneath it.

Read the case study →

Financial services

02

A clinical documentation platform

The scribe produced fluent, well-structured notes — which was precisely the danger.

Read the case study →

Healthcare technology

03

A European industrial-equipment manufacturer

The maintenance copilot could assemble plausible repair advice that appeared in no approved procedure — tolerable in a chat window, unacceptable beside a torque specification on safety-critical machinery.

Read the case study →

Industrial manufacturing

04

A national general insurer

The triage copilot's usefulness was never in doubt; its governability was.

Read the case study →

Insurance

05

A legal-tech research company

Agent runs failed invisibly in the middle of the chain.

Read the case study →

Legal technology

06

A top-20 pharmaceutical company

First-pass literature screening for adverse-event signals had to scale across thousands of items a week without compromising accountability.

Read the case study →

Pharmaceuticals

07

A central-government benefits agency

Case files run to hundreds of pages of medical evidence, employment history, and prior correspondence, and caseworkers spent much of each day assembling a picture of the case before judgement could begin.

Read the case study →

Public sector

08

A multi-brand online retailer

A shopping and service assistant operating across several brands could apply discounts and initiate refunds in free text — in effect, spending company money on its own judgement.

Read the case study →

Retail and e-commerce