Skip to content

Case study

A top-20 pharmaceutical company

A top-20 pharmaceutical company whose pharmacovigilance organisation screens the published literature for adverse-event reports across a broad product portfolio. Its literature-review assistant performs first-pass screening and case triage under medical reviewer oversight, within regulatory timelines.

01 — In their words

“When an inspector asks why an item was closed two years ago, the correct answer is a record, not a recollection. That is the standard we built to.”

Recorded verbatim · Attribution withheld

02 — The work

Challenge

First-pass literature screening for adverse-event signals had to scale across thousands of items a week without compromising accountability. A missed signal is a patient-safety failure and a regulatory finding; a screening decision that cannot be reconstructed under inspection is nearly as serious; and a model update that quietly shifts screening sensitivity is a risk the safety organisation cannot carry unknowingly.

Solution

The literature-review assistant runs on Hyperpriors with the audit trail as the primary artefact: every disposition records the source item, extracted entities, reasoning, and outcome, is replayable under the configuration that produced it, and is signed off by a named medical reviewer. The harness confines the assistant to bounded retrieval and structured-extraction tools and routes uncertainty — ambiguous causality, novel drug-event combinations, low-confidence extractions — to medical reviewers by rule. Evaluation suites built from historically adjudicated cases gate every model and prompt change, weighted towards sensitivity.

03 — The system

HYPERPRIORSCONTROL PLANEHARNESSGUARDRAILSEVAL GATESTRACESLITERATURE FEEDSJOURNALS · ABSTRACTSSAFETY DATABASEHISTORICAL SIGNALSCLAUDE — EXTRACTIONMODEL CALLSJUDGE MODEL — CALIBRATEDMEDICAL REVIEWERSNAMED SIGN-OFFUNCERTAINTY ROUTES TO PEOPLEAUDIT RECORDGxP-STYLE RECORDEVERY STEP, REPLAYABLE
Fig. 1 — System architecturePharmaceuticals · Illustrative topology

04 — Results

A hypothetical deployment scenario.

  1. 01

    Every screening disposition carries a complete, replayable decision trail signed off by a named medical reviewer

  2. 02

    Ambiguous and novel drug-event combinations reach medical reviewers by rule rather than by chance, and the assistant cannot close an item on its own authority

  3. 03

    Model and prompt changes are gated against historically adjudicated cases, so shifts in screening sensitivity are known before deployment rather than discovered after

05 — The record

Industry
Pharmaceuticals
Region
Global, headquartered in Europe
Workloads
Literature screening, pharmacovigilance case triage
Disciplines
Auditing · Harness control · Evals

06 — The full account

The situation

A top-20 pharmaceutical company screens the published literature for adverse-event reports as part of its pharmacovigilance obligations. The volume is unforgiving: thousands of abstracts and articles across the portfolio each week, any of which may contain a case that must be identified, assessed, and processed within regulatory timelines.

The company builds a literature-review assistant to perform first-pass screening — extracting suspected drug-event combinations, assembling candidate cases, and proposing a disposition for each item. The obstacles to production use are not capability but accountability. A missed signal is a patient-safety failure and a regulatory finding; a screening decision that cannot be reconstructed under inspection is very nearly as serious; and a model update that quietly shifts screening sensitivity is a risk the safety organisation cannot carry unknowingly.

What changed

The company runs the assistant on Hyperpriors with the audit trail treated as the primary artefact. Every screening decision records the source item, the extracted entities, the assistant’s reasoning, and the proposed disposition, and any past decision can be replayed under the configuration that produced it. Each accepted disposition is signed off by a named medical reviewer, so accountability rests with a person, never with the pipeline.

The harness confines the assistant to bounded retrieval and structured-extraction tools; it cannot close or dismiss an item on its own authority. Uncertainty is routed rather than resolved: ambiguous causality language, novel drug-event combinations, and low-confidence extractions escalate to medical reviewers by rule, with transient retrieval failures retried before anything is allowed to fall out of the queue.

Validation rests on history. The evaluation suites are built from historically adjudicated cases — items the safety team assessed and closed years earlier, with known outcomes — and judge models are calibrated against those adjudications rather than against intuition. A regression gate runs the suites on every model or prompt change, with particular weight on sensitivity: a change that improves throughput while missing cases the historical record says must be found does not ship.

Where it landed

Illustratively, the effect is a redistribution of expert attention. Medical reviewers spend their time on the genuinely ambiguous items the harness routes to them, rather than on first-pass triage of the entire intake. Inspection preparation changes character: screening decisions are demonstrated from replayable records rather than reconstructed from correspondence. And the safety organisation adopts model improvements deliberately, because the evaluation gate states in advance what a change does to screening behaviour on cases whose correct answer is already known.

This is a hypothetical scenario, illustrating how strict auditing, harness control, and valid evals apply in pharmaceutical pharmacovigilance.

07 — Begin

Make every screening decision inspectable.