Skip to content

Case study

A central-government benefits agency

A central-government agency administering means-tested benefits. Its casework summarisation service condenses lengthy case files into structured, cited summaries for the officials who take entitlement decisions.

01 — In their words

“At tribunal the question is always what was before the decision-maker. Now we can replay the precise summary the caseworker saw, exactly as it was produced.”

Recorded verbatim · Attribution withheld

02 — The work

Challenge

Case files run to hundreds of pages of medical evidence, employment history, and prior correspondence, and caseworkers spent much of each day assembling a picture of the case before judgement could begin. Summarisation assistance was wanted, but the agency's decisions are routinely tested at tribunal, and it could not accept a system whose contribution to a decision was irreproducible on appeal, nor one that might in practice take decisions officials were nominally taking.

Solution

Hyperpriors stores every summary with its complete trace — model version, documents supplied, prompt, and the output the caseworker actually saw — so any summary in scope of an appeal can be replayed exactly as it ran. The harness gives the model read access to the case file and nothing else: no tool can record a determination, update a claim status, or generate correspondence, and files the model cannot summarise confidently are routed to caseworkers unassisted. Golden sets built from tribunal-tested cases, scored by judges calibrated against experienced decision-makers, gate every model and prompt change before it reaches production.

03 — The system

HYPERPRIORSCONTROL PLANEHARNESSGUARDRAILSEVAL GATESTRACESCASE MANAGEMENTEVIDENCE BUNDLESPOLICY LIBRARYSTATUTE · GUIDANCEAPPROVED MODEL — SOVEREIGN HOSTINGMODEL CALLSCASEWORKERSDECISION AUTHORITY RETAINEDUNCERTAINTY ROUTES TO PEOPLEAUDIT RECORDAPPEALS-GRADE TRAILEVERY STEP, REPLAYABLE
Fig. 1 — System architecturePublic sector · Illustrative topology

04 — Results

A hypothetical deployment scenario.

  1. 01

    Every summary that informed a decision can be replayed exactly for appeal and tribunal preparation, turning reconstruction into retrieval

  2. 02

    Entitlement decisions demonstrably taken by officials, with the harness technically incapable of recording a determination

  3. 03

    Summarisation drift on complex cases caught at the regression gate rather than discovered through appeals

05 — The record

Industry
Public sector
Region
United Kingdom
Workloads
Casework summarisation
Disciplines
Auditing · Harness control · Evals

06 — The full account

The situation

A central-government benefits agency processes claims whose case files run to hundreds of pages: medical evidence, employment history, prior correspondence, and earlier decisions. Caseworkers were spending much of each day assembling a picture of the case before any judgement could begin. A summarisation service was proposed: a model would read the file, produce a structured summary with citations back to source documents, and highlight the evidence most relevant to the entitlement criteria.

The proposal met immediate and justified scrutiny. Decisions in this domain affect people’s incomes and are routinely tested at tribunal. Two constraints were set before any deployment: the model must never take, or appear to take, a decision; and every summary that informed a decision must be reproducible in full if that decision were later appealed.

What changed

The agency deployed the service on Hyperpriors, with both constraints made mechanical rather than procedural.

Strict auditing. Every summary is stored with its complete trace: the model version, the documents supplied, the prompt, and the output the caseworker actually saw. When a decision is appealed, the summary can be replayed exactly as it ran, so tribunal preparation starts from the record rather than from recollection. Accountability is specific: which summary, produced by which configuration, was before which official.

Harness control. The harness grants the model read access to the case file and nothing else. It has no tool that can record a determination, update a claim status, or generate correspondence to a claimant. Its output is a summary with citations; the entitlement decision is taken by an official, and the system is incapable of taking it instead. Failed runs retry within a bounded budget, and files the model cannot summarise confidently are routed to caseworkers unassisted, flagged as such.

Valid evals. The golden set was built from cases whose outcomes had been tested at tribunal — cases where the correct reading of the evidence has survived independent scrutiny rather than merely being asserted. Summaries are scored for faithfulness to the file and for whether they surface the evidence a tribunal later found decisive. The judges are calibrated against experienced decision-makers, and a regression gate blocks any model or prompt change that degrades performance on these tribunal-tested cases.

Where it landed

Caseworkers begin from an evidenced summary rather than a cold file, while the decision itself remains visibly theirs. Appeal responses draw on replayable records instead of reconstructed narratives, and the agency can state precisely what the model did and did not contribute to any decision under challenge. Behavioural drift is now a gated event, not a discovery made through appeals.

This is a hypothetical scenario, illustrating how strict auditing, harness control, and valid evaluations apply in public-sector casework.

07 — Begin

Assist the caseworker without ever taking the decision.