Skip to content

Case study

A national general insurer

A national general insurer offering motor and home cover across a single national market. Its claims triage copilot reads first notifications of loss and recommends a handling route, while settlement decisions remain with human claims handlers.

01 — In their words

“When the regulator asked how a particular claim had been triaged, we replayed the exact run in front of them. The conversation moved from assurances to evidence.”

Recorded verbatim · Attribution withheld

02 — The work

Challenge

The triage copilot's usefulness was never in doubt; its governability was. The insurer could not evidence individual triage recommendations to the standard a conduct regulator expects, could not technically guarantee that the copilot would remain advisory rather than acting on claims, and had no early warning if behaviour on edge-case claims — subsidence, disputed liability, vulnerable customers — shifted underneath a provider model update.

Solution

Hyperpriors records every triage run as a complete, replayable trace, so regulatory queries are answered from the record rather than from reconstruction. The harness bounds the copilot to read-only tools, budgets its retries, and escalates low-confidence, high-value, and vulnerability-flagged claims to human handlers, keeping settlement authority with named individuals as a technical property. Evaluation suites built on adjudicated historical claims gate every model and prompt change, and run on a schedule against the live endpoint so drift on edge cases is caught before it reaches the claims queue.

03 — The system

HYPERPRIORSCONTROL PLANEHARNESSGUARDRAILSEVAL GATESTRACESCLAIMS INTAKEFNOL · DOCUMENTSPOLICY ADMINCOVER · HISTORYFRAUD SIGNALSSCORING FEEDCLAUDE — PRIMARYMODEL CALLSPROVIDER B — DIFFCLAIMS HANDLERSSETTLEMENT AUTHORITYUNCERTAINTY ROUTES TO PEOPLEAUDIT RECORDREGULATOR-READY TRAILEVERY STEP, REPLAYABLE
Fig. 1 — System architectureInsurance · Illustrative topology

04 — Results

A hypothetical deployment scenario.

  1. 01

    Regulatory queries answered from complete, replayable traces rather than reconstructed accounts

  2. 02

    Settlement authority demonstrably retained by human handlers, with the harness technically unable to authorise a payment or close a claim

  3. 03

    Behavioural drift on edge-case claims surfaced at the evaluation gate rather than through customer complaints

05 — The record

Industry
Insurance
Region
United Kingdom
Workloads
Claims triage copilot
Disciplines
Auditing · Harness control · Evals

06 — The full account

The situation

A national general insurer handles a high volume of motor and home claims, and the first hours after notification determine much of what follows. A claims triage copilot was introduced to read first notifications of loss, match them against policy wording, flag indicators warranting investigation, and recommend a route: fast-track settlement, standard handling, or referral.

The copilot was useful from the first week. It was also, from the first week, a regulatory question. The insurer operates under conduct rules that require it to explain how customer outcomes are reached, and “a model recommended it” is not an explanation. Three concerns dominated internal review: whether any individual triage could be evidenced to the regulator, whether the copilot could ever act beyond recommendation, and whether its behaviour on awkward claims — subsidence, disputed liability, claims involving vulnerable customers — would quietly change as the underlying models were updated.

What changed

The insurer deployed the copilot inside Hyperpriors, treating the three concerns as three disciplines.

Strict auditing. Every triage run is recorded in full: model version, prompt, the policy wording and claim documents retrieved, and the recommendation produced. Any run can be replayed exactly. When the regulator asks how a claim was triaged, the answer is the trace itself, not a reconstruction assembled weeks later.

Harness control. The copilot’s tools are bounded to reading claim files and policy documents. There is no tool through which it can authorise a payment, close a claim, or contact a customer. Retries are budgeted, and low-confidence recommendations, high-value claims, and any claim flagged as involving a vulnerable customer escalate to a human handler by default. Settlement authority remains with named individuals, and the harness makes that a technical property rather than a policy statement.

Valid evals. A golden set was built from historical claims with settled, adjudicated outcomes, weighted deliberately towards the edge cases that had caused concern. The judge models scoring triage recommendations are calibrated against experienced handlers’ judgements. Every model or prompt change must pass a regression gate before promotion, and the same suite runs on a schedule against the live endpoint, so provider-side drift is caught even when nothing internal has changed.

Where it landed

Conversations with the regulator changed in character: from assurances about process to demonstrations from the record. Handlers report clearer boundaries — the copilot recommends, they decide — and drift on edge-case claims now surfaces at the evaluation gate rather than through complaints. Internally, the arrangement is described as the difference between using a model and being accountable for one.

This is a hypothetical scenario, illustrating how strict auditing, harness control, and valid evaluations apply in general insurance.

07 — Begin

Put evidence, not assurances, in front of your regulator.