Skip to content

Case study

A multi-brand online retailer

A multi-brand online retailer operating several consumer brands from a shared commerce platform, each with its own voice, pricing rules, and service policies. Its conversational assistant handles product discovery, order tracking, returns, and goodwill gestures across the full brand portfolio.

01 — In their words

“The dispute conversations changed completely. We stopped arguing about what the assistant might have said and started reading what it actually said, tool calls included.”

Recorded verbatim · Attribution withheld

02 — The work

Challenge

A shopping and service assistant operating across several brands could apply discounts and initiate refunds in free text — in effect, spending company money on its own judgement. Concessions risked exceeding brand policy, tone drifted between brands with distinct voices, and when customers disputed what the assistant had promised, the retailer held no complete record with which to reconcile the claim.

Solution

The assistant runs inside a Hyperpriors harness in which discounting and refund tools carry per-brand limits and per-order caps enforced outside the model, transient failures are retried, and anything beyond a bound escalates to a human agent with the conversation attached. Golden conversation sets scored by judges calibrated per brand gate every prompt, policy, and model change on tone and policy adherence, and full traces of every exchange and tool call allow disputes to be reconciled from the record.

03 — The system

HYPERPRIORSCONTROL PLANEHARNESSGUARDRAILSEVAL GATESTRACESSTOREFRONTS3 BRANDS · WEB · APPORDER SYSTEMREFUNDS · DISCOUNTSCATALOGUESTOCK · PRICINGCLAUDE — DIALOGUEMODEL CALLSSMALL MODEL — INTENTSERVICE TEAMDISPUTES · OVERRIDESUNCERTAINTY ROUTES TO PEOPLEAUDIT RECORDPER-BRAND EVAL SETSEVERY STEP, REPLAYABLE
Fig. 1 — System architectureRetail and e-commerce · Illustrative topology

04 — Results

A hypothetical deployment scenario.

  1. 01

    Disputed promises are settled from the conversation trace rather than by default concession, in whichever direction the record supports

  2. 02

    Concession exposure is bounded by construction, with out-of-policy requests escalated to human agents rather than resolved by the model

  3. 03

    Brand teams approve releases on per-brand eval evidence, so changes that suit one brand cannot silently degrade another

05 — The record

Industry
Retail and e-commerce
Region
Western Europe
Workloads
Conversational commerce, customer service
Disciplines
Auditing · Harness control · Evals

06 — The full account

The situation

A multi-brand online retailer operates several consumer brands from a shared platform, each with its own voice, pricing rules, and service policies. The company introduces a conversational assistant that handles product discovery, order tracking, returns, and goodwill gestures across all of them.

The commercial risks are immediate. An assistant that can apply discounts and initiate refunds is, in effect, spending company money in free text. Early internal testing surfaces three failure modes: the assistant offering concessions beyond what any brand’s policy permits, tone drifting until a premium brand sounds like its discount sibling, and disputes in which a customer reports a promise the team cannot verify because no complete record of the exchange exists.

What changed

The retailer deploys the assistant inside a Hyperpriors harness with the money-touching tools explicitly bounded. Discounting and refund tools carry per-brand limits and per-order caps enforced outside the model; transient failures in order-management calls are retried automatically; and any request exceeding a bound — an unusually large refund, a repeat concession on the same order — is escalated to a human service agent with the conversation attached, rather than left to the model’s judgement.

Behaviour is validated per brand rather than in aggregate. Golden sets of representative conversations exist for each brand, scored by judges calibrated against the brand teams’ own assessments of tone and policy adherence. A regression gate runs the full suite on every prompt, policy, or model change, so an adjustment that suits one brand cannot silently degrade another.

Every conversation and every tool call is traced end to end. When a customer disputes what the assistant offered, the service team reads the exchange as it actually occurred and reconciles the claim against the recorded tool calls. Proposed changes are replayed against past conversations before release.

Where it landed

The character of the results is operational rather than dramatic. Disputes that previously ended in default concessions are settled from the trace — in the customer’s favour when the assistant did overpromise, and in the retailer’s when it did not. Concession exposure is bounded by construction: the open question is no longer whether the assistant might exceed policy, but whether the bounds are set correctly, and the traces answer it. Brand teams, initially the most sceptical stakeholders, approve releases on eval evidence for their own brand rather than on reassurance, and releases become routine.

This is a hypothetical scenario, illustrating how strict auditing, harness control, and valid evals apply in online retail.

07 — Begin

Bound the tools that spend money on your behalf.