
Latest · 21 Jul 2026
Priors over priors
Why we are called Hyperpriors: what a hyperprior means in Bayesian terms, and why the control plane for production AI is exactly that — beliefs about how a system should hold beliefs.
Safety · Observability
Journal
11 essays · Latest 21 Jul 2026
Essays on harnesses, evaluation, guardrails, and observability — the discipline of shipping AI systems you can stand behind.
01 — Index

Latest · 21 Jul 2026
Why we are called Hyperpriors: what a hyperprior means in Bayesian terms, and why the control plane for production AI is exactly that — beliefs about how a system should hold beliefs.
Safety · Observability
An agent harness is the structured runtime between a model and the world: tool mediation, context management, retries, fallbacks, and human escalation. A prompt and a while-loop is not an architecture.
Harnesses · Safety

A five-stage maturity model for LLM evaluation practice — from ad hoc spot checks to continuous production evaluation — with the failure modes of each stage and the exit criteria that mark genuine progress to the next.
Evals · Whitepaper

The context window is the scarcest resource in an agent system. A whitepaper-length treatment of budgeting it: retrieval allocation, tool-result summarisation, state across steps, and defending against context poisoning.
Prompting · Whitepaper

Evals are the unit test suite of AI systems. How regression gates, golden sets, and honest LLM-as-judge practice keep model behaviour shippable — and why eval suites rot if you let them.
Evals

No single control makes an LLM product safe, and none needs to. A layered architecture — input mediation, capability scoping, output enforcement, human review, audit — works because the layers fail independently. A whitepaper on building it deliberately.
Safety · Whitepaper

Audit trails, data residency, model-change management, validation documentation, and incident response for model-driven features — and how each requirement maps onto control-plane primitives you can actually operate.
Enterprise · Whitepaper

A prompt is the interface contract between your deterministic system and a probabilistic component. That means version control, review, separation of concerns, and eval coverage — the same discipline you apply to any other interface.
Prompting

How to make evals block merges the way tests do: golden set construction, threshold design, flake management, and the gate discipline that makes provider migrations survivable.
Evals

Every enterprise deploying LLM features ends up building an eval harness, a trace store, and a guardrail layer. Most build them badly, twice. Where the ownership boundary should actually sit.
Enterprise

A production AI system does not need to be right every time; it needs to know when it might be wrong and hand those cases to a person. Confidence signals, threshold calibration, escalation ergonomics, and the discipline of not crying wolf.
Safety

02 — Signal
There is no email list on this site yet — we will not put your address in a query string and pretend a subscription happened. Read the journal here, or follow the RSS feed.