agent = harness(model)
Harness Engineering for Production AI Systems
Sensors, Gates, and Bounds for LLM Applications That Have to Work When the Model Doesn’t.
Every component in it is one of four things: a sensor that observes, a comparator that checks against a standard, a gate that decides whether a proposal becomes an action, or a bound that caps how far the system can go. If a component is none of these, it is not harness. It is application code.
Read a sample with Amazon’s Look Inside on either listing. All the code is open: the companion repository on GitHub, one tag per chapter.
The problem
Your LLM application works in the demo and fails in production.
The first fix everyone reaches for is a better model. That reflex is understandable and it is usually wrong. The gap between the inputs a system was shown and the inputs it will meet is a property of the input distribution, not of the model, so it does not close when the model improves. Work in that gap is handled by the code around the model, or it is not handled at all.
Almost all the published evidence that harnesses work comes from coding agents, and that is a selection effect rather than a coincidence: coding is the one domain with a free, immediate, unambiguous verdict on every run. Most domains have no such oracle. This book is written for those domains.
Why not “guardrails”
The book does not use the word, because it names a feeling rather than a mechanism. A postmortem asks which component could have prevented the action. “It had guardrails” does not answer that question. “Five components, none of which could prevent it” does.
The running example
One system carries the whole book: a billing support agent that reads tickets, looks
up invoices, drafts replies, applies credits, issues refunds, and closes tickets, with
a human reviewing every reply by day and nobody there overnight. issue_refund
moves money and has no undo. It is a constructed scenario, and the book says so.
The spine of the book
Four roles. Defined by power, not by position.
Every chapter from Part II onward opens by naming which of the four roles it serves. Open a role to see what it can do, what the book builds for it, and how it fails without raising anything.
Sensor Observes, and shapes what is observed what it is · what the book builds · how it fails
What it is
Its output is a record or a transformed input. It cannot stop anything. What a sensor drops is what the model will never see: a context assembler that trims the oldest turns to fit a budget has just decided what the system knows.
What the book builds
Ch 4, the input policy · Ch 5, the context budget spec · Ch 6, the signal schema.
How it fails silently
The context assembler drops what the model needed to see. Answers get worse. Nothing raises.
Comparator Checks an observation against a standard what it is · what the book builds · how it fails
What it is
It emits a verdict and cannot stop anything; a judgment is inert. The schema validator is a comparator, as long as it only reports. Give it the power to reject the request and it becomes a gate.
What the book builds
Ch 7, the repair ladder · Ch 8, the resilience config.
How it fails silently
The schema validator returns passed on its own exception. The system permits more. Nothing raises.
Gate Decides whether a proposal becomes an action what it is · what the book builds · how it fails
What it is
It reads the proposal, and it can refuse. The approval tier on
issue_refund is a gate, and so is a human checkpoint. Because it reads the
proposal, it can only refuse the proposals its author knew how to read.
What the book builds
Ch 9, the gate policy · Ch 10, the tool-safety matrix · Ch 11, the injection-resistant merge gate.
How it fails silently
Stuck open, the refund gate allows everything it is shown and the system permits more. Stuck closed, it escalates everything, permits nothing, and looks like a busy night.
Bound Caps how far the system can go what it is · what the book builds · how it fails
What it is
It does not read the proposal. It reads a counter, or a clock, and it throws. Gates hold against the failures you enumerated. Bounds hold against the ones you did not, and the second set is larger and is the one that pages you.
What the book builds
Ch 12, the checkpoint spec · Ch 13, the budget config · Ch 14, the authority band table.
How it fails silently
The tool-call ceiling’s counter stops being incremented. The system permits more. Nothing raises.
A component that is none of the four is application code. Chapter 17 is about the failure all four share: a control that has stopped working produces, on traffic that does not need it, exactly the output of a control that is working.
Named concepts
Tests you can run in a meeting.
The book names thirty-eight concepts and defines every one in Appendix A. Six of them, quoted from their definitions.
Ch 1
The Demo Cliff
The gap between the inputs a system was shown and the inputs it will meet. It is a property of the input distribution, not of the model, so it does not close when the model improves.
Ch 3
The Role Test
A three-question procedure that classifies any component as a sensor, comparator, gate, bound, or application code. Can it prevent an action; does its decision read the proposal’s content; does it emit a verdict against a standard.
Ch 9
Consequence Gating
A gate does not decide whether an action is correct, because nobody in the system can. It decides whether the system is allowed to be wrong about it.
Ch 10
The Expressible Set
The set of actions a model can propose at all. The cheapest safety work available is moving actions out of the set rather than controlling them inside it.
Ch 16
The Oracle Deficit
The property of a test-suite verdict that a substitute gives up. Every scorer forfeits at least one, and which one it forfeits is what its number may not be used for.
Ch 17
The Silent Pass
A control that has stopped working produces, on traffic that does not need it, exactly the output of a control that is working. Production cannot distinguish a live control from a dead one. Only exercise can.
How it is built
Seventeen chapters, seventeen artifacts you can copy.
Every numbered chapter ships one named artifact: twelve declarative policy files, three Python modules, one test file, and one figure. Every one that is a file is enforced by code in the same repository.
Every control appears three times, in this order: as a plain object with no framework import, as the artifact it reads, and as the wiring, the LangGraph node that puts it in a graph. Read only the first two passes and you have a harness.
| Ch | Artifact | File | Lines |
|---|---|---|---|
| 9 | The gate policy | policies/gate-policy.toml | 135 |
Every measurement is pinned to LangGraph 1.2.11 and Python 3.13, in September 2026.
Who this book is for
You run an LLM application in production, and it is not a coding agent.
AI, ML, and platform engineers who run an LLM application in production, or are about to, and are not building a coding agent.
If you are an engineering leader deciding what “reliable” has to mean before you sign off, Chapters 1, 3, 9, and 14 are the four to read.
What it is not
If you want Claude Code or Codex internals, this is the wrong book.
Contents
Four parts, seventeen chapters, a closer, four appendices.
Part I
The Harness Is a Control System
- 1 The Demo Cliff Is Not a Model Problem
- 2 Agent = Model + Harness, and the Trust Boundary
- 3 Sensors, Comparators, Gates, Bounds
Part II
The Four Roles
- 4 Input Defense and Normalisation
- 5 Context Assembly as Managed Infrastructure
- 6 Signals, and What the Harness Must Emit
- 7 Validation and Repair
- 8 Retry, Fallback, Circuit Breaking
- 9 Gated Execution
- 10 Deterministic Constraint Systems
- 11 The Gate Had Nothing to Refuse
- 12 State, Checkpoints, Resume
- 13 Cost, Depth, Width
- 14 Authority Bands
Part III
Constructing the Oracle
- 15 Testing the Harness Without the Model, and With a Bad One
- 16 Attributing Failures to a Layer
Part IV
When the Harness Fails
- 17 Failure Modes of the Control Layer Itself
- · The Framework Seam
The companion repository
One tag per chapter. No API key, no network, no model.
One Python package, harness, grows across the book. There is a git tag
per chapter, so checking out a tag gives you the package exactly as the book describes
it at the end of that chapter, with the tests that were passing then. At the
release tag, code-v1.0, the suite stands at 389 passing
tests. MIT License.
# Check out a chapter's end state git checkout ch01-demo-cliff # Install and run its tests pip install -e ".[dev]" pytest
The test suite runs with no API key, no network, and no model, at every tag. If the model is a seam you can cut, most of what you need to know about a harness is knowable without it.
The author
Ranjan Kumar
Ranjan Kumar builds and writes about production AI systems. He holds an M.Tech in Artificial Intelligence from IIT Jodhpur and has 20 years of experience in the field. He writes continuing series on harness engineering, agent security architecture, RAG engineering, and the operational side of running agents at scale.
This book is the third in his Agentic AI Series:
Errata and feedback
A book that prints its own source, and a repository that keeps moving, produce one characteristic error between them: a listing that was true when it went to press and is not true at the tip today. Appendix D lists every known case.
If you find another, or anything else that is wrong, the fastest route is an issue on the companion repository.