Ranjan Kumar Agentic AI Series · Book 3 · First edition
Harness Engineering for Production AI Systems, by Ranjan Kumar

agent = harness(model)

Harness Engineering for Production AI Systems

Sensors, Gates, and Bounds for LLM Applications That Have to Work When the Model Doesn’t.

The harness is a control system.

Every component in it is one of four things: a sensor that observes, a comparator that checks against a standard, a gate that decides whether a proposal becomes an action, or a bound that caps how far the system can go. If a component is none of these, it is not harness. It is application code.

Read a sample with Amazon’s Look Inside on either listing. All the code is open: the companion repository on GitHub, one tag per chapter.

17Chapters
17Artifacts
4Control roles
389Passing tests
0API keys to run them

The problem

Your LLM application works in the demo and fails in production.

The first fix everyone reaches for is a better model. That reflex is understandable and it is usually wrong. The gap between the inputs a system was shown and the inputs it will meet is a property of the input distribution, not of the model, so it does not close when the model improves. Work in that gap is handled by the code around the model, or it is not handled at all.

Almost all the published evidence that harnesses work comes from coding agents, and that is a selection effect rather than a coincidence: coding is the one domain with a free, immediate, unambiguous verdict on every run. Most domains have no such oracle. This book is written for those domains.

Why not “guardrails”

The book does not use the word, because it names a feeling rather than a mechanism. A postmortem asks which component could have prevented the action. “It had guardrails” does not answer that question. “Five components, none of which could prevent it” does.

The running example

One system carries the whole book: a billing support agent that reads tickets, looks up invoices, drafts replies, applies credits, issues refunds, and closes tickets, with a human reviewing every reply by day and nobody there overnight. issue_refund moves money and has no undo. It is a constructed scenario, and the book says so.

The spine of the book

Four roles. Defined by power, not by position.

Every chapter from Part II onward opens by naming which of the four roles it serves. Open a role to see what it can do, what the book builds for it, and how it fails without raising anything.

In the order Part II builds them
Sensor Observes, and shapes what is observed what it is · what the book builds · how it fails

What it is

Its output is a record or a transformed input. It cannot stop anything. What a sensor drops is what the model will never see: a context assembler that trims the oldest turns to fit a budget has just decided what the system knows.

What the book builds

Ch 4, the input policy · Ch 5, the context budget spec · Ch 6, the signal schema.

How it fails silently

The context assembler drops what the model needed to see. Answers get worse. Nothing raises.

Comparator Checks an observation against a standard what it is · what the book builds · how it fails

What it is

It emits a verdict and cannot stop anything; a judgment is inert. The schema validator is a comparator, as long as it only reports. Give it the power to reject the request and it becomes a gate.

What the book builds

Ch 7, the repair ladder · Ch 8, the resilience config.

How it fails silently

The schema validator returns passed on its own exception. The system permits more. Nothing raises.

Gate Decides whether a proposal becomes an action what it is · what the book builds · how it fails

What it is

It reads the proposal, and it can refuse. The approval tier on issue_refund is a gate, and so is a human checkpoint. Because it reads the proposal, it can only refuse the proposals its author knew how to read.

What the book builds

Ch 9, the gate policy · Ch 10, the tool-safety matrix · Ch 11, the injection-resistant merge gate.

How it fails silently

Stuck open, the refund gate allows everything it is shown and the system permits more. Stuck closed, it escalates everything, permits nothing, and looks like a busy night.

Bound Caps how far the system can go what it is · what the book builds · how it fails

What it is

It does not read the proposal. It reads a counter, or a clock, and it throws. Gates hold against the failures you enumerated. Bounds hold against the ones you did not, and the second set is larger and is the one that pages you.

What the book builds

Ch 12, the checkpoint spec · Ch 13, the budget config · Ch 14, the authority band table.

How it fails silently

The tool-call ceiling’s counter stops being incremented. The system permits more. Nothing raises.

A component that is none of the four is application code. Chapter 17 is about the failure all four share: a control that has stopped working produces, on traffic that does not need it, exactly the output of a control that is working.

Named concepts

Tests you can run in a meeting.

The book names thirty-eight concepts and defines every one in Appendix A. Six of them, quoted from their definitions.

Ch 1

The Demo Cliff

The gap between the inputs a system was shown and the inputs it will meet. It is a property of the input distribution, not of the model, so it does not close when the model improves.

Ch 3

The Role Test

A three-question procedure that classifies any component as a sensor, comparator, gate, bound, or application code. Can it prevent an action; does its decision read the proposal’s content; does it emit a verdict against a standard.

Ch 9

Consequence Gating

A gate does not decide whether an action is correct, because nobody in the system can. It decides whether the system is allowed to be wrong about it.

Ch 10

The Expressible Set

The set of actions a model can propose at all. The cheapest safety work available is moving actions out of the set rather than controlling them inside it.

Ch 16

The Oracle Deficit

The property of a test-suite verdict that a substitute gives up. Every scorer forfeits at least one, and which one it forfeits is what its number may not be used for.

Ch 17

The Silent Pass

A control that has stopped working produces, on traffic that does not need it, exactly the output of a control that is working. Production cannot distinguish a live control from a dead one. Only exercise can.

How it is built

Seventeen chapters, seventeen artifacts you can copy.

Every numbered chapter ships one named artifact: twelve declarative policy files, three Python modules, one test file, and one figure. Every one that is a file is enforced by code in the same repository.

Every control appears three times, in this order: as a plain object with no framework import, as the artifact it reads, and as the wiring, the LangGraph node that puts it in a graph. Read only the first two passes and you have a harness.

Appendix B lists all seventeen. One row:
ChArtifactFileLines
9The gate policypolicies/gate-policy.toml135

Every measurement is pinned to LangGraph 1.2.11 and Python 3.13, in September 2026.

Who this book is for

You run an LLM application in production, and it is not a coding agent.

AI, ML, and platform engineers who run an LLM application in production, or are about to, and are not building a coding agent.

If you are an engineering leader deciding what “reliable” has to mean before you sign off, Chapters 1, 3, 9, and 14 are the four to read.

What it is not

If you want Claude Code or Codex internals, this is the wrong book.

Contents

Four parts, seventeen chapters, a closer, four appendices.

Part I

The Harness Is a Control System

  • 1 The Demo Cliff Is Not a Model Problem
  • 2 Agent = Model + Harness, and the Trust Boundary
  • 3 Sensors, Comparators, Gates, Bounds

Part II

The Four Roles

  • 4 Input Defense and Normalisation
  • 5 Context Assembly as Managed Infrastructure
  • 6 Signals, and What the Harness Must Emit
  • 7 Validation and Repair
  • 8 Retry, Fallback, Circuit Breaking
  • 9 Gated Execution
  • 10 Deterministic Constraint Systems
  • 11 The Gate Had Nothing to Refuse
  • 12 State, Checkpoints, Resume
  • 13 Cost, Depth, Width
  • 14 Authority Bands

Part III

Constructing the Oracle

  • 15 Testing the Harness Without the Model, and With a Bad One
  • 16 Attributing Failures to a Layer

Part IV

When the Harness Fails

  • 17 Failure Modes of the Control Layer Itself
  • · The Framework Seam
A · Vocabulary Reference B · The Seventeen Artifacts C · Sources D · The Companion Repository

The companion repository

One tag per chapter. No API key, no network, no model.

One Python package, harness, grows across the book. There is a git tag per chapter, so checking out a tag gives you the package exactly as the book describes it at the end of that chapter, with the tests that were passing then. At the release tag, code-v1.0, the suite stands at 389 passing tests. MIT License.

# Check out a chapter's end state
git checkout ch01-demo-cliff

# Install and run its tests
pip install -e ".[dev]"
pytest

The test suite runs with no API key, no network, and no model, at every tag. If the model is a seam you can cut, most of what you need to know about a harness is knowable without it.

github.com/ranjankumar-gh/harness-engineering-code

The author

Ranjan Kumar

Ranjan Kumar builds and writes about production AI systems. He holds an M.Tech in Artificial Intelligence from IIT Jodhpur and has 20 years of experience in the field. He writes continuing series on harness engineering, agent security architecture, RAG engineering, and the operational side of running agents at scale.

This book is the third in his Agentic AI Series:

Errata and feedback

A book that prints its own source, and a repository that keeps moving, produce one characteristic error between them: a listing that was true when it went to press and is not true at the tip today. Appendix D lists every known case.

If you find another, or anything else that is wrong, the fastest route is an issue on the companion repository.