X-Ops

The agent harness: how to build control planes that turn an LLM into a production-ready system

Vinoth Govindarajan, an OpenAI engineer working on data and AI infrastructure, delivered at InfoQ one of the most operational talks of the year on agentic systems in production. The central thesis is direct: an AI agent in production is not a model with a prompt. It is a harness — a system around the model that decides what executes, in what order, with what authority, and leaves a verifiable receipt of what happened. The phrase that summarises the entire talk and that any platform engineering team should write on their office wall is this: "A model proposes, the harness commits, and the receipts proves it." For teams putting agents into production in 2026 — from SREs automating incident response to customer support with conversational agents — this distinction between model and harness is what separates a demo from a system you can operate.

What is a harness and why the model is not enough

Govindarajan starts with a metaphor that runs through the entire talk: a car. The model is the engine. It matters — nobody buys a production car looking only at horsepower — but what you really care about is steering, brakes, dashboard, black box and transmission. Agents are the same: the model gives you capability, but the harness gives you control. A powerful engine with no brakes is not autonomy, it is a liability with good acceleration.

The operational definition of harness he offers is the following: the system around the model that lets the model output more safely in the real world. That system has to decide whether a model proposal belongs to the write state, whether the mutation is ordered, whether the work is bounded, whether the authority is valid, and whether the outcome can be proven. Five questions, five decisions the harness has to make before any model action touches persistent state.

The analogy with distributed systems is not decorative. Govindarajan comes from building distributed systems at Uber and Apple, and the rest of the talk is structured as a conversation between what reliability engineers already know (idempotency, retries, locks, ordering, state boundaries) and what agents change. The short answer: the failure modes are the same, but agents make them easier to trigger and harder to explain. A traditional distributed system fails in known, reproducible ways; an agentic system fails in ways that depend on context, the model's internal state, and the order in which multiple agents decide to cooperate.

The three operational rules you have to remember

If you can only take three ideas from the talk, Govindarajan formulates them like this:

First: own the state. A fact needs one owner and one replay path. If your system cannot reconstruct the fact later, it does not really own the fact. This is the rule with the most impact on design: in traditional systems, state lives in a database with ACID transactions; in agentic systems, "state" includes not only the database, but the agent's memory, the conversation transcript, workflow annotations, and side effects the agent already triggered. All those components have to be replayable from a known point, not just restorable from a backup.

Second: order the mutation. Concurrency is fine. Accidental interleaving is not. Agents can turn out work in parallel, read in parallel, call sub-agents or tools in parallel, but the shared state needs one mutable commit path. This sounds like standard distributed systems advice, and it is. The difference in agentic systems is that "order" does not only come from the database — it also comes from the logical order of the conversation. If an agent responds to a user and then updates the database, but the response to the user is confirmed before the database is updated, there is a window where the user sees a state the system does not have. That window is invisible to traditional consistency mechanisms.

Third: prove the action. The transcript is not the receipt. The transcript can tell you what the agent said or what the model intended. You need to know what was attempted, what was approved, and what was committed at the user-visible edge. The difference between "the model said it would remember something" and "the system remembered something" is exactly the difference between a demo and a production-ready system.

The OpenClaw case: when the system forgets what the user asked

Govindarajan uses OpenClaw as a public case study. The example that opens the talk is one that any team operating agents has seen in production: the user asks the agent to remember something for the next turn — a refund, a preference, a context. The agent responds "I will remember it." The user sees the response, everything appears to have worked. But behind the scenes, the system could not reliably reconstruct the future for the next turn. The action happened, but the record that should have made part of the agent's durable memory did not. For production agents, this kind of bug matters more than a crash, because the transcript can look coherent but the reality changed in the wrong order — or was not recorded at all.

Govindarajan is specific about why this kind of failure is worse than a crash. A crash is annoying, but at least it gives you a boundary. You see something stopped. You see an error. You can replay from the last known good point. Silent success is worse: it is a lie. The channel says success, the user sees something happened, but the operator has no reason to doubt and the system lost part of its memory. The next time the agent responds, it will do so with an incomplete context, possibly confidently, possibly incorrectly.

The specific OpenClaw example he cites has exactly that shape: the user-visible delivery path looked healthy while the turn was not recorded in the persistent path the future context would depend on. The future context inherited a hole. It may answer confidently, but it is reasoning over an incomplete record. This is why he does not want to start with a model benchmark: a benchmark can tell you about model behaviour, it cannot tell you whether the persistent edge and the delivery edge agree.

The three production questions you have to ask yourself

Once a system can send messages, update databases, run commands or trigger workflows, the production questions change. Govindarajan reduces them to three:

Who owned the state? Which memory pipeline or workflow owned the state, which was the source of truth, who committed first. When two events arrive together, you know which decided the order. If the answer is "I don't know, it depends on the agent," you have a problem.

Who can show what happened? Not what the model intended, not what the agent said. What the user-visible edge persisted. This translates in practice to: your system has an audit log that is not the model's transcript. It is a separate log, signed, with timestamps, that says exactly what state changed, when, and for what user or system action.

These are not model questions. They are production questions. The distinction matters because many teams, when starting to deploy agents, try to answer them with model tools: better prompt, better fine-tuning, better RAG. Those tools can help, but they do not solve state ownership problems. The model does not know what it owns; the harness does, or should.

How this translates into code you can write today

There are five concrete patterns the talk implies and that any team can start implementing without waiting for a new framework.

Single commit path. Any mutation to persistent state goes through a single function, registered, with verified idempotency. If the agent wants to update the database, it does not do it directly: it does it through a service that records who requested the change, what validated the authority, and what was confirmed. The commit service is both the bottleneck (for concurrency control) and the source of truth (for audit). Practical implementation: an HTTP endpoint or a queue with exactly one consumer, not multiple.

Stable identifiers for each fact. A fact is not a row in a table: it is a stable identifier that survives migrations, column renames, schema reorganisations. When an agent executes an action, that action has an ID that is maintained in logs, in metrics, in alerts, and in the transcript. If in six months you need to answer "what happened on March 14 at 10:32," the ID is in all the relevant places.

Verifiable receipt for each action. Each action executed by an agent produces a receipt that includes: the action ID, the input that generated it, the output it produced, the timestamp, the identifier of the user who authorised it (or of the system that approved it automatically), and the cryptographic signature of the executor. The receipt is what allows you to reproduce what happened without having to trust the model's transcript. It is the equivalent of a financial transaction receipt: if you have the receipt, you can prove the operation; if you do not, it does not exist.

Reorder buffer for concurrent mutations. When multiple sub-agents or tools execute in parallel, their mutations are serialised through a buffer that applies logical order before applying physical order. This avoids race conditions that appear when two agents modify the same record simultaneously, and allows reproducing state at any point in logical order, not only chronological order.

Approval policies as code. The decisions of what an agent can do without human approval and what requires confirmation are encoded in versioned files, not in model prompts. This allows reviewing who changed which policy, enforcing decisions consistently, and separating business logic (what is important to approve) from model behaviour (how to reason about the approval).

What your team should measure

The talk ends with a list of metrics that any platform engineering team should start tracking if they operate agents in production. Govindarajan does not enumerate them with exact names, but they can be inferred from the rest of the argument.

Silent success vs confirmed success rate. What fraction of agent actions completed according to the transcript but were not confirmed in the persistent log. A non-zero rate indicates a harness bug, not a model bug. This metric is built by cross-referencing the transcript log with the persistent log, and should be monitored like an API error rate is monitored.

Latency between transcript and persistence. How much time passes between the agent emitting a response to the user and the system confirming the state was persisted. A high or variable latency indicates the persistent path has bottlenecks or dependencies that may fail. In systems where consistency between what the user sees and what the system remembers is important, this metric is a leading indicator of incidents.

Replay re-execution rate. What fraction of agent actions had to be re-executed during a replay to reconstruct state. A high rate indicates the replay path is not well defined or there is state that is not replayable. This metric is built by observing incidents or disaster recovery drills, not normal operation.

Audit coverage. What fraction of actions executed by the agent have a complete receipt. Coverage below 100% indicates there are mutation routes that do not go through the harness. It is a security and compliance hole, not just an operational metric.

Why this talk matters even if you don't use OpenAI

OpenClaw is the case study, but the lesson is not about OpenClaw. It is about the fact that any sufficiently complex agentic system — any product that uses LLMs to make decisions affecting persistent state — will face the same problems of ownership, ordering and proof. The talk does not sell an OpenAI product; it sells a way of thinking about the problem. Govindarajan explicitly says this is not an OpenClaw pitch or any product pitch, and that he uses OpenClaw as a public case study because its harness and open nature made the harness visible. The rest of the argument applies equally to Anthropic, Google, Cohere, or an internal system your team is building with LangChain, LlamaIndex or a custom framework.

The three takeaways — own the state, order the mutation, prove the action — predate agents. They are classic distributed systems principles. What the talk contributes is the specific application to the agentic context: how those principles manifest when the decision-maker is not deterministic, is not repeatable, and does not always reason over the same context. That is the piece most teams deploying agents for the first time have not internalised. And it is the piece that separates a demo from a production-ready system.

The common mistake this talk dismantles

The most common mistake when deploying agents in production is treating them as if they were traditional APIs. The implicit reasoning is: the model is like a function, I pass it inputs and it returns outputs, and the rest of the system is like any backend. The reality is different. The model is not deterministic; it can give different answers to the same input. The model is not idempotent; the same call twice can have different effects if the context changed. The model is not auditable the same way an API is; the model's transcript is interpretive, not declarative.

The harness is what turns the model into something operable. It is what takes probabilistic behaviour and exposes it as an API with clear semantics: what it does, what can fail, what guarantees it offers, what is needed to call it correctly. Without the harness, you have a demo. With the harness, you have a system.

Fact verification and sources

- Original presentation by Vinoth Govindarajan at InfoQ: https://www.infoq.com/presentations/ai-agent-harness/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global - Full transcript of the presentation available at the above link - Govindarajan's previous work on distributed systems (Uber, Apple) per his biography in the presentation

The three operational takeaways ("own the state," "order the mutation," "prove the action") are textual formulations from the speaker during the talk. The names of technical patterns (single commit path, stable identifiers, verifiable receipts, reorder buffer, policies as code) are editorial interpretations based on the principles described, not textual claims by the speaker. The suggested metrics are operational derivatives, not explicitly listed in the talk.

Closing

If your team is deploying agents in production — or about to — this talk is the required reading or viewing of the week. Not because it has concrete answers to all your problems, but because it has the right mental framework for formulating the questions. The difference between a team that operates agents successfully and one that struggles with unexplained incidents is, most of the time, whether that team understood that the model is the engine and the harness is the car. If your team is still treating the model as if it were the entire system, this is the moment to correct the mental model. The next time an agent does something it shouldn't, the difference between a manageable incident and an operational disaster will be whether you had the right harness or not.