ixsoftum
Harness engineering vs prompt, context, and loop engineering
AI-assisted developmentComparison

Harness engineering vs prompt, context, and loop engineering

Prompt, context, loop, and harness engineering optimize four different layers of an AI agent. Here's what separates them, and what a real harness is made of.

Ava Harlan·Published 20 Aug 2026·9 min read
Last verified 21 Aug 2026

Harness engineering is one of at least four distinct practices now grouped loosely under "AI engineering," alongside prompt engineering, context engineering, and loop engineering. Each optimizes a different layer of an AI coding agent: the single instruction, the information window, the run-until-done iteration cycle, or the execution environment built around the model. Harness engineering is the last one. Martin Fowler's own framing of the term is precise: "Agent = Model + Harness." If you're running Claude Code against a real CLAUDE.md file, a couple of MCP servers, and a command allowlist, you already have one.

The four terms get used almost interchangeably in casual conversation, which is exactly why they're worth separating properly.

Four disciplines, four different layers

DisciplineWhat it optimizesDefinitionSource
Prompt engineeringA single instruction, for one model call"The practice of creating the single most effective, optimized instruction for an AI model."IBM
Context engineeringWhat the model actually sees during inference"The set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference."Anthropic
Loop engineeringThe reason, act, observe, adjust cycle an agent runs with minimal supervision"Designing agentic workflows... that iteratively guide AI agents toward completing user-defined goals with minimal human intervention."IBM
Harness engineeringThe execution environment, tools, and guardrails around the agent"Agent = Model + Harness."Martin Fowler

These four sources don't fully agree on how the pieces fit together, and it's worth saying so plainly rather than picking one framing and presenting it as settled. Fowler's article, written by Birgitta Bockeler, treats the harness as the umbrella term, everything that isn't the model, with context management and tools living inside it. IBM's own explainer treats harness engineering as one specific practice nested inside loop engineering, alongside context engineering as a separate, parallel component.

Neither source is wrong. They're describing the same real system from two different starting points, one from the agent's total surface area inward, the other from the iteration cycle outward. What doesn't change between the two framings is the actual list of things you build: instructions, tools, checks, and a loop that runs them.

A fifth, related term worth flagging so it doesn't get confused with these four: spec-driven development, the practice of writing a detailed specification before an agent touches code, covered in depth in Does agentic engineering kill Scrum? That's a requirements-definition practice, a different axis entirely from any of the four above.

What a harness is actually made of

Per Fowler's article, a harness splits into two halves.

HalfWhen it runsWhat it doesSmall-team examples
GuidesBefore the agent acts (feedforward)Steers behavior, aims to get the output right on the first tryCLAUDE.md/AGENTS.md, coding standards, a bootstrap script
SensorsAfter the agent acts (feedback)Catches problems, lets the agent self-correctTests, linters, a second agent doing code review

Both halves come in two flavors: computational (deterministic and fast, a type checker or a test suite) and inferential (an LLM doing semantic review, slower and non-deterministic but able to catch what a linter can't). A small team's harness usually leans computational for good reason. It's cheaper to run, and it doesn't drift the way a semantic review pass can.

See the CLAUDE.md file vs .cursor/rules vs AGENTS.md for what actually belongs in the guide half and what to gitignore instead. See MCP servers explained for the tool-access half, what actually lets an agent open a pull request or query a database instead of just describing what it would do.

Why harness engineering showed up now

Coding agents used to run one turn at a time: a person typed a prompt, the model answered, the person judged the answer. Loop engineering changed that shape. An agent now reasons, acts, observes the result, and adjusts, repeatedly, often for hours, with a human checking in only at the start and the end. IBM's own breakdown of that cycle lists the components a well-built loop needs to stay "efficient, reliable, and bounded": scheduling, hooks, context engineering, tool access, worktrees, skills, subagents, and a persistent state IBM calls the spine.

Once an agent is running unsupervised for that long, the prompt stops being the thing that determines whether it can be trusted. The environment around it does: what tools it can reach, what it's told not to do, what catches a mistake before a human sees it. That environment is the harness.

The term names a real shift in where the engineering effort actually goes: from the prompt to the environment around it. That's likely why both OpenAI and IBM published their own explainers on it within the same year, arriving at the concept from different products and different scales independently.

How OpenAI built its own harness

OpenAI's own Codex team, running a five-month internal build with no manually written code, hit a version of this at real scale: three engineers growing to seven, roughly 1,500 pull requests merged. Their first instinct was one large AGENTS.md file covering everything.

It failed for reasons that generalize past their team size. A giant instruction file crowds out the actual task, and once everything in it is flagged as important, the agent stops navigating intentionally and starts pattern-matching locally. Their fix was to shrink AGENTS.md to roughly 100 lines and treat it as a map, pointing into a structured docs/ directory that holds the actual rules. The principle, that the guide file's job is to point somewhere rather than contain everything, holds at two engineers just as much as it did for their team of seven.

Their other transferable lesson is about where to spend enforcement effort: boundaries centrally, autonomy locally. A command allowlist should stop an agent from running a destructive command or pushing straight to main. Stylistic choices inside those boundaries stay the agent's call.

That's the guardrail scope a small team actually needs from a harness, lighter than OpenAI's full architecture-fitness function suite. It's also the same shift from code writer to orchestrator already describes at a smaller scale: the human's job moves from writing the line to deciding where the line needs a check.

The real advantages, and the real costs

The advantage that matters is throughput holding up as the team grows, a harder bar than a fast start on day one. OpenAI's team merged roughly 3.5 pull requests per engineer per day, and that rate increased as headcount went from three to seven, the opposite of what usually happens when more people touch the same codebase. That's the payoff of a harness doing its job: rules enforced mechanically apply to every agent run at once, instead of degrading as more people (or more agents) touch the code.

The costs are real too, and IBM names three worth taking seriously before scaling a harness past a small team.

  • Comprehension debt: the gap between how much code exists and how much of it any human actually understands. It grows quietly, because agent-generated code usually passes its tests even when nobody could explain why it works.
  • Intent debt: what's lost when the reasoning behind a decision lives only in a chat transcript instead of a committed file. Agents can't infer intent that was never written down.
  • Cognitive surrender: the point past comprehension debt where a team stops reviewing agent output at all and starts accepting it by default.

None of these are reasons to avoid building a harness. They're reasons to keep a human genuinely engaged with the loop the harness is automating, reviewing real output on a real cadence.

Where harness engineering actually gets used

The two real examples above sit at opposite ends of the same practice. A 2-8 person team's harness is CLAUDE.md/AGENTS.md plus a couple of MCP servers plus a command allowlist, three files, built in an afternoon, sufficient for one or two engineers driving Claude Code on a real codebase.

OpenAI's harness is the same three categories (guides, tools, guardrails) at a scale that needed a structured docs/ directory, custom architecture linters, and a recurring cleanup agent to keep seven engineers and their agents from drifting apart. The components don't change between the two. Only how much machinery each component needs does.

Is this actually a new discipline, or a new name for old practice?

For a team already running Claude Code with a real CLAUDE.md file, a couple of MCP servers, and a required review pass, harness engineering just names a setup that already exists. A Reddit thread debating the term landed on close to the same read from several different commenters, more than one called it a rebrand of something practitioners were already doing well before the phrase showed up.

Where the term earns its weight past semantics is enforcement. OpenAI's team didn't just document its architecture rules, it encoded them as custom linters that block a pull request mechanically. At agent-driven throughput, a rule that only lives in a doc doesn't hold. Agents replicate whatever pattern already exists in the repo, good or bad, so an unenforced convention gets copied right along with everything else.

A 2-8 person team doesn't need OpenAI's machinery yet. But the day a CLAUDE.md file stops being enough and a real rule needs to become a CI check instead of a suggestion, that's the actual moment harness engineering stops being a label and starts being a decision worth making on purpose.

Frequently asked questions

What's the difference between prompt engineering and harness engineering? Prompt engineering optimizes a single instruction for one model call. Harness engineering builds the tools, guardrails, and checks that surround an agent across many calls, so the agent's output can be trusted without a human reviewing every step.

Is harness engineering the same as context engineering? No, though they overlap. Anthropic's own definition scopes context engineering specifically to curating what tokens the model sees during inference. Harness engineering is broader: it includes context management but also covers tool access, guardrails, and the checks that run after the agent acts.

Do I need a formal harness for a small side project? Usually not a formal one. A CLAUDE.md or AGENTS.md file, one or two MCP servers, and a command allowlist already function as a minimal harness for a 2-8 person team. The heavier machinery (custom architecture linters, a dedicated cleanup agent) only earns its cost once a team and its agent-generated pull request volume both grow past what one person can review by hand.

How is loop engineering different from harness engineering? Loop engineering, per IBM's definition, designs the iterative reason-act-observe-adjust cycle an agent runs to reach a goal. Harness engineering is one of the components that makes a loop work reliably, specifically the tools and guardrails layer, alongside other components like context engineering and persistent state.

Related