skip to content
$empowered.guru

AI & Machine Learning

Harness, Loop, Graph: The Three Layers Behind Reliable AI Agents

Most AI agents do not fail because the model is too weak. They fail because the surrounding system lacks a durable environment, evidence-based feedback cycle, or explicit control flow.

August 10, 202612 min read
S

Staff Writer

Published August 10, 2026 · Updated September 30, 2026last updated dates

Harness, Loop, Graph: The Three Layers Behind Reliable AI Agents

The real reason AI agents fail

AI agents often look impressive in a demo: give the model a goal, expose a few tools, and watch it complete a task. Production is less forgiving. Tools time out. State disappears between sessions. A nearly-correct result gets accepted without verification. A workflow branches in ways no one can inspect, and a retry simply repeats the same mistake.

When that happens, teams often upgrade the model first. That is frequently the wrong diagnosis. A serious agent is not just a model; it is a system built around the model. That system has at least three distinct engineering layers: the harness, the loop, and the graph.

“The harness gives the model a place to work. The loop gives the work a feedback cycle. The graph gives the process an explicit route.” - Adapted from the source thread by rari 1

These layers overlap, and one codebase may contain all three. They are nevertheless different levers for reliability:

Layer Core question Primary responsibility Typical failure
Harness Where can the agent work, and what can it safely access? Tools, context, state, permissions, persistence, checkpoints, traces, and approvals The agent lacks a capability, loses progress, or operates with unsafe or ambiguous access
Loop How does the agent know whether its work is good enough? Evidence, feedback, bounded retries, and stopping rules The agent repeats itself, stops on confidence, or continues after success
Graph What is allowed to run next? Nodes, branches, joins, parallel work, handoffs, recovery paths, and legal cycles The process is hidden inside ad hoc orchestration and cannot be inspected or resumed

The shortest version is worth memorizing: harness equals environment, loop equals feedback, and graph equals flow. The distinction becomes important as soon as an agent touches real files, APIs, customers, money, or production code.

1. Harness engineering: build the environment around the model

A raw model can transform an input into an output. It cannot independently maintain project state, run a test suite, inspect a browser, write files safely, enforce permissions, or resume tomorrow where it stopped today. The harness supplies that machinery.

A useful mental model is this: the model provides intelligence, while the harness turns that intelligence into an operational capability. If you remove the model from the architecture diagram, everything that remains is probably part of the harness.

A production harness usually has six responsibilities:

Responsibility What it includes Why it matters
Context System instructions, retrieved knowledge, conversation state, task policies, and operating procedures The agent needs the right information without replaying an entire history
Action surfaces APIs, browser control, code execution, databases, MCP tools, and specialist agents The agent must be able to act through narrow, well-defined interfaces
Persistence Files, checkpoints, session state, progress logs, git history, and long-term memory Long-running work must survive interruptions and context-window boundaries
Execution control Timeouts, retry limits, cost budgets, model routing, handoffs, and approval gates The system needs predictable limits rather than unlimited autonomy
Safety Isolated environments, least-privilege permissions, allow lists, secret handling, and human authorization The agent should have only the access required to complete its job
Observability Tool inputs and outputs, traces, state transitions, cost, latency, and evaluation results Operators need to reconstruct what happened and why

Harness engineering earns its keep when a task lasts longer than one conversation. A coding agent that works for hours cannot rely on chat history alone. It needs durable artifacts that another session-or another person-can understand.

A practical long-running setup might include an initializer that inspects the workspace, a progress file that records what is complete and what remains, commits that preserve working states, checkpoints before risky actions, and verification tools that produce clear evidence. That is not a better prompt. It is a better working environment.

Start with the harness when the agent cannot access the right capability, loses progress between sessions, has permissions that are too broad, behaves differently across environments, cannot be paused or resumed, or produces failures nobody can reconstruct.

2. Loop engineering: make the work iterative and verifiable

Every tool-using agent already has a small internal cycle:

model → action → observation → model

Loop engineering begins when you deliberately design the cycles around that behavior. The goal is not to make the agent repeat itself forever. The goal is to turn a one-shot attempt into a managed process with evidence, useful feedback, and a bounded exit.

The most valuable outer loop is a verification loop:

BUILD
  ↓
CHECK AGAINST EVIDENCE
  ↓
PASS? ── yes ──> STOP
  │
  no
  ↓
RETURN SPECIFIC FEEDBACK
  ↓
RETRY WITH A LIMIT

The check may be deterministic: tests pass, a schema validates, links resolve, numbers reconcile, or files compile. It may also require a reviewer who evaluates whether the argument is complete, the tone fits the audience, the evidence supports the conclusion, or the change is correctly scoped.

The rule is the same in both cases: do not loop on confidence; loop on evidence. “The agent says it is finished” is not proof. Passing tests, resolving sources, reconciling numbers, or receiving approval on a diff is proof.

A useful production loop has seven parts:

Part Design question
Trigger What starts another cycle: a request, schedule, webhook, failed test, new document, or evaluator result?
Goal What measurable state must be reached?
State What must the next attempt know without replaying the entire history?
Action policy What may the agent change, call, delegate, or spend?
Evidence Which tests, citations, diffs, metrics, schemas, or human reviews prove progress?
Feedback What compact explanation tells the next attempt what failed and what must change?
Stopping rule When does the system stop because it succeeded, timed out, exhausted its budget, hit a hard error, or needs a human?

Loops can stack. An agent loop performs the work. A verification loop checks the work. An event loop wakes the system when new work arrives. A trace-improvement loop studies production behavior and changes the harness itself.

This is why loop engineering is larger than prompt engineering. A prompt defines what should happen during one model call. A loop defines what the system does after that call.

Loops have a cost: every retry, grader, and reviewer adds latency and spend. Add a loop when the expected cost of failure is higher than the cost of verification. A high-impact publishing workflow may justify several checks; a low-risk formatting task may not.

3. Graph engineering: make control flow explicit

Graph engineering asks a different question: not “How should the agent work?” but “What is allowed to run next?” Work becomes nodes, allowed transitions become edges, and state moves through the graph.

A graph can represent fixed sequences, conditional branches, parallel fan-out, joins, bounded cycles, recovery paths, and human interrupts. The important decisions are architectural rather than cosmetic:

Graph decision What must be made explicit
Node boundaries Which work belongs in ordinary code, an LLM call, a specialist agent, or a human review step?
State schema What can each node read or update, and how are parallel results merged?
Routing conditions Which evidence moves work forward, backward, sideways, or into escalation?
Concurrency What can run in parallel, and what must wait for a join?
Cycles and exits Where are retries legal, how many attempts are allowed, and what makes the cycle safe?
Durability Where is execution checkpointed, and how does it resume after interruption?

A graph is worth the ceremony when the process contains meaningful branches, parallel specialists, approvals, recovery routes, or stateful handoffs. It is not worth adding merely because a workflow has several steps. If one capable agent with three tools can solve the task, a graph may add structure without adding value.

The opposite mistake is also common: teams formalize the workflow before they understand the work. The result is a beautiful diagram that encodes the wrong assumptions. Start with a simple harness, study real traces, and formalize the paths that remain stable.

How the three layers nest in a real system

Consider a research-and-publishing agent that produces a factual industry briefing. The harness provides browser and search tools, source storage, a writing workspace, citation checking, permissions, approval rules, checkpoints, and traces.

The graph controls the route:

RESEARCH
   ↓
DRAFT
   ↓
FACT CHECK ── fail ──> RESEARCH
   │
  pass
   ↓
EDITORIAL REVIEW ── fail ──> DRAFT
   │
  pass
   ↓
HUMAN APPROVAL
   ↓
PUBLISH

The loops live inside that route. The research node may search until source coverage is sufficient. The drafting node may revise until a style grader passes. The fact-check node may return the exact unsupported claims instead of a vague rejection.

This nesting is the key idea. The graph runs inside the harness. The loops run inside parts of the graph. The harness supplies the tools, state, and evidence those loops need. The layers overlap because real software layers overlap, but they still give you three different levers when the system fails.

Diagnose the failure before changing the architecture

When an agent fails, first identify which layer owns the failure. Changing the model or adding orchestration before doing that usually increases complexity without addressing the cause.

Symptom Start with Likely fix
The agent cannot access the right data safely Harness Improve the tool contract, permissions, sandbox, or context injection
The agent forgets progress between sessions Harness Add durable state, checkpoints, progress artifacts, and compaction
The first attempt is close but unreliable Loop Add deterministic tests or an external grader with actionable feedback and bounded retry
The agent continues after success or stops before proof Loop Define evidence-based terminal states and budget-aware stopping rules
Specialists must run in a controlled order Graph Add explicit nodes, edges, routing conditions, and joins
A multi-step failure is impossible to locate Graph + harness Align stateful traces with nodes and transitions
The process changes too quickly for a fixed diagram Simpler harness Keep planning model-driven and delay graph formalization

This diagnostic approach is more useful than arguing about terminology. Find the layer that owns the failure, then fix that layer first.

The expensive mistakes behind weak agent systems

Building the graph too early. Do not convert an imagined business process into dozens of nodes before watching a strong agent perform the work. Trace first; formalize second.

Letting the maker grade itself. Self-review is useful, but it shares many blind spots with the original attempt. Prefer deterministic checks where possible, use an isolated reviewer context for subjective checks, and require human approval for high-impact actions.

Defining the loop as “keep trying.” An unbounded retry is not reliability; it is a cost leak. Every cycle needs fresh evidence, a maximum attempt count, and a named escalation path.

Turning the harness into a warehouse. More tools do not automatically create a better agent. A crowded toolset increases selection errors, noisy context increases confusion, and broad permissions increase risk. Give the agent the smallest environment that can complete the job.

Blaming the model for orchestration failures. A stronger model cannot reliably repair stale state, broken APIs, ambiguous tool schemas, or missing exit conditions. Prove that the model is the problem before upgrading it.

A production-ready checklist

Before shipping an agent, review each layer independently and then review how the layers interact.

Layer Questions to answer
Harness Are the tools narrow, documented, and observable? Is state durable? Are permissions least-privilege? Can an operator pause, inspect, and resume the run? Can important actions be reconstructed from traces?
Loop What evidence proves success? What feedback follows failure? How many retries are allowed? What happens when the budget is exhausted? Where is human judgment required?
Graph Which paths must be deterministic? Where can work run in parallel? What state is shared? Where are the joins, approvals, and recovery routes? Which cycles are legal, and how do they terminate?
Evaluation Can the team replay real traces? Can versions be compared on the same tasks? Can an improvement be attributed to a specific change?
Operations Are cost and latency monitored? Is failure rate visible by node and tool? Is human intervention measured? Is task-level success measured in production?

The framework in one sentence

Harness engineering makes the model operational. Loop engineering makes the work iterative and verifiable. Graph engineering makes complex execution explicit and controllable.

None replaces the others. A perfect graph cannot save an agent that loses its state. A perfect harness still wastes money if the loop has no evidence or stopping rule. A strong loop becomes difficult to operate when branches, parallelism, and approvals are hidden inside ad hoc code.

Reliable agents appear when all three layers are designed together - and when every layer has one clear job:

ENVIRONMENT → HARNESS
FEEDBACK    → LOOP
FLOW        → GRAPH

That is the practical shift from prompt engineering to agent-systems engineering.


Source note: This article is an original adaptation of rari’s thread, “LOOP vs GRAPH vs HARNESS ENGINEERING”, published July 28, 2026. The structure and wording have been edited for a blog audience; credit remains with the original author for the framework and source ideas.

References

Filed under

AI AgentsAgent EngineeringAI & Machine LearningSoftware ArchitectureAutomation
$empowered.guru --book-session

Keep exploring

Turn the next insight into a shipped product.

Bring us the product, architecture, or delivery problem you are working through. We will help you find the clearest path forward.