A coding agent and a delivery harness solve related but different problems. The agent reasons about a change and uses tools to make it. The harness decides when that work may run, what context and authority it receives, and how the result is evaluated.
Separate the worker from the rules
An agent can propose a patch, run a command, and explain a result. A delivery system must also handle duplicate events, unavailable machines, competing tasks, exhausted budgets, and stale verification. Asking the agent to remember every rule does not give those rules a durable enforcement point.
Consider a worker that disappears after creating a commit. Another worker needs to know whether the task is still owned, whether execution can resume, and what evidence already exists. That is a state-management problem, not simply a prompt-writing problem. Durable task records and leases make the decision explicit.
Why the distinction matters for teams
When policy lives outside the model conversation, it can be reviewed independently. The same repository rules can apply across different providers. A change in model does not need to redefine who may merge a PR or which tests are mandatory. Likewise, provider success should not be able to overwrite a failed verification result.
The tradeoff is complexity. A harness has configuration, credentials, storage, and operational responsibilities. For a one-off local edit, that overhead may be unnecessary. It becomes more valuable when multiple repositories, repeatable verification, or auditability matter.
What to ask when evaluating a harness
- Can you identify the exact commit that passed verification?
- Is there an explicit boundary between drafting and approving?
- Can work recover from a lost runner without losing ownership history?
- Are budgets and retries enforced outside the model response?
- Can the system admit that a result is blocked or unverified?
ForgeLoop is intended to be that surrounding system, not a replacement for every editor assistant. Its feature overview describes the implementation boundaries, while the repository checklist records remaining release work. Evaluate both; a polished dashboard alone is not delivery evidence.