A repair loop should make new progress, not merely repeat a request. If the third attempt has the same input, environment, and failure as the first, the system may be consuming budget without learning anything useful.
Classify the failure first
Compilation errors, failed assertions, missing tools, provider outages, and denied permissions require different responses. A model can often fix a code defect. It cannot make an unavailable Docker daemon appear or grant itself a missing organization role. Retrying every failure through the same prompt obscures the real problem.
Separate transient execution failures from permanent configuration failures and failed verification. Preserve an attempt history so a later operator can see whether the system tried a new hypothesis or repeated the same action.
Give repair work a bounded package
A useful repair package contains the failing gate, relevant output, expected behavior, and commit context. It should not indiscriminately include every log and secret-bearing environment variable. Give the repair task a clear scope and retain the original acceptance criteria so a “fix” cannot redefine success.
For example, if a retry test fails because the button remains disabled, the repair should address that behavior and rerun the relevant checks. Deleting the test or weakening its assertion may make a gate green while violating the requirement. Independent review should consider the test changes as part of the patch.
Limits are part of the design
Bound the number of provider attempts and repair cycles. Apply time and cost limits as well as a count. When those limits are reached, keep the failed state and evidence visible, then escalate. A clear blocked result is preferable to an endless “working” indicator.
Recovery is not the same as repair
A runner disappearing creates an ownership problem: another process may need to resume or reclaim work. A failed acceptance test creates a correctness problem: the implementation needs a change. Mixing the two makes duplicate execution and misleading histories more likely. Use task leases and durable state for ownership recovery, and bounded repair work for code correction.
ForgeLoop exposes both concepts in its delivery model. The goal is not to avoid every failure; it is to leave the system and its operator with an honest, actionable account of what happened. Explore verification features for the surrounding workflow.