Skip to content

Loops & verdicts

Two step options work together to build quality gates:

  • Verdicts (Pass if output matches / Fail if output matches) turn a step’s output into a pass or fail, even when the process itself exited successfully.
  • Feedback loops (On failure, loop back to + Max loops) re-run an upstream step, and everything after it, when a step fails, with the failing step’s output available as feedback.

Both are in the step’s Quality gate & feedback loop section.

UI label JSON Effect
Pass if output matches (regex) passPattern If set and the output does not match, the step fails: Verdict: output did not match pass pattern /…/
Fail if output matches (regex) failPattern If the output matches, the step fails: Verdict: output matched fail pattern /…/

Rules:

  • Patterns are JavaScript regular expressions, evaluated case-insensitively and in multiline mode (^ and $ match at line boundaries). They’re matched anywhere in the output.
  • They’re checked only when the attempt otherwise succeeded (exit code 0 / successful API call). A crashed or timed-out step fails regardless.
  • failPattern is checked first. If both are set, the step passes only if the fail pattern doesn’t match and the pass pattern does.
  • They’re matched against the step’s recorded output: the agent’s final answer, or a shell step’s combined stdout/stderr (the last 20,000 characters).
  • They apply to agent and shell steps, not approval steps.
  • A verdict failure is not retried. Retries are for transient failures; the same output would just fail again. It goes straight to the loop-back (if any) or fails the step.
  • Invalid regexes are caught by validation: Step "Name": invalid regex ….

In JSON, backslashes must be escaped: "passPattern": "VERDICT:\\s*APPROVE".

The built-in review prompt asks the model to end with exactly one line, VERDICT: APPROVE or VERDICT: CHANGES_REQUESTED, and the step uses:

passPattern: VERDICT:\s*APPROVE

Prefer a pass pattern over a fail pattern for gates: if the model forgets the verdict line or the output is empty, a pass pattern fails safe, while a fail pattern would let it through.

Other useful patterns:

Goal Pattern
Reviewer approves VERDICT:\s*APPROVE (pass)
Reviewer requests changes CHANGES_REQUESTED (fail)
Only the last verdict counts VERDICT:\s*APPROVE(?![\s\S]*VERDICT) (pass)
Test runner reported no failures \b0 failed\b (pass) or \b[1-9]\d* failed\b (fail)
Agent flagged a blocker ^BLOCKED: (fail)
Classifier says “yes” ^\s*yes\s*$ (pass)
UI label JSON Default
On failure, loop back to loopBackTo — fail the step —
Max loops maxLoops 2 (UI range 1–10)

When a step with a loop-back target fails for any reason (verdict, non-zero exit, timeout, rejected approval) and its loop budget isn’t used up:

  1. The step’s loop counter increases (run.loops[<step id>]).
  2. The target step and every step downstream of it are reset to Pending, and so is the failing step itself. Each reset step’s iteration counter increases by one.
  3. The run log shows ↺ <Step name> requested another pass: looping back to "<target>" (1/2).
  4. The scheduler runs the reset steps again in dependency order.

The failing step doesn’t fire a step.failed notification when it loops.

When the budget is used up, the step fails with its error plus (feedback loop limit of N reached), and the run fails as usual (unless the step has Continue on error). With maxLoops: 2, the target runs at most three times: the first pass plus two loops.

In the editor, steps with a loop-back show a ↺ icon in the step list, and the graph draws the loop as an arc underneath.

Step outputs persist across iterations. When implement re-runs because review failed, {{steps.review.output}} still holds the review that sent it back. On the first pass it’s empty because review hasn’t run yet. That makes the standard pattern work:

Implementation plan:
{{steps.plan.output}}
{{#if steps.review.output}}
A reviewer requested changes on the previous attempt. Address every point:
{{steps.review.output}}
{{/if}}

Other values useful inside a loop:

  • {{loop.iteration}}: how many times this step has been re-entered by a loop (0 on the first pass).
  • {{steps.test.output | truncate:4000}}: failing test output, truncated so it doesn’t swamp the prompt.
  • {{steps.approve.output}}: a rejecting approver’s comment, when an approval step loops back. (Use .output, not .comment: the loop clears the approval record of the reset approval step, including comment and approvedBy, but its output keeps the comment.)

Each step with a loop-back has its own counter. In the Feature template, both test and review loop back to implement with maxLoops: 2, so implement can run up to five times in one run (1 + 2 from tests + 2 from review). Counters reset when you Retry a run.

{
"id": "review",
"name": "Cross-model review",
"type": "agent",
"harnessId": "codex",
"dependsOn": ["test"],
"prompt": "Review `git diff {{run.baseBranch}}...HEAD` for {{issue.key}}. Do not modify files.\nEnd your answer with exactly one line:\nVERDICT: APPROVE\nor\nVERDICT: CHANGES_REQUESTED\nfollowed by a numbered list of required changes if any.",
"passPattern": "VERDICT:\\s*APPROVE",
"loopBackTo": "implement",
"maxLoops": 2,
"timeoutMinutes": 20
}

A shell step fails on a non-zero exit code, so a test step needs no pattern:

{ "id": "test", "name": "Run tests", "type": "shell", "dependsOn": ["fix"],
"command": "{{vars.testCommand | raw}}", "loopBackTo": "fix", "maxLoops": 2 }

The fixer prompt includes the failure when there is one:

{{#if steps.test.output}}
The previous fix failed the test suite:
{{steps.test.output | truncate:4000}}
{{/if}}

Give an approval step a loop-back, and a rejection re-runs the implementer with the approver’s comment:

{ "id": "approve", "type": "approval", "name": "Human approval", "dependsOn": ["review"],
"approvalMessage": "Approve {{issue.key}}, or reject with what to change.",
"loopBackTo": "implement", "maxLoops": 3 }
{{#if steps.approve.output}}
The human reviewer rejected the previous attempt with this feedback:
{{steps.approve.output}}
{{/if}}
  • Retries run before the loop. If a step has both Retries on failure and a loop-back, a non-verdict failure (for example a test command exiting non-zero) is first retried unchanged, and only loops back after the retries are exhausted. For test steps that loop back, leave retries at 0.
  • Stale feedback. Outputs are only replaced when a step runs again. If implement is sent back by test after an earlier pass was sent back by review, {{steps.review.output}} still contains that earlier review. Word prompts so this is harmless (“if a reviewer requested changes…”), or check {{loop.iteration}}.
  • Parallel siblings. Steps that are still running or waiting for approval when a loop fires are not reset. They finish with results from the previous iteration. Put loop owners after a join (as the built-in templates do) so nothing downstream of the target is still running when they fail.
  • Loop target placement. The target should be upstream of the looping step. Everything downstream of the target is reset, including other branches, so a loop to an early step re-runs a lot.
  • Continue on error vs. loops. A step with a loop-back still loops before Continue on error matters; Continue on error only applies once it finally fails.