Flaky test hunter
Every night, Jimothy runs your tests five times in a fresh worktree. When the results disagree, a condition step hands the failure logs to Claude Code, which finds why the test is flaky (shared state, timing, ordering, real race conditions) and fixes the cause instead of adding retries. The suite has to pass five times in a row before you’re asked to approve the PR.
- Run tests repeatedly Shell
- Flaky? Condition
- Fix the flake Claude Code · Sonnet
- Verify fix Shell
- Approve fix You
- Open pull request Shell
What each step does
Run tests repeatedly
ShellRuns the suite several times and keeps the logs of the runs that failed.
Flaky?
ConditionContinues only when some runs passed and some failed. All passing means nothing to fix; all failing is a real break, not a flake.
Fix the flake
Claude Code · SonnetFinds why the test is flaky and fixes the cause, without retries or skips.
Verify fix
ShellRuns the suite again, the same number of times. Every run has to pass.
Approve fix
YouShows you the cause and the fix before anything is pushed.
Open pull request
ShellPushes the branch and opens a pull request.
From copy to first run
- Copy the pipeline JSON with the button above.
- In Jimothy, open Pipelines → New pipeline, paste it under “Or import a pipeline JSON” and click Import.
- Under Workspace, set Local repository to your clone.
- Fill in the variables below under Pipeline settings → Variables.
- Triggers → Add trigger → Schedule, cron 30 2 * * * (02:30 every night).
What to fill in
The prompts and commands read these as {{vars.<name>}}. The examples are placeholders: replace them with your own.
| Variable | What to put in it | Example |
|---|---|---|
testCommand | The command that runs your test suite. | npm test |
runs | How many times to run it. More runs catch rarer flakes but take longer. | 5 |
Show the pipeline JSON
{
"jimothyPipeline": 1,
"name": "Flaky test hunter",
"description": "Runs the test suite several times; when results disagree, fixes the flaky test and opens a PR.",
"icon": "bug",
"color": "#ea580c",
"repo": {
"mode": "worktree",
"localPath": "",
"baseBranch": "main",
"branchTemplate": "jimothy/flaky-{{run.number}}"
},
"concurrency": 1,
"variables": {
"testCommand": "npm test",
"runs": "5"
},
"steps": [
{
"id": "hunt",
"name": "Run tests repeatedly",
"type": "shell",
"description": "Runs the suite several times and keeps the logs of the runs that failed.",
"timeoutMinutes": 90,
"command": "pass=0; fail=0; for i in $(seq 1 {{vars.runs}}); do if sh -c {{vars.testCommand}} > \".jimothy-run-$i.log\" 2>&1; then pass=$((pass+1)); else fail=$((fail+1)); echo \"=== Run $i failed\"; tail -120 \".jimothy-run-$i.log\"; fi; rm -f \".jimothy-run-$i.log\"; done; echo \"PASSED $pass FAILED $fail\""
},
{
"id": "flaky",
"name": "Flaky?",
"type": "condition",
"description": "Continues only when some runs passed and some failed. All passing means nothing to fix; all failing is a real break, not a flake.",
"dependsOn": [
"hunt"
],
"condition": {
"match": "all",
"rules": [
{
"value": "{{steps.hunt.output}}",
"op": "matches",
"compare": "PASSED [1-9]\\d* FAILED [1-9]"
}
]
}
},
{
"id": "fix",
"name": "Fix the flake",
"type": "agent",
"description": "Finds why the test is flaky and fixes the cause, without retries or skips.",
"harnessId": "claude-code",
"model": "sonnet",
"dependsOn": [
"flaky"
],
"timeoutMinutes": 45,
"prompt": "The test suite ({{vars.testCommand}}) passed on some runs and failed on others. Logs of the failing runs:\n\n{{steps.hunt.output | truncate:15000}}\n\nFind which tests are flaky and why: shared state between tests, test order, time and time zones, randomness, network, or a real race condition in the code. Fix the cause. Do not add retries, skip or quarantine tests, or loosen assertions. If the flakiness exposes a real bug in the code, fix the code.\n\nCommit with a message explaining the cause. Then summarise the cause and fix in a few sentences.{{#if steps.verify.output}}\n\nThe suite is still not passing reliably after your last fix:\n{{steps.verify.output | truncate:6000}}{{/if}}"
},
{
"id": "verify",
"name": "Verify fix",
"type": "shell",
"description": "Runs the suite again, the same number of times. Every run has to pass.",
"dependsOn": [
"fix"
],
"timeoutMinutes": 90,
"passPattern": "FAILED 0$",
"loopBackTo": "fix",
"maxLoops": 2,
"command": "pass=0; fail=0; for i in $(seq 1 {{vars.runs}}); do if sh -c {{vars.testCommand}} > \".jimothy-run-$i.log\" 2>&1; then pass=$((pass+1)); else fail=$((fail+1)); echo \"=== Run $i failed\"; tail -120 \".jimothy-run-$i.log\"; fi; rm -f \".jimothy-run-$i.log\"; done; echo \"PASSED $pass FAILED $fail\""
},
{
"id": "approve",
"name": "Approve fix",
"type": "approval",
"description": "Shows you the cause and the fix before anything is pushed.",
"dependsOn": [
"verify"
],
"approvalMessage": "Open a PR with this flaky-test fix?\n\n{{steps.fix.output}}\n\n{{steps.verify.output | trim}}"
},
{
"id": "pr",
"name": "Open pull request",
"type": "shell",
"description": "Pushes the branch and opens a pull request.",
"dependsOn": [
"approve"
],
"timeoutMinutes": 5,
"command": "git push -u origin HEAD && gh pr create --base {{run.baseBranch}} --title \"Fix flaky tests\" --body {{steps.fix.output}}"
}
]
}More for engineering, and beyond
Incident postmortem draft
Turns an incident timeline into a blameless postmortem, traced to the commits that caused it, and opens it as a pull request.
EngineeringRelease notes from a tag
When you push a release tag, writes release notes from the commits since the last one and publishes them to the GitHub release.
EngineeringSecurity advisory triage
Runs your dependency audit weekly, works out which advisories actually affect your code, and opens a PR fixing the ones that do.