CI Sweeper Workflow Design¶
Workflow and domain design for the ci-sweeper loop.
| Layer | Document |
|---|---|
| Platform | Multi-Branch Loops Design |
| Caller shell | Loop Caller Workflows Design |
| Invariants | Loop Engineering Design |
Artifacts: on-loop-ci-sweeper.yaml · skill ci-sweeper · ci-sweeper/scripts/detect_ci_failures.sh · ci-sweeper/scripts/update_run_ledger.sh
Shared caller keys: Loop Caller Inputs Reference.
Purpose¶
Automated minimal repair when CI fails on integration branches and/or open PR heads. One engine; no separate loop-pr-ci-healer package.
Supported use cases¶
- Integration branch CI failure (
main,develop,release/*, …) → open a bot fix PR to the watch branch; L3 enables GitHub auto-merge on that fix PR - Open PR head CI failure → open a bot fix PR to the PR head branch (not
main); post a marker comment on the human PR with fix summary and link to the bot fix PR; L3 auto-merges the bot fix PR into the head branch - Classify failures as Fix / Watch / Escalate; apply small lint, workflow, shell, or doc edits when actionable
- Dedupe by
workflow_run_idledger; skip infra/flake/env failures asoutcome: watch workflow_runon repair-target CI workflows (dogfood default);workflow_dispatchfor manual /gh run listscan; see ops checklist
Dogfood uses open_pr for all repair paths. Direct branch push (push / push_head) is a platform exception path — not the ci-sweeper default. See Finalize strategy.
Out of scope¶
Entry skill design intent for failure kinds deferred via Failure kind defer (B) — coverage threshold, dependency breakage.
- Infra outages, secrets, runner capacity, or persistent flakes (Watch — no code edit when Skill recognizes them)
- Large refactors (>5 files), auth/payment/credential paths
- Auto-merging the human's open PR (only the bot fix PR is auto-merged at L3)
- Re-running CI in the checker (semantic fit against log excerpt only)
- Manual interactive debugging as a substitute for the loop
- Separate
loop-pr-ci-healerpackage - Coverage-threshold and test-gap repair — defer (B) until a domain skill exists
- Dependency-breakage repair — defer (B); bot PR heads excluded in dogfood (
pr_include_bots: "") - Per-PR opt-in labels (
pr_require) — removed; usepr_excludeonly
Skill execution boundaries: ci-sweeper SKILL.md (USE FOR / DO NOT USE FOR).
Execute — responsibility split (A' / B)¶
Distributable ci-sweeper skill stays repository-neutral. Do not hardcode consumer skill names in skill references/. Named dispatch belongs in caller agent_maker_instructions (dogfood: on-loop-ci-sweeper.yaml).
| Layer | Input | Role |
|---|---|---|
| Detect | detect_ci_failures.sh |
failures[], failure_type hint; optional future stack_hint |
| Entry skill | ci-sweeper |
Generic orchestration: classify, follow ## Instructions for skill dispatch, fix one, validate |
| Caller | agent_maker_instructions |
Stack routing (A') — workflow/stack → named domain skills for this repo |
| Caller | agent_checker_instructions |
Failure kind defer (B): appendix REJECT rules |
| Domain skills | Consumer .agents/skills/* |
Invoked per caller routing table |
See CI failure repair — layered responsibilities.
Failure contexts¶
Two independent watch paths. Both use open_pr finalize; level selects human review (L2) vs GitHub auto-merge on the bot fix PR (L3).
| Context | Trigger example | Bot fix PR target (to.branch) |
|---|---|---|
integration |
CI fails on main after direct push |
main |
pull_request |
CI fails on PR head hotfix/0001 → main |
hotfix/0001 (PR head) |
End-to-end flows¶
Integration (main CI failure)¶
L2:
1. main CI fails
2. Bot opens loop/* → main fix PR
3. Human reviews and merges fix PR
L3:
1. main CI fails
2. Bot opens loop/* → main fix PR
3. GitHub auto-merge on fix PR (allowlist + branch protection)
Pull request (main ← hotfix/0001, head CI failure)¶
L2:
1. Human PR #1 (hotfix/0001 → main) — CI fails on hotfix/0001
2. Bot opens loop/* → hotfix/0001 fix PR (#2)
3. Bot comments on human PR #1: fix summary, link to #2, merge or close guidance
4. Human merges #2 into hotfix/0001 → CI on #1 goes green → human merges #1
L3:
1–3. Same as L2, except bot fix PR (#2) is auto-merged into hotfix/0001
4. Human merges #1 when CI is green
The human PR is never auto-merged by the loop. L3 auto-merge applies only to the bot fix PR.
Caller inputs¶
Keys are passed in on-loop-ci-sweeper.yaml via with: on ci-loop-caller.yaml (alphabetically ordered). Multiline values (agent_checker_instructions, agent_maker_instructions) are defined inline in the caller workflow.
Shared semantics: Loop Caller Inputs Reference. Platform branch/finalize caps: canonical table.
Dogfood minimum (autonomy + delivery + PR watch):
delivery: open_pr
level: L2
may_edit: true
pr_enabled: true
pr_exclude: fork,draft,label:no-loop
write_target: fix
Git landing (open_pr / push / push_head) is derived from delivery inside loop-detect. Advanced overrides use git_landing_* on the action only — not caller inputs. See Level × finalize matrix.
| Input / JSON key | Description | Dogfood value |
|---|---|---|
additional_commit_paths |
Extra paths included in finalize commit (ledger file). | .loop/state-ci-sweeper-run-ledger.json |
agent_maker_max_turns |
Max maker agent turns per loop attempt (one Agent→Verify cycle). | 100 |
agent_maker_model |
Maker model ID. Cursor: agent --list-models. |
claude-sonnet-5 |
agent_maker_effort |
Maker reasoning effort. Only engines with an effort flag (claude) consume it | medium |
agent_loop_max_attempts |
Max Agent→Verify retry cycles before finalize records failure. | 3 |
agent_checker_instructions |
Checker APPROVE/REJECT rubric. Requires fix addresses logged CI failure; minimal diff; allowlist/denylist respected. | Inline in caller workflow |
agent_checker_max_turns |
Max checker agent turns per verification. | 100 |
agent_checker_model |
Checker model ID. Cursor: agent --list-models. |
claude-opus-5 |
agent_checker_effort |
Checker reasoning effort. Only engines with an effort flag (claude) consume it | medium |
allowlist |
Comma-separated globs the maker may modify. | .github/**,.apm/packages/**,scripts/**,apm.yml,mise.toml,renovate/**,docs/**/*.md,README.md,mkdocs.yml |
branch_match |
Comma-separated integration branch patterns to poll for failed CI. | main |
branch_state |
Branch for .loop/* persistence, state migration, and watch fallback. |
main |
budget_max_runs_per_day |
Daily run cap keyed by loop_name. Caller input; .loop/loop-budget.json overrides when present. |
5 (caller); effective 50 via .loop/loop-budget.json |
budget_max_tokens_per_day |
Daily aggregated token cap across loops. | 1000000 |
denylist |
Not set — the platform default (**/.env,**/credentials*,**/secrets*) applies. |
Omitted |
Reusable workflow (uses:) |
ci-loop-caller.yaml — detect job requires actions: write, contents: read, and pull-requests: read; caller permissions must include actions: write. |
./.github/workflows/ci-loop-caller.yaml |
detect_script |
Domain detect script path. Uses gh run list per watch branch / PR head. |
.agents/skills/ci-sweeper/scripts/detect_ci_failures.sh |
domain_persistence_script |
Bash script for loop-finalize domain persistence (run ledger updates). |
.agents/skills/ci-sweeper/scripts/update_run_ledger.sh |
engine |
AI engine (claude, copilot, codex, cursor). Maps AGENT_TOKEN to engine env. |
claude |
level |
Autonomy: L2 (human merges bot fix PR) or L3 (GitHub auto-merge on bot fix PR). |
L2 |
loop_name |
Loop identifier; state file .loop/state-ci-sweeper.json. |
ci-sweeper |
max_targets_per_schedule |
Max targets per cron tick after priority filters. | 3 |
no_changes_verdict |
APPROVE or REJECT when maker produces no file diff on actionable CI failure. |
REJECT |
pr_body |
Optional static prefix (dogfood: ""). loop-finalize composes agent Overview/Summary + mechanical sections. See Loop PR Body Readable Design. |
"" |
pr_enabled |
Watch open PR heads for failed CI. | true |
pr_exclude |
PR exclusion tokens: fork, draft, label:<name>, wip_title. |
fork,draft,label:no-loop |
pr_include_bots |
Comma-separated bot logins to include when scanning PRs. Empty = exclude all bots. | "" |
pr_title |
PR title when finalize strategy is open_pr. |
chore(ci-sweeper): automated CI repair (loop-ci-sweeper) |
agent_maker_instructions |
Domain instructions: classify Watch vs Fix; minimal diff; run validation skills. | Inline in caller workflow |
agent_maker_skill_name |
Skill package to invoke. | ci-sweeper |
Removed from dogfood (do not set): pr_require, finalize_integration, finalize_pull_request (replaced by delivery).
Domain detect environment (detect_domain_env_json)¶
| JSON key | Description | Dogfood value |
|---|---|---|
CI_SWEEPER_LEDGER_FILE |
JSON ledger for workflow_run_id dedupe |
.loop/state-ci-sweeper-run-ledger.json |
CI_SWEEPER_REJECT_MAX_RETRIES |
Max retries per run ID when policy is limited |
"3" |
CI_SWEEPER_REJECT_RETRY_POLICY |
block, retry, or limited |
block |
Event keys (embed when workflow_run trigger is enabled on the caller):
detect_domain_env_json: ${{ format('{{"CI_SWEEPER_EVENT_HEAD_BRANCH":"{0}","CI_SWEEPER_HEAD_BRANCH":"{0}","CI_SWEEPER_HEAD_SHA":"{1}","CI_SWEEPER_WORKFLOW_RUN_ID":"{2}"}}', github.event.workflow_run.head_branch || 'main', github.event.workflow_run.head_sha || github.sha, github.event.workflow_run.id || '') }}
| JSON key | Description |
|---|---|
CI_SWEEPER_EVENT_HEAD_BRANCH |
Stable failed-run head (not rewritten per scan). With WORKFLOW_RUN_ID, scopes loop-detect to that head only. |
CI_SWEEPER_HEAD_BRANCH |
Initial head; loop-detect rewrites per scan context |
CI_SWEEPER_HEAD_SHA |
Failed run head SHA |
CI_SWEEPER_WORKFLOW_RUN_ID |
Failed run ID for ledger dedupe + trigger-aware target scoping |
CI_SWEEPER_WORKFLOW_NAME |
Failed workflow display name |
CI_SWEEPER_RUN_URL |
HTML URL of failed run |
Execute-only inputs¶
| Input | Dogfood value |
|---|---|
additional_commit_paths |
.loop/state-ci-sweeper-run-ledger.json |
domain_persistence_script |
.agents/skills/ci-sweeper/scripts/update_run_ledger.sh |
When CI_SWEEPER_WORKFLOW_RUN_ID is set, loop-detect enumerates only the failed head (matching integration branch or open PR). See Trigger-aware priority. Detect script binaries are always pinned from the job checkout (branch_state), never from a stale PR worktree.
Detect¶
Stable mechanical filters (detect script + platform)¶
Detect applies only stable gates — not semantic failure classification:
| Filter | Layer |
|---|---|
Ledger dedupe (workflow_run_id) |
detect script |
workflow_run caller workflows: list |
caller trigger |
workflow_run → single failed head |
loop-detect scope |
Pin DETECT_SCRIPT absolute path |
loop-detect |
| Event run branch vs scan branch | detect script |
| PR exclusion (fork, draft, bot, label) | loop-detect + script |
Workflow concurrency (loop-state-main) |
on-loop-*.yaml |
| Budget / circuit breaker | loop-detect |
failure_type from grep heuristics is an optional hint for the Skill — not a detect gate. Default is regression when the log is not infra/env/flake. See Execute — responsibility split.
Integration mode¶
Per watch branch, loop-detect sets context; script uses gh run list --branch <watch_branch> --status failure.
- Range filter via
targets["integration:<branch>"].last_sha - Dedupe:
state-ci-sweeper-run-ledger.jsonkeyed byworkflow_run_id(entries pruned after 30 days on ledger update)
Pull request mode¶
Requires pr_enabled: true.
- Failed runs where
head_branchmatches an eligible open PR - Emit
pr_number,base.branchintarget_json.to - Apply PR exclusion rules
Detect truth source¶
| Mode | Primary cursor | Secondary |
|---|---|---|
| integration | workflow_run_id + ledger |
last_sha in state (advance on finalize) |
| pull_request | workflow_run_id + ledger |
head_ref per PR in state |
Rule: Do not rely on last_sha alone to skip a still-failing workflow run. Ledger outcome pr-created or watch blocks re-processing same run ID under block policy.
Run ledger and REJECT retry policy (caller / detect)¶
Owned by the caller and detect scripts — not by skill triage. Configure via detect_domain_env_json:
| Variable | Default (portable detect default) | Dogfood | Description |
|---|---|---|---|
CI_SWEEPER_LEDGER_FILE |
ci-sweeper-run-ledger.json |
.loop/state-ci-sweeper-run-ledger.json |
Repo-relative ledger keyed by workflow_run_id |
CI_SWEEPER_REJECT_RETRY_POLICY |
block |
block |
block: skip any ledgered run. retry: skip only pr-created. limited: skip rejected after max retries. Aliases a/b/c |
CI_SWEEPER_REJECT_MAX_RETRIES |
3 |
3 |
Used when policy is limited |
update_run_ledger.sh (via domain_persistence_script) prunes runs older than 30 days. Detect places skipped runs in ignored[]; the skill only notes them in Overview.
no_changes_verdict: REJECT¶
When the Skill classifies Fix but the maker produces no changes → REJECT → outcome: rejected.
When the Skill classifies Watch (infra/flake/env) with no code edit → outcome: watch (not REJECT). See Outcome enum.
PR exclusion (pr_exclude)¶
| Rule | Default | Token |
|---|---|---|
| Fork PR | Exclude | fork |
| Draft | Exclude | draft |
| Label opt-out | Exclude | label:<name> |
| Bots | Exclude | use pr_include_bots to opt in |
| WIP title | Optional | wip_title |
No label opt-in (pr_require) — eligible open PRs passing pr_exclude are watched when pr_enabled: true.
After finalize on pull_request targets, loop-notify-pr posts or updates a marker comment on the human PR (target_json.to.pr_number), including the bot fix PR URL when finalize creates one. See loop-notify-pr Specification.
Execute¶
- Worktree:
target.from(see platform table) verifier_context: failed job log excerpt from detectresult— always wired (integration and pull_request)
with:
target_json: ${{ toJson(matrix.target.target_json) }}
verifier_context: ${{ matrix.target.verifier_context }}
Checker criteria (caller-owned)¶
CI sweeper criteria require the fix to address the logged failure (semantic fit against verifier_context). This does not violate the generic rule “checker does not re-run CI” — see Loop Engineering — Verify.
Finalize strategy¶
PR body is composed by loop-finalize from agent ## Overview / ## Summary (skill-owned) plus mechanical sections. Dogfood sets pr_body: "". See Loop PR Body Skill Contract.
Platform rule for dogfood loops (changelog, docs-updater, ci-sweeper): target.finalize is always open_pr. level controls review vs auto-merge on the bot fix PR.
| Mode | L2 | L3 |
|---|---|---|
integration |
Bot fix PR → to.branch; human merge |
Bot fix PR → to.branch; GitHub auto-merge |
pull_request |
Bot fix PR → PR head; comment on human PR | Bot fix PR → PR head; auto-merge; comment on human PR |
Reference: Finalize strategy matrix.
| Persistence | Mechanism |
|---|---|
| State | state-ci-sweeper.json (targets map) on branch_state |
| Run ledger | domain_persistence_script → update_run_ledger.sh |
| Run log | loop-run-log in loop-finalize chain |
Dogfood: level=L2 until L3 promotion gate.
Implementation Checklist¶
Shared platform: Multi-Branch — Shared platform checklist.
Loop-specific¶
- [x]
ci-sweeper/scripts/detect_ci_failures.sh(facts output) - [x]
on-loop-ci-sweeper.yamldogfood caller viaci-loop-caller - [x]
verifier_contexton execute path (build_verifier_context_from_result.failuresbranch) - [x] Single detect path via
loop-detect(no caller re-run) - [x]
pr_enabled - [x] Ledger via
domain_persistence_scriptinloop-finalize - [x]
outcome: watchfor Skill Watch classification - [x]
loop-notify-pron human PR forpull_requestmode - [x]
open_prgit landing for PR head targets (derived fromdelivery: open_pr)
workflow_run Operational Checklist¶
Dogfood on-loop-ci-sweeper.yaml enables workflow_run. Keep these gates when changing the trigger or workflows: list:
- [x] zizmor / security review complete (
zizmor: ignore[dangerous-triggers]documented on the trigger) - [x]
workflow_run.workflowslists only CI workflows to repair (noton-loop-*/ci-loop-*callers) - [x] Concurrency prevents overlapping runs on same workflow
- [x] Event path sets
CI_SWEEPER_*keys indetect_domain_env_json - [x] Failed workflow is not the sweeper’s own finalize push (
[skip ci]on ledger commits) - [x] Fork / draft / bot exclusions active (
pr_exclude) - [x] Job
if:limits event runs tofailure/startup_failure - [x] 1 failure event → 1 target (no unbounded matrix from one event) —
loop-detectscopes watch list to failed head whenCI_SWEEPER_WORKFLOW_RUN_IDis set; verify when expandingworkflows:
Dependency update (caller filter + domain skill)¶
Tier 3 dependency-update behavior is a domain skill plus caller PR filters (pr_include_bots, pr_exclude) under ci-sweeper — not a separate loop package. Defer via Failure kind defer (B) until the skill exists.
Cross-Loop Note¶
CI failure on integration:main is serialized with docs-updater and changelog via workflow concurrency. Limit recursion with workflow_run.workflows (caller allowlist), run ledger (workflow_run_id), and daily budget — not workflow-name exclude lists in detect.
workflow_dispatch (no event run ID) uses gh run list on the watch branch (SCAN_BRANCH_RUN_LIMIT, default 100), then ledger and since range filters. Skill classifies infra/env/flake failures as Watch.