Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -132,6 +132,8 @@ jobs:
statuses: write
pull-requests: write
issues: write
# #487: `gh run rerun --failed` of the pinned failed fixture run (PB_GH_FIXTURE_FAILED_RUN).
actions: write
steps:
- uses: actions/checkout@v4

Expand Down Expand Up @@ -160,6 +162,9 @@ jobs:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
PB_GH_FIXTURE_PR: ${{ vars.PB_GH_FIXTURE_PR || github.event.pull_request.html_url }}
PB_GH_ALLOW_WRITES: ${{ (github.event_name != 'pull_request' || github.event.pull_request.head.repo.full_name == github.repository) && '1' || '' }}
# #487: a run of ci-fixture-fails.yml (dispatch-only, always fails) that the rerun seam
# reruns for real — a public run id, never a credential. Needs `actions: write` above.
PB_GH_FIXTURE_FAILED_RUN: ${{ vars.PB_GH_FIXTURE_FAILED_RUN }}
# tests/test_attach_pr_gh_402.py: the attach edge's `pr_identity` seam (#402), REAL in the
# seam ratchet. It needs no `br`, so it runs here, where the fixture is.
# tests/test_publish_gate_real.py: the publish-gate evaluators (`waits_for`: npm registry
Expand Down
10 changes: 10 additions & 0 deletions changelog.d/487.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
- **A red PR's failed CI jobs are rerun once before a coder fix round is spent (#487).**
A flaky job used to cost a model run and a `ci_fix_max` unit on a diff that was fine. Now
the CI reconcile first runs `gh run rerun <id> --failed` for each failing GitHub Actions
run in the PR's checks, stamps the card `ci-rerun:<sha>:<n>` and spends nothing. Green
after the rerun logs one `CI flake` line naming the checks that failed and clears the
stamp; red again at the same head bounces into a fix round as before. A new push gets a
fresh allowance, and the stamp on the bead means a restart never reruns a head twice. With
no Actions run behind the red checks, or a `gh` refusal, the card bounces at once. The new
`ci_rerun_max` (default `1`, `0` turns it off) sets how many reruns a head gets. Rerunning
needs `actions: write` on the board's `gh` token.
13 changes: 12 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ list — so an undocumented knob cannot be added quietly.

**`· YAML only`** marks a key the Settings UI cannot edit: it is absent from
`protoagent.plugin.yaml`'s schema, so `POST /api/settings` refuses it and the console
never renders it. **29 of 79 keys are in this state, including `coders` and `projects`** —
never renders it. **29 of 80 keys are in this state, including `coders` and `projects`** —
the two you must set for a multi-repo board. Edit
`~/.protoagent/<instance>/config/langgraph-config.yaml` directly, then restart.

Expand Down Expand Up @@ -332,11 +332,22 @@ The gates between a green build and main.
| `auto_merge_max` | `3` | reload |
| `merged_verify_max` | `5` | reload |
| `ci_fix_max` | `2` | reload |
| `ci_rerun_max` | `1` | reload |
| `review_fix_max` | `2` | reload |
| `review_gate_timeout_s` | `1800` | reload |
| `external_review` | `—` | reload **· YAML only** |
| `auto_merge` | `False` | live |

`ci_rerun_max` is how many times a red PR's failed GitHub Actions jobs are rerun per PR head
before `ci_fix_max` spends a coder fix round on them (#487). A flaky job is not the coder's bug.
On the first red, the board runs `gh run rerun <id> --failed` for each failing Actions run,
stamps the bead `ci-rerun:<sha>:<n>` and spends nothing. If the rerun comes back green, it
logs one `CI flake` line naming the checks that failed and clears the stamp. If it is red again
at the same head, the card bounces into a fix round as before. A new push is a new head and
gets a fresh allowance. With no Actions run behind the red checks (only a non-Actions required
status failed), or when `gh` refuses the rerun, the card bounces at once. Rerunning needs the
board's `gh` token to hold `actions: write` on the repo. `0` turns reruns off.

`review_gate_timeout_s` is a hard cap on one review-gate model call: the host workflow run,
or the a2a reviewer fallback. The cap is the board's own, not the host model client's. That
client's `request_timeout` did not apply to a hung stream (protoAgent#3699), and one review
Expand Down
10 changes: 10 additions & 0 deletions loop/_common.py
Original file line number Diff line number Diff line change
Expand Up @@ -446,6 +446,15 @@ def _ci_failure_reason(summary: str, max_chars: int = 500) -> str:
return reason[:max_chars]


def _ci_failed_check_names(summary: str) -> str:
"""The failing check names from a ``pr_ci_status`` summary (``- <name>: <CONCLUSION>``
lines above the log excerpt), joined for one log line — what a CI rerun (#487) was
spent on. ``""`` when the summary names none."""
head = (summary or "").split("\n\nFailing log", 1)[0]
lines = [ln[2:].strip() for ln in head.splitlines() if ln.startswith("- ")]
return ", ".join(ln.rpartition(":")[0].strip() or ln for ln in lines)


_PR_URL_RE = re.compile(r"github\.com/([^/]+/[^/]+)/pull/(\d+)")


Expand Down Expand Up @@ -1485,6 +1494,7 @@ def _inbox_db_path():
"_is_test_path",
"_is_code_path",
"_CI_SIGNAL_RE",
"_ci_failed_check_names",
"_ci_failure_reason",
"_PR_URL_RE",
"_parse_pr_url",
Expand Down
8 changes: 8 additions & 0 deletions loop/core.py
Original file line number Diff line number Diff line change
Expand Up @@ -175,6 +175,11 @@ def __init__(self, cfg: dict, *, gap_reporter: setup_check.GapReporter | None =
# feature is blocked for human triage (a real bug, not a self-fixable nit).
self.ci_poll = bool(self.cfg.get("ci_poll", self.merge_poll))
self.ci_fix_max = max(0, int(self.cfg.get("ci_fix_max", 2)))
# #487: before a red rollup spends a `ci_fix_max` unit, rerun its failed GitHub
# Actions jobs up to this many times per PR head — a flaky job is not the coder's
# bug. Recorded on the bead (`ci-rerun:<sha>:<n>`), so a restart can't rerun a
# head twice; a new push re-arms it. 0 disables (bounce on the first red, as before).
self.ci_rerun_max = max(0, int(self.cfg.get("ci_rerun_max", 1)))
# Auto-rebase a stale/conflicting in_review PR onto base. Parallel PRs branch
# off the SAME base, and the hot-file guard serializes DISPATCH not the branch
# BASE — so each merge re-stales the others (a sibling's change lands in the
Expand Down Expand Up @@ -473,6 +478,9 @@ def __init__(self, cfg: dict, *, gap_reporter: setup_check.GapReporter | None =
# The prompt-feedback dicts (_ci_feedback/_ci_prior_diff/_review_prior)
# stay memory-only: they enrich the next prompt, they never gate anything.
self._ci_feedback: dict[str, str] = {}
# #487: the failed check names a CI rerun was spent on, for the flake log line
# when the rerun comes back green (in memory only; a restart just names fewer).
self._ci_rerun_checks: dict[str, str] = {}
self._ci_prior_diff: dict[str, str] = {}
self._ci_fix_attempts: dict[str, int] = {}
# Pre-PR goal-verify gap re-dispatches so far (fid → count), same-tier.
Expand Down
83 changes: 83 additions & 0 deletions loop/reconcile.py
Original file line number Diff line number Diff line change
Expand Up @@ -1886,8 +1886,14 @@ async def _reconcile_ci(self, store, fid: str, pr_url: str, repo: str, feature:
if await worktree.pr_state(pr_url, cwd=repo) != "OPEN":
return # merged/closed since the poll started -> never dispatch a CI fix
status, summary = await worktree.pr_ci_status(pr_url, cwd=repo)
if status == "passing":
await self._settle_ci_rerun(store, fid, pr_url, repo, feature)
if status != "failing":
return
# #487: a red rollup may be a flake. Rerun its failed Actions jobs once per head
# before spending a fix round; the next pass sees the rerun's verdict.
if await self._rerun_ci_once(store, fid, pr_url, repo, feature, summary):
return
# Carry the lesson: the CI error + the diff that failed it (best-effort).
self._ci_feedback[fid] = summary
self._ci_prior_diff[fid] = await worktree.pr_diff(pr_url, cwd=repo)
Expand Down Expand Up @@ -1935,6 +1941,83 @@ async def _block(reason: str):
await worktree.reap_feature_worktree(repo, self.root, fid)
log.warning("[project_board] reconcile → blocked (CI fails, %d attempt(s) exhausted): %s", attempts, fid)

async def _rerun_ci_once(self, store, fid: str, pr_url: str, repo: str, feature: dict | None, summary: str) -> bool:
"""Rerun a red PR's failed GitHub Actions jobs before a coder fix round (#487).

Returns True when a rerun was started: the caller then spends nothing and requeues
nothing, and a later pass reads the rerun's verdict (green → ``_settle_ci_rerun``
logs the flake; red again at the same head → the bounce below runs as it always has).

At most ``ci_rerun_max`` reruns per PR head, counted on the bead's
``ci-rerun:<sha>:<n>`` label so a restart can't rerun a head again; a new push is a
new head and gets a fresh allowance. Returns False — bounce exactly as before — when
reruns are off (``ci_rerun_max: 0``), this head's allowance is spent, the head can't
be read to prove otherwise, or nothing was rerun (no Actions run behind the red
checks, e.g. only a non-Actions required status failed, or ``gh`` refused)."""
cap = int(getattr(self, "ci_rerun_max", 0) or 0)
if cap <= 0:
return False
stamped, used = store_mod.ci_rerun_from_labels((feature or {}).get("labels"))
head = ""
if stamped:
# Only a stamp needs the head up front: without one this is the head's first
# rerun whatever it is, and the head is read after the rerun to stamp it.
head = await worktree.pr_head_sha(pr_url, cwd=repo)
if not head:
return False # can't prove this head still has an allowance → bounce as before
if head[: store_mod.SHORT_SHA_LEN] != stamped:
used = 0 # a new push since the last rerun: a fresh allowance
elif used >= cap:
return False # this head was already rerun and is red again → a real failure
run_ids = await worktree.rerun_failed_ci(pr_url, cwd=repo)
if not run_ids:
return False
if not head:
head = await worktree.pr_head_sha(pr_url, cwd=repo)
short = head[: store_mod.SHORT_SHA_LEN] if head else "an unreadable head"
self.__dict__.setdefault("_ci_rerun_checks", {})[fid] = _ci_failed_check_names(summary)
if head:
try:
await asyncio.to_thread(store.record_ci_rerun, fid, head, used + 1)
except Exception as exc: # noqa: BLE001 — the rerun is already running; a failed
# stamp only means a later red at this head may be rerun once more.
log.warning("[project_board] %s: could not stamp the CI rerun at %s: %s", fid, short, exc)
log.info(
"[project_board] %s CI red at %s — rerunning failed jobs once before a fix round: %s (%s)",
fid,
short,
", ".join(run_ids),
pr_url,
)
return True

async def _settle_ci_rerun(self, store, fid: str, pr_url: str, repo: str, feature: dict | None) -> None:
"""A green rollup on a card with a ``ci-rerun:`` stamp (#487): the rerun passed with
no fix round. Log ONE flake line naming the checks that failed the first time (when
the head is still the one that was rerun; a green new head is just a fixed PR), and
clear the stamp. Best-effort: an unreadable head leaves the stamp for the next poll."""
checks_by_fid = self.__dict__.setdefault("_ci_rerun_checks", {})
stamped, _used = store_mod.ci_rerun_from_labels((feature or {}).get("labels"))
if not stamped:
checks_by_fid.pop(fid, None)
return
head = await worktree.pr_head_sha(pr_url, cwd=repo)
if not head:
return
checks = checks_by_fid.pop(fid, "") or "the failed checks (names not kept across a restart)"
if head[: store_mod.SHORT_SHA_LEN] == stamped:
log.info(
"[project_board] %s CI flake: the rerun at %s passed with no fix round spent — failed first: %s (%s)",
fid,
stamped,
checks,
pr_url,
)
try:
await asyncio.to_thread(store.record_ci_rerun, fid, "")
except Exception as exc: # noqa: BLE001 — a stale stamp only costs a rerun allowance
log.warning("[project_board] %s: could not clear the CI rerun stamp: %s", fid, exc)

async def _request_review(self, fid: str, pr_url: str):
"""Hand the PR to the reviewer (an a2a delegate, e.g. quinn). Best-effort:
a review-dispatch failure doesn't block the feature — CI + the merge
Expand Down
7 changes: 7 additions & 0 deletions protoagent.plugin.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -193,6 +193,11 @@ config:
ci_fix_max: 2 # cap on CI-fix re-dispatches before the feature is Blocked for
# human triage (a real bug, not a self-fixable nit). 0 = never
# auto-fix (block on the first CI failure).
ci_rerun_max: 1 # #487: rerun a red PR's failed GitHub Actions jobs this many times
# per PR head BEFORE a ci_fix_max fix round — a flaky job is not
# the coder's bug. Green after the rerun → a logged flake, nothing
# spent; red again at the same head → the usual bounce. A new push
# re-arms it. Needs `actions: write` on the gh token. 0 = off.
auto_rebase: true # keep stale in_review PRs mergeable (v0.14.0, bd-2gu). Parallel
# PRs branch off the SAME base and the hot-file guard serializes
# DISPATCH not the branch BASE, so each merge re-stales the others.
Expand Down Expand Up @@ -329,6 +334,8 @@ settings:
description: "Poll CI for in-review PRs and re-dispatch failures instead of leaving red PRs parked forever." }
- { key: ci_fix_max, label: "CI fix attempts", type: number, minimum: 0, maximum: 20, group: "Project Board", tab: review, restart: true, depends_on: { key: ci_poll },
description: "Automatic CI-fix re-dispatches before the card blocks for human triage. 0 blocks on the first failure." }
- { key: ci_rerun_max, label: "CI reruns before a fix", type: number, minimum: 0, maximum: 5, group: "Project Board", tab: review, restart: true, depends_on: { key: ci_poll },
description: "Rerun a red PR's failed GitHub Actions jobs this many times per head before spending a CI fix attempt on a possible flake. 0 turns reruns off." }
- { key: auto_rebase, label: "Rebase stale PRs", type: bool, group: "Project Board", tab: review, restart: true,
description: "Cleanly rebase behind PRs and re-dispatch genuine conflicts so parallel work remains mergeable." }
- { key: rebase_fix_max, label: "Conflict fix attempts", type: number, minimum: 0, maximum: 20, group: "Project Board", tab: review, restart: true, depends_on: { key: auto_rebase },
Expand Down
41 changes: 41 additions & 0 deletions store.py
Original file line number Diff line number Diff line change
Expand Up @@ -616,6 +616,13 @@ def split_breadth(files, patterns) -> tuple[list[str], list[str]]:
# and a new push is a new head, which re-arms it. On the bead, not in memory, so a
# restart can't re-bounce a head the last process already bounced.
LABEL_EXTERNAL_REVIEW_BOUNCED_PREFIX = "ext-review-bounced:"
# The PR head whose failed CI jobs the reconcile already RERAN (#487) —
# `ci-rerun:<sha>:<n>`, replaced (never accumulated), SHORT for the 50-char label cap. A
# red rollup is rerun (`gh run rerun --failed`) up to `ci_rerun_max` times per head before
# a coder fix round is spent on what may be a flake; `<n>` is how many reruns this head has
# had. On the bead, not in memory, so a restart can't rerun a head the last process already
# reran. A new push is a new head, which re-arms the allowance; a green rollup clears it.
LABEL_CI_RERUN_PREFIX = "ci-rerun:"
# The ORIGINATING GitHub issue (#97) — a structured `source-issue: owner/repo#N`
# metadata line in the bead `notes` field, beside the files_to_modify path lines.
# NOT a label: beads' label validator only allows alphanumeric/hyphen/underscore/
Expand Down Expand Up @@ -671,6 +678,19 @@ def budgets_from_labels(labels) -> dict[str, int]:
return out


def ci_rerun_from_labels(labels) -> tuple[str, int]:
"""``(short_head, reruns)`` from a bead's ``ci-rerun:<sha>:<n>`` label (#487), or
``("", 0)`` when there is none. A malformed count reads as 1: the label only exists
once a rerun was spent, so it must never re-arm an allowance it cannot parse."""
for label in labels or []:
if not str(label).startswith(LABEL_CI_RERUN_PREFIX):
continue
head, _, num = str(label)[len(LABEL_CI_RERUN_PREFIX) :].partition(":")
if head:
return head, int(num) if num.isdigit() and int(num) > 0 else 1
return "", 0


def replace_prefixed_label_args(labels, prefix: str, desired: str) -> list[str]:
"""`br update` args that REPLACE the single ``<prefix>…`` label with ``desired`` —
the one correct spelling of the single-label-replaced pattern (`diff:`, `gens:`,
Expand Down Expand Up @@ -3976,6 +3996,27 @@ def record_merged_verified(self, fid: str, sha: str) -> dict:
self._run(*args)
return self.get_feature(fid)

# ── CI rerun stamp (#487) ─────────────────────────────────────────────────
def record_ci_rerun(self, fid: str, head: str = "", n: int = 1) -> dict:
"""Stamp the PR head whose failed CI jobs were rerun, and how many times (#487) —
a single, replaced ``ci-rerun:<sha>:<n>`` label (the ``merged-verified:`` pattern),
SHORT-abbreviated for beads' 50-char label cap. ``head=""`` CLEARS the stamp (the
rerun came back green, or the head moved on). No-op when there is nothing to change,
so a green poll never burns a ``br`` write."""
f = self._require(fid)
labels = f.get("labels") or []
short = str(head or "").strip()[:SHORT_SHA_LEN]
if short:
args = replace_prefixed_label_args(
labels, LABEL_CI_RERUN_PREFIX, f"{LABEL_CI_RERUN_PREFIX}{short}:{max(1, int(n))}"
)
else:
args = [a for l in labels if str(l).startswith(LABEL_CI_RERUN_PREFIX) for a in ("--remove-label", str(l))]
if not args:
return f
self._run("update", fid, *args)
return self.get_feature(fid)

# ── review-verdict head stamp (#328) ──────────────────────────────────────
def record_reviewed_head(self, fid: str, sha: str) -> dict:
"""Stamp the PR head sha the active review verdict was rendered against (#328)
Expand Down
10 changes: 10 additions & 0 deletions tests/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -142,6 +142,7 @@ def _no_network(url, timeout=0.0):
"worktree.untagged_release_head": _worktree_mod.untagged_release_head,
"worktree.refresh_base_checkout": _worktree_mod.refresh_base_checkout,
"worktree.pr_review_state": _worktree_mod.pr_review_state,
"worktree.rerun_failed_ci": _worktree_mod.rerun_failed_ci,
}


Expand Down Expand Up @@ -177,6 +178,15 @@ async def _no_release_gap(_slug, _base, _patterns, *, cwd="."):
monkeypatch.setattr(_worktree_mod, "open_pr_heads", _no_prs)
monkeypatch.setattr(_worktree_mod, "active_workflow_runs", _no_runs)
monkeypatch.setattr(_worktree_mod, "untagged_release_head", _no_release_gap)

# #487: the CI reconcile reruns a red PR's failed Actions jobs before a fix round. The
# unit tier never asks GitHub to rerun anything: by default nothing is rerun, so every
# pre-existing CI-bounce test bounces on the first red exactly as before. A test of the
# rerun edge injects its own; tests/test_publish_gate_real.py restores the real seam.
async def _no_rerun(_pr_url="", *, cwd=".", run_ids=None, slug=""):
return []

monkeypatch.setattr(_worktree_mod, "rerun_failed_ci", _no_rerun)
_gates_mod.reset_cache()
_freeze_mod.reset_state()
yield
Expand Down
Loading
Loading