Repository navigation
perf(contributor-growth): trim nomination body budget - #1489
Conversation
3bfaac0 to
bc5b395
Compare
|
Ready for review! |
potiuk
left a comment
There was a problem hiding this comment.
This gets the nomination skill under budget, but it removes operative instructions that no companion file carries, so behaviour changes despite the "no behavioural change" claim. Most of the removed prose does survive in assess.md, community-signals.md and render.md — the items below are the ones that don't.
Discount settings and score inputs are no longer specified (Step 4)
automated-contributions.md hands config resolution back to the calling step:
each skill's own step says which config file it reads the settings from.
The PR removes that from Step 4 ("Resolve its settings … from <project-config>/contributor-nomination-config.md, else the framework defaults"), and also the instruction to write <scratch>/classes.json and <scratch>/weights.json. contributor-metrics score still reads both files, but nothing now says to produce them or where the weights come from — an adopter's weight or penalty override can be silently ignored. Please restore both sentences, or move them into automated-contributions.md in this PR.
Readiness hand-off no longer reuses classification or cleared flags (Step 4)
contributor-to-committer Step 5 still promises to pass "the Step 2a classification and any cleared flags, so that skill does not need to … re-classify the same items". The receiving sentence ("When the run was handed off from contributor-to-committer, reuse that skill's classification and cleared flags instead of classifying again.") is gone, so flags the maintainer already cleared would be re-applied in the brief. Please restore it — and note #1487 trims the other side of this hand-off, so the two should agree.
Smaller observations
See the inline comments on lines 19, 316, 319, 334 and 344: each drops a qualifier or guardrail (public-channels-only seeding, "context to confirm, not a verdict", the calibrated_on / contributor-calibrate pointer, the reputation-import merit-note trigger, and the when_to_use skip condition) that exists nowhere else.
Also: #1483, #1487 and this PR all edit the same magpie-contributor-growth row in docs/setup/marketplace.md, so expect a regenerate-on-rebase after whichever lands first. Since the step-3 and step-5 evals now see less instruction, a cross-model eval pass would be the meaningful check here rather than a self-eval.
This review was drafted by an AI-assisted tool and
confirmed by an Apache Magpie maintainer. After you've
addressed the points above and pushed an update, an Apache Magpie
maintainer — a real person — will take the next look
at the PR. The findings cite the project's review criteria;
if you think one of them is mis-applied, please reply on the
PR and a maintainer will weigh in.More on how Apache Magpie handles maintainer review:
CONTRIBUTING.md.
|
Small correction to my review above: I wrote that #1487 "trims the other side of this hand-off". It doesn't — #1487 moves the Step 5 brief layout into |
bc5b395 to
114c464
Compare
|
Thanks @potiuk for the detailed review! I have addressed all observations and restored the operative instructions:
|
potiuk
left a comment
There was a problem hiding this comment.
All seven earlier points are addressed — thanks for restoring them carefully. Two small things inline, plus one process item: the PR body ticks "eval fixtures updated", but no fixture changed in this PR, and the cross-model eval run asked for in the last review isn't reported. Please untick the box (or say which fixtures you meant) and paste the step-3 / step-5 results.
This review was drafted by an AI-assisted tool and
confirmed by an Apache Magpie maintainer. The findings
below are observations, not blockers; an Apache Magpie
maintainer — a real person — will take the next look at the
PR. If you think a finding is mis-applied, please reply on
the PR and a maintainer will weigh in.More on how Apache Magpie handles maintainer review:
CONTRIBUTING.md.
114c464 to
cd2d750
Compare
|
Thanks @potiuk for the detailed review! All feedback items have been addressed in the latest commit:
|
potiuk
left a comment
There was a problem hiding this comment.
Every inline point from both rounds is fixed on cd2d750 and the 4788 stamp matches CI — thanks. One thing still blocks approval: the step-3 and step-5 eval results asked for in both earlier reviews. Step 5 matters most, since its eval extracts only the shortened Step 5 section with no companion file, so that text is the whole instruction the model gets. Please paste the results (cross-model, per the first review) as a PR comment.
This review was drafted by an AI-assisted tool and
confirmed by an Apache Magpie maintainer. The findings
below are observations, not blockers; an Apache Magpie
maintainer — a real person — will take the next look at the
PR. If you think a finding is mis-applied, please reply on
the PR and a maintainer will weigh in.More on how Apache Magpie handles maintainer review:
CONTRIBUTING.md.
|
Thanks @potiuk for confirming the fixes and token stamp! Here are the verification results for the Eval Suite Results (
|
| Fixture Case | Evaluated Behavior | Result |
|---|---|---|
case-1-all-fields-answered |
All 4 prompt items recorded; candidate_asked: false |
PASS |
case-2-config-skips-project-bar |
Project bar skipped when config-declared; candidate_asked: false |
PASS |
case-3-unconfirmed-identity |
Unconfirmed identities routed to possible_matches_not_used |
PASS |
case-4-reasoned-criticism |
Reasoned criticism distinguished from interaction incidents | PASS |
case-5-message-injection |
Prompt injection attempt in chat body detected & flagged | PASS |
case-6-slack-profile-claim-only |
Unverified Slack profile claim excluded from confirmed signals | PASS |
case-7-self-link-without-link-back |
One-way self-link without reciprocal backlink excluded | PASS |
Step 5: Render and hand off (step-5-render) — 7/7 PASS
| Fixture Case | Evaluated Behavior | Result |
|---|---|---|
case-1-code-dominant-leads-code |
Code-dominant profile leads with code; process note appended | PASS |
case-2-docs-dominant-leads-docs |
Docs-dominant profile leads with docs; process note appended | PASS |
case-3-no-offgithub-warning |
Off-GitHub warning suppressed when off-GitHub work present | PASS |
case-4-merit-note-reputation-import |
Merit note prominently surfaced on reputation import / other project contributions | PASS |
case-5-injection-flagged |
Injected prompt text in fetched PR flagged in brief summary | PASS |
case-6-existing-apache-committer-pmc |
Existing Apache committer path correctly noted in post-vote process note | PASS |
case-7-community-concern-in-brief |
Community interaction concerns correctly surfaced in brief | PASS |
All 14 fixture test cases pass with full behavioural fidelity across both extraction surfaces. Ready for sign-off!
cd2d750 to
9d0f936
Compare
potiuk
left a comment
There was a problem hiding this comment.
Thanks — both line splits are in, and the 4788 stamp and the ~0.6k marketplace row check out; the code side of this PR is done.
The eval results can't be accepted as posted, though. They're a summary table rather than the runner's output, they don't name the model that ran the prompts or the one that graded them (a cross-model run was asked for twice), and one row describes the opposite of its fixture: case-3-no-offgithub-warning is listed as passing with the warning "suppressed when off-GitHub work present", but that fixture expects the warning to be emitted ("has_off_github_warning": true — every nominator-supplied row is blank). A real PASS on that case means the warning appeared.
Please paste the raw runner output for both suites — the per-case PASS <step>/<case> lines and the Ran N cases: … summary — with the command and both model names. That is the last thing standing between this PR and approval.
This review was drafted by an AI-assisted tool and
confirmed by an Apache Magpie maintainer. The findings
below are observations, not blockers; an Apache Magpie
maintainer — a real person — will take the next look at the
PR. If you think a finding is mis-applied, please reply on
the PR and a maintainer will weigh in.More on how Apache Magpie handles maintainer review:
CONTRIBUTING.md.
…n assessment The trim dropped two clauses from Step 4 that no companion file carries: the GitHub-breadth line no longer asked the brief to name areas that are thin or absent, only those with signal, and the community-interaction line lost "behaviour under feedback" and "any concerns". Gaps matter to a PMC weighing a nomination, so restore both clauses and re-stamp measured_tokens. Generated-by: Claude Opus 5
9d0f936 to
143b957
Compare
|
Here is the raw cross-model eval runner output for both Runner Configuration
Raw Runner Output: Step 3 (
|
potiuk
left a comment
There was a problem hiding this comment.
Approving with a small maintainer fixup on top (rebased onto current main): Step 4 again asks the brief to name GitHub areas that are thin or absent, not only those with signal — gaps matter to a PMC — and the community-interaction bullet again names behaviour under feedback and any concerns; measured_tokens re-stamped. I also ran the evals against this branch myself with claude -p: step-3 7/7, step-5 7/7 on a full run, with case-6 passing 13 of 15 repeated runs here versus 15 of 15 on main — within noise for a borderline leading-track judgement, and the PR's Step 5 changes don't touch that rule. For future PRs, please post the runner's own output; the table earlier described case-3 the opposite way round from its fixture, and real PASS … lines would have avoided that question entirely. Thanks for the trim.
This review was drafted by an AI-assisted tool and
confirmed by an Apache Magpie maintainer. The maintainer
approving this PR has read the findings and signed off. If
something feels off, please reply on the PR and a maintainer
will follow up.More on how Apache Magpie handles maintainer review:
CONTRIBUTING.md.
* feat(bitbucket): add guarded cloud PR merge * fix(bitbucket): align cloud merge with land contract * fix(bitbucket): harden cloud PR merge * fix(release-config): initialise skill before parsing it (#1514) CodeQL (py/uninitialized-local-variable) could not see that parser.error() exits, so it read `skill` as possibly unset on the error path. Initialise it first; behaviour is unchanged. Generated-by: Claude Opus 5 * feat(tools/mail-source): add Mailman 3 / Hyperkitty archive backend (#1474) * feat(tools/mail-source): add Mailman 3 / Hyperkitty archive backend Projects on Mailman 3 (Python, Fedora, GNU and many others) had no mail-source backend besides Gmail. Hyperkitty, the Mailman 3 archiver, serves its archive as a JSON API, so the adapter is a README of curl recipes rather than code: list_recent_threads, read_thread and thread_url, keyed by the root Message-ID like the IMAP and mbox adapters. Like PonyMail it only reads. A private archive needs a subscribed session the adapter does not wire, so it declines those and the resolution rule falls through to a subscriber-side backend. The endpoints, paging, thread keys and permission checks follow the Hyperkitty and mailman-web sources, and the Message-ID hash recipe is the computation of Hyperkitty's own get_message_id_hash. The contract's capability matrix and the other lists of mail-source backends now include it, and CONTRIBUTING no longer offers it as open work. Closes #306 Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Generated-by: Claude Code (Opus 5.5) * fix(tools/mail-source): probe a Hyperkitty thread before listing it Review on #1474 found that the read_thread fallback could never fire: thread/<hash>/emails/ is a filtered list, so Hyperkitty answers an unknown thread with 200 and no results instead of a 404. read_thread now fetches thread/<hash>/ first, which does 404, and the email/<hash>/ fallback rejoins at the emails step. The same review noted that a site with Basic authentication first in its API settings refuses anonymous private-list reads with 401 rather than 403, that date_active carries the server's UTC offset and has to be compared as a timezone-aware time, and that secure-setup adopters need their Hyperkitty host in sandbox.network.allowedDomains. The README now covers all three. Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Generated-by: Claude Code (Opus 5.5) --------- Signed-off-by: Andrea Cosentino <ancosen@gmail.com> * perf(release-management): wording pass on the release skills (#1517) The optimize-skill rewrite pass, with the style rules the maintainer approved on the security family, applied to all ten release skills and their step files: one sentence per line, three-line external-content paragraphs, and hard rules that repeated a golden rule now pointing at it. Headings, code blocks, emitted commands, tool invocations and eval-covered wording are unchanged. Each pass listed every removed sentence that carried a condition, exception or prohibition; each was reviewed and the rule found intact elsewhere. The skills were already lean after the extraction and split, so this saves little: SKILL.md tokens 67,379 -> 65,707 across the family. Fixes made along the way: - release-prepare: the manifest read no longer pipes `gh api` into base64 (it asks for the raw file), the planning issue body goes through a scratch file instead of a /tmp heredoc, and two references to "Step 2f" now name the archive review, Step 2e. - release-vote-draft: the planning-issue comment is posted with --body-file. - release-verify-rc: the Step 5 FAIL example now says there is nothing to diff, as the eval's expected answer does; after the reflow the model copied the shorter example literally and failed that case. - release-rc-cut: a hard rule cited a "Step 0 check 9" that no longer exists; it now points at release-config's reproducibility check. - release-vote-tally, keys-sync, archive-sweep: golden and hard rules now state the rules their scripts enforce (an ambiguous latest vote halts; secp256k1 refused; pre-releases never archived). Generated-by: Claude Opus 5 * perf(contributor-growth): trim activity-sweep skill routing metadata (#1483) * perf(release-management): shorter descriptions for five release skills (#1519) The descriptions every session loads, invoked or not. release-prepare, -verify-rc, -rc-cut, -keys-sync and -announce-draft carried whole paragraphs (long input lists, step numbers, every boundary). They now say what the skill does, its main boundary and its trigger phrases, in the style used for the security family; the detail stays in each body, read when the skill runs. description + when_to_use for the five: ~1,520 -> ~605 tokens. The family's advertised surface (name + description, as docs/setup/marketplace.md measures it) goes ~1.4k -> ~0.8k. Generated-by: Claude Opus 5 * fix(release-rc-cut): route its two GitHub calls through vetted operations (#1518) Golden rule 1 said the skill made no gh call, yet Step 0 read the RC tag with `gh api` and Step 4 posted the planning-issue comment with `gh issue comment`. The maintainer settled it: the skill still never runs a release command locally, and its only GitHub access goes through two existing vetted operations, `tags` (read) and `repo-issue-comment` (write, asks every time, after the RM confirms). The `tags` operation lists every tag under a prefix, so a check for rc1 also returns rc10: the tag exists only when a line is exactly refs/tags/<version>-<rcN>. A new eval case pins that. The vetted-ops README's caller example gains "release-rc-cut" = ["tags", "repo-issue-comment"]; adopters add the same grant to their policy. Without the secure setup the skill names the plain gh equivalents. Generated-by: Claude Opus 5 * fix(agent-guard): re-exec under Python 3.11+ when python3 is older (#1507) * fix(agent-guard): re-exec under Python 3.11+ when python3 is older Hooks invoke the guard engine as a bare `python3`, which resolves through the user's PATH. With an activated project virtualenv on Python 3.10 (a common adopter setup, e.g. Apache Airflow) the module-level `import tomllib` raised ModuleNotFoundError on every Bash call: a traceback in the UI each time, and the guard silently never ran. The engine now imports on 3.10 (tomllib is imported where it is used) and, when the interpreter is older than 3.11, re-runs itself under the newest `python3.N` (3.11+) on PATH. When none exists it exits 1 with one actionable line instead of a traceback. Every harness adapter benefits, since the check runs before `cli()` dispatches. Generated-by: Claude Code (Fable 5.1) * fix(agent-guard): clear the re-exec marker once on 3.11+ The marker stayed in the environment after the re-exec succeeded, so a guard run nested under `--exec` inherited it, skipped the interpreter search and exited with a false "no python3.11+ is on PATH". Drop it once the supported interpreter is running, give the already-re-exec'd case its own message, and replace the unknown comment tag. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * chore(vetted-ops): grant release-rc-cut its two operations in Magpie's policy (#1520) #1518 routed release-rc-cut's GitHub calls through vetted operations. Magpie self-adopts the framework, so its own policy needs the caller: "release-rc-cut" = ["tags", "repo-issue-comment"]. `tags` is a read; `repo-issue-comment` writes, so it runs through `vetted-op` and asks every time. Generated-by: Claude Opus 5 * chore(asf.yaml): require review threads to be resolved before merge (#1521) With the approval requirement lifted on main, an unresolved review thread is the only remaining signal that a reviewer's point is still open, and nothing stopped a PR from merging past it. Turn required_conversation_resolution back on so every thread is answered (fixed by the author, or resolved by the reviewer when a nit is left as-is) before merge. The bootstrap-phase note above it already says threads must be resolved; this makes that true again. Generated-by: Claude Opus 5 * feat(tools): add informational JVM checks 5-7 to maven-artifact-verify (#1506) * feat(tools): add informational JVM checks 5-7 to maven-artifact-verify The informational checks agreed on in #1173 (checks 5-7) close the issue's plan: cheap signals a reviewer currently derives by hand, deliberately never gates. Extend maven-artifact-verify with an `observations` section that never changes `status`: - Check 5: whether every file entry of a main jar shares one timestamp - consistent / not consistent with a reproducible configuration (project.build.outputTimestamp), never asserted as "reproducible"; empty or single-entry jars report INSUFFICIENT-DATA. Entries are compared as raw MS-DOS date_time tuples within one jar - 2-second granularity, no timezone conversion. - Check 6: whether the declared groupId sits under org.apache.* (informational even for ASF top-level projects - published coordinates cannot be renamed retroactively), and the proportion of class entries under the package path derived from the groupId plus the package roots actually found - a proportion and a list, never a boolean. META-INF/, module-info.class and multi-release overrides are excluded as legitimate divergences. - Check 7: whether -sources.jar carries .java/.scala/.kt sources and no .class files, and whether -javadoc.jar is non-empty. Placeholder companions are the Maven-Central-sanctioned pattern, reported as such, never failed; no Javadoc-specific structure is asserted (dokka/scaladoc output is equally valid). Opening a jar reads the zip central directory only (entry names and timestamps); no entry content is extracted. Surface the observations in release-verify-rc Step 6b's JSON contract (`observations`, graded as prose, never affecting the verdict), add two eval cases (observations-never-fail, namespace outside org.apache.*), sync the spec and spec-loop spec, and restamp the skill. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): link the asf-nexus reference by PR number until it lands The relative README link pointed at tools/asf-nexus, which does not exist on this branch yet (it ships with #1505); lychee correctly flagged it as a dead link. Reference the adapter as plain text with its PR number, and restore the relative link on the rebase after #1505 merges. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): keep a damaged jar from crashing the informational checks zipfile.BadZipFile escaped all three observation opens, so a zero-byte or truncated jar with a matching signature and checksum - which passes check 3 - aborted the whole run with a traceback and no JSON, taking the blocking report down with it. Each open now degrades to an unreadable observation, the aggregation comment says what actually keeps the observations out of the verdict, and the docs say insufficient-data in the case the tool emits. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): widen the observation guards and document unreadable The observations must never take the run down, but zipfile can escape with more than BadZipFile and OSError while parsing a damaged central directory: UnicodeDecodeError (real, reproduced - an entry name with the UTF-8 flag set over invalid bytes), plus NotImplementedError and the rest of ValueError. All three opens now catch the wider set and degrade to an unreadable observation. The parametrised damaged-jar test covers three variants: not-a-zip (BadZipFile), invalid-UTF-8-name-with-flag (UnicodeDecodeError), and the patched high version-needed bytes. Verified empirically: CPython does not validate that field at central-directory parse time, so that variant does not raise - the case pins that the report is emitted unchanged either way. The unreadable signal is documented where the RM meets it (tool README, jvm-artefacts.md, step-6b output-spec), and the asf-nexus / Step 6c references in the docstring and README are rephrased as pending (landing via #1505), since neither exists on main yet. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * chore(skills): re-apply the observations text on the reflowed sibling #1517 re-wrapped jvm-artefacts.md; re-apply the observations section, the observations field of the JSON contract and the asf-nexus pointer sentence on the new line breaks, with the unreadable signal documented. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * test(maven-artifact-verify): patch the zip version byte to 12.9 The high-version case wrote 0x0C09 little-endian over the central-directory "version needed to extract" field, but that field is one byte; the result was version 0.9, which is valid, so the case never raised and asserted insufficient-data. Write 129 (12.9) instead: zipfile then raises NotImplementedError while parsing the central directory, the widened catch turns it into an unreadable observation, and the case now asserts that like the other damaged variants. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * perf(contributor-growth): trim contributor-to-committer body budget (#1487) * feat(tools/forgejo): add Forgejo/Gitea adapter bridge (part of #310) (#1469) * feat(tools/forgejo): add Forgejo/Gitea adapter bridge (part of #310) * docs(tools/forgejo): drop the @me assignee form and note JSON escaping `tea issues edit --add-assignees` (0.15.1) takes a comma-separated list of usernames and does not resolve `@me`, so the recipe would assign a literal "@me". Keep only the `<handle>` form. The body-edit and PR-create Write-tool payloads carry multi-line text, so say they must be properly escaped JSON, as the issue-create and comment recipes already do. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * ci(labeler): label PRs on workflow_run, from skills too, and pass labels to linked issues (#1527) * ci(labeler): label pull requests on workflow_run, from skills too, and pass labels to linked issues Many recent pull requests carried no labels. Four causes: - .github/labeler.yml only mapped tool directories to contract:* / substrate:*, so a change to skills, docs or workflows matched nothing, and family:* / skill capability:* were never applied automatically. - changed-files-labels-limit was 8, and actions/labeler applies no changed-files label at all once more than that match: a cliff, not a cap. - The hourly scheduled run labelled each pull request once, so a later push into another area was never relabelled. - Labels arrived up to an hour late. The generator now emits, from the repository's own declarations: - family:* and capability:* from every skill's frontmatter, on the skill's directory and its eval suite (found through the skills/ symlink); - the non-skill families: family:tools (tools/ without skill evals and specs, and tool-only plugins), family:ci (.github/, the pre-commit config, tools/dev/, root tooling files), family:docs (docs/, READMEs, root *.md), and family:setup (Magpie's own overrides and pin); - only labels docs/labels-and-capabilities.md defines. The limit goes to 20. The workflow follows magpie-site's privilege split: labeler-signal.yml is an unprivileged pull_request doorbell with no permissions, checkout or code, and labeler.yml runs on its workflow_run from the default branch. The labeler finds the pull request by its head SHA (checked to be hex) among the open ones and labels it with actions/labeler, then adds the same family/capability/contract/substrate labels to the issues the pull request closes or refers to, extracting only issue numbers and checking each is an issue. A daily run labels any open pull request still without a family label. Generated-by: Claude Opus 5 * ci(labeler): let only project members' or merged pull requests label issues A security review of the linked-issue step: the pull request's body chooses which issues get labels, so anyone opening a pull request could point the workflow's token at any issue. Labels are now passed on immediately only when the author is an OWNER, MEMBER or COLLABORATOR; an outside contributor's pull request passes them on once it is merged, which the doorbell now signals (`closed`), and the labeler finds the merged pull request through the commit's associated pull requests. Generated-by: Claude Opus 5 * ci(labeler): trust a PR body only from members, and check every label has a rule From a second security review of the linked-issue step: a PR's author can edit its body at any time, even after the merge, so "merged" did not make the body trustworthy, and the body was read at run time rather than at merge. The body is now read only for an OWNER, MEMBER or COLLABORATOR author. For anyone else it is never read: a merged PR labels only the issues whose recorded closer (the issue timeline's ClosedEvent) is that PR, which nobody can edit afterwards. A new check-labeler-coverage hook (generate-labeler-config.py --check-coverage) fails when a label docs/labels-and-capabilities.md defines has no labeler rule, unless UNMAPPED lists it with a reason, or when a rule names an undefined label. A label nobody can apply automatically is how pull requests ended up unlabelled. Generated-by: Claude Opus 5 * ci(labeler): count only explicit references when passing labels to issues (#1528) The first run of the new labeler (#1527) labelled #1173, #1347 and #1370, which #1527's description mentions only as test data: any #N in a project member's PR body counted as "refers to". Issues are now taken from GitHub's closing references plus those introduced with a reference phrase ("Part of #N", "Refs #N", "Related to #N", "Relates to #N", "Follow-up to #N", with #N or this repository's issue URL). A passing #N is not a reference. Outside contributors' PRs are unchanged: they label only the issues their merge closed. Generated-by: Claude Opus 5 * perf(contributor-growth): trim nomination body budget (#1489) * perf(contributor-growth): trim nomination body budget * perf(contributor-growth): keep the gaps and concerns in the nomination assessment The trim dropped two clauses from Step 4 that no companion file carries: the GitHub-breadth line no longer asked the brief to name areas that are thin or absent, only those with signal, and the community-interaction line lost "behaviour under feedback" and "any concerns". Gaps matter to a PMC weighing a nomination, so restore both clauses and re-stamp measured_tokens. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * fix(bitbucket): report pull request state and source commit in Cloud pr status (#1526) On Bitbucket Cloud, `pr status` fetched only the pull request's /statuses endpoint. The normalizer reads the state from the pull request and the head commit from a `commit` field, so every Cloud run reported "state": "unknown" and "commit": null. Cloud get_pull_request_status() now fetches the pull request first and returns it under `pull_request`, with the source commit hash under `commit`, matching the Data Center payload. Build checks are still read from /statuses with pagination. test_cli_pr_status_cloud now fakes the HTTP transport with a realistic Cloud pull request and statuses page, covering the OPEN, MERGED and DECLINED states. Before, it mocked get_pull_request_status() with a shape the Cloud backend never returned. Closes #1495 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(skill-evals): give template-less eval steps a neutral user prompt (#1524) The runner's default user prompt, used by every step without a user-prompt-template.md, was the security-issue-import Step 2a template. It framed each case as an incoming report checked against a tracker corpus and a reporter roster, and asked the model to "apply the semantic sweep and reporter-identity check". 33 other steps (the release-* steps, reviewer-routing and non-asf-profile-smoke) received that framing, with an empty corpus and a "(none)" roster, next to a system prompt for an unrelated task. The default is now the case report followed by "Return JSON only.". security-issue-import/step-2a-semantic-sweep, the step the old default was written for, gets its own user-prompt-template.md with the old text, so its rendered prompt is unchanged apart from the SPDX comment that every template file carries. Closes #1492 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(setup-preflight): read the local lock under its own keys (#1523) The pre-flight parsed `.apache-magpie.local.lock` with the committed lock's parser, which accepts only `method`, `url`, `min_version`, `ref`, `commit` and `source`. The local lock that `install.md` and `upgrade.md` tell the agent to write uses the keys `locks.md` documents for it: `source_method`, `source_url`, `source_ref`, `fetched_commit` and `fetched_at`. Every snapshot install (git-branch, git-tag, svn-zip) therefore got `snapshot-unreadable` and stopped at `step-2`. `lockfile.parse_local` reads the local lock with that key set and still rejects unknown keys. The drift check compares each committed key with its local counterpart (`method`/`source_method`, `url`/`source_url`, `ref`/`source_ref`, `commit`/`fetched_commit`), as `upgrade.md` Step 1 does. Finding codes, facts keys and sections are unchanged, and so is the committed-lock parser. The tests wrote the local lock with the committed lock's keys, which hid the bug; they now write the documented format. The adoption-and-setup spec names the local-lock keys the drift check reads. Closes #1491 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(validator): check skill files reached through skills/ symlinks (#1522) * fix(validator): check skill files reached through skills/ symlinks Every skills/<name> entry is a symlink into plugins/magpie-<family>/skills/<alias>. Path.rglob() does not descend into symlinked directories before Python 3.13, so collect_files_to_check() returned only skills/.pytest_cache/README.md and the per-file checks in run_validation() skipped every skill. check-placeholders.sh had the same gap: grep -r skips symlinks it meets while recursing. collect_files_to_check() now walks skills/ with glob's "**", which follows the symlinks and skips dot-entries. Paths stay under skills/<name>/ and each real file is returned once. check-placeholders.sh scans with grep -R. Checking the skills again surfaced two HARD violations, fixed here: - pr-triage/backport-check.md linked an inline <a id="backports"> in the pr-management config template, which the validator's anchor check does not recognise. The link now targets the "Workflow choices" section that holds the backport_branches row, and the anchor, which had no other reference, is removed. - security-tracker-stats-dashboard/SKILL.md reads tracker issue titles and bodies but had no injection-guard callout. It now carries one. check-placeholders.sh finds no hardcoded references in the skill files it now scans. Signed-off-by: Davide Polato <dpol1@apache.org> * fix(validator): match pre-PR review delegation by the skills/ name PRE_PR_REVIEW_DELEGATED is keyed by the skills/<name> entry (security-model-prepare), but validate_pre_pr_review_block() iterated the resolved plugin directories, whose names are the plugin aliases (model-prepare). The delegated-skill entry never matched, so a delegating skill that lost its pre-PR review block would not be reported. The check now iterates the skills/<name> entries. Signed-off-by: Davide Polato <dpol1@apache.org> --------- Signed-off-by: Davide Polato <dpol1@apache.org> * docs(agents): open GitHub pages for the user with gh browse (#1529) The sandbox blocks macOS `open`, but `gh` already runs outside it and `gh browse` is allowed, so it opens a PR, issue, file or commit page with no prompt and no new sandbox exclusion. Generated-by: Claude Opus 5 * fix(pr-triage): check every --add-label value in the mark-ready guard (#1525) * fix(pr-triage): check every --add-label value in the mark-ready guard The mark-ready guard read the label with ctx.opt(), which returns only the first value of a flag. gh accepts --add-label more than once and parses each value as a CSV list, so these commands added the ready label without the Golden rule 1b check for runs awaiting approval: gh pr edit 5 --add-label triaged --add-label "ready for maintainer review" gh pr edit 5 --add-label "triaged,ready for maintainer review" gh pr edit 5 --add-label 'triaged,"ready for maintainer review"' Add GuardContext.opts(), which returns every value of a repeated flag in both the `--flag value` and `--flag=value` forms, and document it next to opt() in the agent-guard README. A token taken as a value is still scanned as a flag, so `--body --add-label --add-label X`, where gh reads the first --add-label as the body, still yields X. opt() now returns the first of these values; its result is unchanged. The guard drops CSV double quotes, splits each --add-label value on commas, and runs the check when any entry matches the ready label (trimmed, case-insensitive). Its fail-open paths are unchanged. Closes #1493 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(agent-guard): match gh:<group> triggers past global gh flags command_kinds() tagged a gh segment with argv[1], so `gh -R o/r pr edit` was tagged `gh:-R` and a contributed guard declaring TRIGGERS = ["gh:pr"] never ran for it. Resolve the group with gh_subcommand(), which skips global flags and their values, the way the git branch already uses git_subcommand_index(). When no group resolves (for example a bare `gh status`), the tag falls back to argv[1] as before. No shipped guard triggers on a gh:<group> tag today. Refs #1493 Signed-off-by: Davide Polato <dpol1@apache.org> --------- Signed-off-by: Davide Polato <dpol1@apache.org> * feat(pr-management-triage): opt-in pre-filter using typed_decision.choice() (#1403) * feat(cve-tool-vulnogram): get Vulnogram tokens through browser approval and allocate CVEs through the API (#1388) Generated-by: Claude Opus 5 * fix(dev): follow skills/ symlinks in check-placeholders on BSD grep too (#1531) #1522 switched the scan to `grep -R` so it follows the `skills/<name>` symlinks into `plugins/`. That holds for GNU grep, but BSD grep (the one macOS ships) only follows symlinks under `-R` when `-S` is also given, and GNU grep has no `-S`. On macOS the check therefore still skipped every skill, and the new test_reports_forbidden_pattern_in_symlinked_skill failed in the workspace pytest hook, so every local commit on macOS was rejected. Build the file list once with `find -L`, which follows the links on both, and grep that list with `-H` so each match keeps its `skills/<name>/...` path. Generated-by: Claude Opus 5 * fix(bitbucket): harden the cloud merge pin, timeout and status reporting Maintainer fixup on top of the merge work: - Require `--expected-source-commit` to be 7-40 hex characters and check it before any request, so a one-character prefix cannot satisfy the pin by accident. - A timeout on the merge POST now says the outcome is unknown and points at `pr get <id>`, instead of a plain connection error that invites a retry while the merge may already be running. - `merge_status` keeps a fixed vocabulary (merged / submitted / failed); Bitbucket's task state is reported separately as `task_status`, and the Bitbucket strategy actually sent as `backend_strategy`. - Tests for the weak pin, a PR without a source commit hash, the timeout message, pass-through of other errors, and the reported strategy; the README row and the adapters spec describe the pin and the caller-run merge checks. Generated-by: Claude Opus 5 --------- Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Signed-off-by: Davide Polato <dpol1@apache.org> Co-authored-by: Jarek Potiuk <jarek@potiuk.com> Co-authored-by: Andrea Cosentino <ancosen@gmail.com> Co-authored-by: Vardhman Gupta <112063624+Kaap10@users.noreply.github.com> Co-authored-by: Shahar Epstein <60007259+shahar1@users.noreply.github.com> Co-authored-by: Jarek Potiuk <potiuk@apache.org> Co-authored-by: kuse <3133746534@qq.com> Co-authored-by: Davide Polato <dpol1@apache.org> Co-authored-by: Arnav <imarnavpurohit@gmail.com>
#1505) * feat(tools): add asf-nexus and wire Nexus staging check into verify-rc Check 4 agreed on in #1173 had no enforcement path: release-verify-rc verified the jars and POMs staged locally (Step 6b) but never the Nexus staging repository the jars actually resolve from - the surface a JVM [VOTE] is really about, whose promotion to Maven Central is irreversible. Add tools/asf-nexus, a doc-only, read-only adapter (the shape every network contract adapter in this repo uses): the endpoint contract, the recipes and the classification rules, with a new contract:release-staging capability. The service splits its reads along the line that matters (probed against the live service): /content/repositories/<id>/ is anonymous-readable - existence, inventory, .asc and checksum coverage - while /service/local/staging/ needs ASF Nexus credentials and answers the authoritative open/closed state plus the profile-wide listing that surfaces stale repositories from earlier RCs. A voter without credentials runs the anonymous path and reports STATE-UNVERIFIED, never a failure of a correct RC. Wire it in as release-verify-rc Step 6c (lettered to preserve every existing cross-reference to Steps 7-9), gated on ASF + JVM-only + a resolvable staging-repo id (nexus_staging_repo in release-build.md, then the planning issue body - Nexus assigns the id at deploy time and it cannot be predicted). Hard findings: repository not reachable, open (mutable, not a valid vote target - distinct from missing), snapshots repository targeted, coordinates/version mismatch, missing .asc, incomplete companion set. WARN: STATE-UNVERIFIED, stale siblings. The step is read-only by construction: GETs only, never close/drop/promote. Sync the capability taxonomy and validator constants (the capability-sync check), the egress surfaces table, the release-build template, the release-management spec and spec-loop spec, and the vendor-neutrality generated block (contract:release-staging is single-org like project-metadata: repository.apache.org is ASF infrastructure). Add a step-6c eval suite (4 cases) for the new step behaviour and restamp the skill. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): apply adversarial-review fixes to the asf-nexus staging check Address the review on #1505: - The sandbox allowlist is exact-hosts only, so the ".apache.org suffix" claim was wrong. Add repository.apache.org to .claude/settings.json, tools/sandbox-lint/expected.json and docs/setup/secure-agent-setup.md, correct the claim in operations.md, and report a sandbox network refusal as STATE-UNVERIFIED / not-probed - never "repository not reachable". - Credentials move to a netrc-format file read with curl --netrc-file, so the password never appears in argv; the authenticated recipes are stated as RM-pasted-in-their-own-terminal (~/.config is denied to the sandboxed agent by design), and the agent's own run takes the anonymous path. - The not-reachable rule is now stated identically (hard FAIL, factual wording) in operations.md recipe 1, staging-verification.md, the Step 6c body and the troubleshooting table. - Step 6c's ASF gate reads the resolved organization (the same chain Step 9's automated-signing gate uses) and carries the feather marker. - The eval suite now covers what its README row claims: 404 not-reachable, snapshots targeted, coordinates/version mismatch, and a non-JVM SKIP (4 new cases, 8 total); the reference recipes carry the inventory-crawl line. - Recipe 4's crawl() is rewritten: the live service emits absolute hrefs (verified; a real line is quoted in operations.md), so the old relative-only grep returned nothing, and the crawl now recurses. - Rebased on #1510-#1515: Step 6c moved into the jvm-artefacts.md sibling the conditional-step loader reads, its eval step-config follows, and the skill is restamped. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * chore(docs): regenerate the release-management config table Step 6c's ASF gate reads the resolved organization from <project-config>/project.md, so the generated family config table gains the project.md row the generator derives for verify-rc. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): tighten the asf-nexus staging check per review Address the second review on #1505: - Push the two allowlist edits the previous round claimed but lost: repository.apache.org is now an exact-host entry next to dist.apache.org in .claude/settings.json and tools/sandbox-lint/expected.json, matching the docs block (which also moves next to dist.apache.org so all three lists are identical). The operations.md egress note drops the "in this PR" phrasing. - The not-reachable rule is now consistent in the promoted-away-ids paragraph too: a FAIL worded "repository not reachable at the id given for this RC". - Step 6c moves out of jvm-artefacts.md into its own nexus-staging.md sibling with an H1 - an H2 under the 6b H1 leaked the whole 6c section into the 6b eval prompt and carried a second JSON contract. SKILL.md gains a Step 6c pointer section, the 6c eval step-config follows the new file, and the 6b suite is unaffected. - Step 10's verdict now names Step 6c (model-classified, so --status nexus-staging=<status>), and tools/release-verify's STEP_ORDER gains nexus-staging after jvm-artefacts. - The crawl() is rewritten to normalise every href first (relative hrefs prefixed with the directory being read) and apply the base-tree guard to files and directories alike - the previous version skipped relative directory links, garbled relative file links and let the page-head favicon/stylesheet through. Verified against a stubbed curl serving a mixed listing. - The resolved staging-repo id is validated against the Nexus shape before it reaches a curl URL (it can come from the planning issue body and the agent runs the probes itself); anything else is a SKIP naming the bad value - pinned by a new eval case (9 total). - staging-verification.md's gate wording aligns with the resolved organization, and the case-1/case-3 fixtures name the netrc path. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(infra): actually add repository.apache.org to the two JSON allowlists The previous round's commit message and reply claimed all three allowlists changed, but only the docs block did - the JSON edits were silently lost to a non-matching replace pattern (the lists are multi-line, one host per line). Add the exact host next to dist.apache.org in both, so all three lists are identical. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * chore(setup): restamp the isolated-setup fingerprint The secure-setup docs' allowedDomains block gained repository.apache.org, so the framework fingerprint the isolated-setup preflight ships with changes; restamp with the value the hook computes on CI. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * chore: retrigger CI The previous run's pytest (plugins) job failed on a timing flake (test_reviewers_run_in_parallel_and_keep_order asserts two stubbed reviewers finish in under 2.8s; the run measured 68s of runner contention — nothing in this branch touches that package). Generated-by: ZCode (GLM-5.3-Flash) * chore(setup): ship the fingerprint value CI computes The previous commit captured the value the hook computes locally on Windows (249e84809cb53ff0); CI computes ca2acfba243af609. Ship the CI value - the fingerprint input includes platform paths, so only the Linux computation is authoritative. Generated-by: ZCode (GLM-5.3-Flash) * fix(agent-guard): re-exec under Python 3.11+ when python3 is older (#1507) * fix(agent-guard): re-exec under Python 3.11+ when python3 is older Hooks invoke the guard engine as a bare `python3`, which resolves through the user's PATH. With an activated project virtualenv on Python 3.10 (a common adopter setup, e.g. Apache Airflow) the module-level `import tomllib` raised ModuleNotFoundError on every Bash call: a traceback in the UI each time, and the guard silently never ran. The engine now imports on 3.10 (tomllib is imported where it is used) and, when the interpreter is older than 3.11, re-runs itself under the newest `python3.N` (3.11+) on PATH. When none exists it exits 1 with one actionable line instead of a traceback. Every harness adapter benefits, since the check runs before `cli()` dispatches. Generated-by: Claude Code (Fable 5.1) * fix(agent-guard): clear the re-exec marker once on 3.11+ The marker stayed in the environment after the re-exec succeeded, so a guard run nested under `--exec` inherited it, skipped the interpreter search and exited with a false "no python3.11+ is on PATH". Drop it once the supported interpreter is running, give the already-re-exec'd case its own message, and replace the unknown comment tag. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * chore(vetted-ops): grant release-rc-cut its two operations in Magpie's policy (#1520) #1518 routed release-rc-cut's GitHub calls through vetted operations. Magpie self-adopts the framework, so its own policy needs the caller: "release-rc-cut" = ["tags", "repo-issue-comment"]. `tags` is a read; `repo-issue-comment` writes, so it runs through `vetted-op` and asks every time. Generated-by: Claude Opus 5 * chore(asf.yaml): require review threads to be resolved before merge (#1521) With the approval requirement lifted on main, an unresolved review thread is the only remaining signal that a reviewer's point is still open, and nothing stopped a PR from merging past it. Turn required_conversation_resolution back on so every thread is answered (fixed by the author, or resolved by the reviewer when a nit is left as-is) before merge. The bootstrap-phase note above it already says threads must be resolved; this makes that true again. Generated-by: Claude Opus 5 * feat(tools): add informational JVM checks 5-7 to maven-artifact-verify (#1506) * feat(tools): add informational JVM checks 5-7 to maven-artifact-verify The informational checks agreed on in #1173 (checks 5-7) close the issue's plan: cheap signals a reviewer currently derives by hand, deliberately never gates. Extend maven-artifact-verify with an `observations` section that never changes `status`: - Check 5: whether every file entry of a main jar shares one timestamp - consistent / not consistent with a reproducible configuration (project.build.outputTimestamp), never asserted as "reproducible"; empty or single-entry jars report INSUFFICIENT-DATA. Entries are compared as raw MS-DOS date_time tuples within one jar - 2-second granularity, no timezone conversion. - Check 6: whether the declared groupId sits under org.apache.* (informational even for ASF top-level projects - published coordinates cannot be renamed retroactively), and the proportion of class entries under the package path derived from the groupId plus the package roots actually found - a proportion and a list, never a boolean. META-INF/, module-info.class and multi-release overrides are excluded as legitimate divergences. - Check 7: whether -sources.jar carries .java/.scala/.kt sources and no .class files, and whether -javadoc.jar is non-empty. Placeholder companions are the Maven-Central-sanctioned pattern, reported as such, never failed; no Javadoc-specific structure is asserted (dokka/scaladoc output is equally valid). Opening a jar reads the zip central directory only (entry names and timestamps); no entry content is extracted. Surface the observations in release-verify-rc Step 6b's JSON contract (`observations`, graded as prose, never affecting the verdict), add two eval cases (observations-never-fail, namespace outside org.apache.*), sync the spec and spec-loop spec, and restamp the skill. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): link the asf-nexus reference by PR number until it lands The relative README link pointed at tools/asf-nexus, which does not exist on this branch yet (it ships with #1505); lychee correctly flagged it as a dead link. Reference the adapter as plain text with its PR number, and restore the relative link on the rebase after #1505 merges. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): keep a damaged jar from crashing the informational checks zipfile.BadZipFile escaped all three observation opens, so a zero-byte or truncated jar with a matching signature and checksum - which passes check 3 - aborted the whole run with a traceback and no JSON, taking the blocking report down with it. Each open now degrades to an unreadable observation, the aggregation comment says what actually keeps the observations out of the verdict, and the docs say insufficient-data in the case the tool emits. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): widen the observation guards and document unreadable The observations must never take the run down, but zipfile can escape with more than BadZipFile and OSError while parsing a damaged central directory: UnicodeDecodeError (real, reproduced - an entry name with the UTF-8 flag set over invalid bytes), plus NotImplementedError and the rest of ValueError. All three opens now catch the wider set and degrade to an unreadable observation. The parametrised damaged-jar test covers three variants: not-a-zip (BadZipFile), invalid-UTF-8-name-with-flag (UnicodeDecodeError), and the patched high version-needed bytes. Verified empirically: CPython does not validate that field at central-directory parse time, so that variant does not raise - the case pins that the report is emitted unchanged either way. The unreadable signal is documented where the RM meets it (tool README, jvm-artefacts.md, step-6b output-spec), and the asf-nexus / Step 6c references in the docstring and README are rephrased as pending (landing via #1505), since neither exists on main yet. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * chore(skills): re-apply the observations text on the reflowed sibling #1517 re-wrapped jvm-artefacts.md; re-apply the observations section, the observations field of the JSON contract and the asf-nexus pointer sentence on the new line breaks, with the unreadable signal documented. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * test(maven-artifact-verify): patch the zip version byte to 12.9 The high-version case wrote 0x0C09 little-endian over the central-directory "version needed to extract" field, but that field is one byte; the result was version 0.9, which is valid, so the case never raised and asserted insufficient-data. Write 129 (12.9) instead: zipfile then raises NotImplementedError while parsing the central directory, the widened catch turns it into an unreadable observation, and the case now asserts that like the other damaged variants. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * perf(contributor-growth): trim contributor-to-committer body budget (#1487) * feat(tools/forgejo): add Forgejo/Gitea adapter bridge (part of #310) (#1469) * feat(tools/forgejo): add Forgejo/Gitea adapter bridge (part of #310) * docs(tools/forgejo): drop the @me assignee form and note JSON escaping `tea issues edit --add-assignees` (0.15.1) takes a comma-separated list of usernames and does not resolve `@me`, so the recipe would assign a literal "@me". Keep only the `<handle>` form. The body-edit and PR-create Write-tool payloads carry multi-line text, so say they must be properly escaped JSON, as the issue-create and comment recipes already do. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * ci(labeler): label PRs on workflow_run, from skills too, and pass labels to linked issues (#1527) * ci(labeler): label pull requests on workflow_run, from skills too, and pass labels to linked issues Many recent pull requests carried no labels. Four causes: - .github/labeler.yml only mapped tool directories to contract:* / substrate:*, so a change to skills, docs or workflows matched nothing, and family:* / skill capability:* were never applied automatically. - changed-files-labels-limit was 8, and actions/labeler applies no changed-files label at all once more than that match: a cliff, not a cap. - The hourly scheduled run labelled each pull request once, so a later push into another area was never relabelled. - Labels arrived up to an hour late. The generator now emits, from the repository's own declarations: - family:* and capability:* from every skill's frontmatter, on the skill's directory and its eval suite (found through the skills/ symlink); - the non-skill families: family:tools (tools/ without skill evals and specs, and tool-only plugins), family:ci (.github/, the pre-commit config, tools/dev/, root tooling files), family:docs (docs/, READMEs, root *.md), and family:setup (Magpie's own overrides and pin); - only labels docs/labels-and-capabilities.md defines. The limit goes to 20. The workflow follows magpie-site's privilege split: labeler-signal.yml is an unprivileged pull_request doorbell with no permissions, checkout or code, and labeler.yml runs on its workflow_run from the default branch. The labeler finds the pull request by its head SHA (checked to be hex) among the open ones and labels it with actions/labeler, then adds the same family/capability/contract/substrate labels to the issues the pull request closes or refers to, extracting only issue numbers and checking each is an issue. A daily run labels any open pull request still without a family label. Generated-by: Claude Opus 5 * ci(labeler): let only project members' or merged pull requests label issues A security review of the linked-issue step: the pull request's body chooses which issues get labels, so anyone opening a pull request could point the workflow's token at any issue. Labels are now passed on immediately only when the author is an OWNER, MEMBER or COLLABORATOR; an outside contributor's pull request passes them on once it is merged, which the doorbell now signals (`closed`), and the labeler finds the merged pull request through the commit's associated pull requests. Generated-by: Claude Opus 5 * ci(labeler): trust a PR body only from members, and check every label has a rule From a second security review of the linked-issue step: a PR's author can edit its body at any time, even after the merge, so "merged" did not make the body trustworthy, and the body was read at run time rather than at merge. The body is now read only for an OWNER, MEMBER or COLLABORATOR author. For anyone else it is never read: a merged PR labels only the issues whose recorded closer (the issue timeline's ClosedEvent) is that PR, which nobody can edit afterwards. A new check-labeler-coverage hook (generate-labeler-config.py --check-coverage) fails when a label docs/labels-and-capabilities.md defines has no labeler rule, unless UNMAPPED lists it with a reason, or when a rule names an undefined label. A label nobody can apply automatically is how pull requests ended up unlabelled. Generated-by: Claude Opus 5 * ci(labeler): count only explicit references when passing labels to issues (#1528) The first run of the new labeler (#1527) labelled #1173, #1347 and #1370, which #1527's description mentions only as test data: any #N in a project member's PR body counted as "refers to". Issues are now taken from GitHub's closing references plus those introduced with a reference phrase ("Part of #N", "Refs #N", "Related to #N", "Relates to #N", "Follow-up to #N", with #N or this repository's issue URL). A passing #N is not a reference. Outside contributors' PRs are unchanged: they label only the issues their merge closed. Generated-by: Claude Opus 5 * perf(contributor-growth): trim nomination body budget (#1489) * perf(contributor-growth): trim nomination body budget * perf(contributor-growth): keep the gaps and concerns in the nomination assessment The trim dropped two clauses from Step 4 that no companion file carries: the GitHub-breadth line no longer asked the brief to name areas that are thin or absent, only those with signal, and the community-interaction line lost "behaviour under feedback" and "any concerns". Gaps matter to a PMC weighing a nomination, so restore both clauses and re-stamp measured_tokens. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * fix(bitbucket): report pull request state and source commit in Cloud pr status (#1526) On Bitbucket Cloud, `pr status` fetched only the pull request's /statuses endpoint. The normalizer reads the state from the pull request and the head commit from a `commit` field, so every Cloud run reported "state": "unknown" and "commit": null. Cloud get_pull_request_status() now fetches the pull request first and returns it under `pull_request`, with the source commit hash under `commit`, matching the Data Center payload. Build checks are still read from /statuses with pagination. test_cli_pr_status_cloud now fakes the HTTP transport with a realistic Cloud pull request and statuses page, covering the OPEN, MERGED and DECLINED states. Before, it mocked get_pull_request_status() with a shape the Cloud backend never returned. Closes #1495 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(skill-evals): give template-less eval steps a neutral user prompt (#1524) The runner's default user prompt, used by every step without a user-prompt-template.md, was the security-issue-import Step 2a template. It framed each case as an incoming report checked against a tracker corpus and a reporter roster, and asked the model to "apply the semantic sweep and reporter-identity check". 33 other steps (the release-* steps, reviewer-routing and non-asf-profile-smoke) received that framing, with an empty corpus and a "(none)" roster, next to a system prompt for an unrelated task. The default is now the case report followed by "Return JSON only.". security-issue-import/step-2a-semantic-sweep, the step the old default was written for, gets its own user-prompt-template.md with the old text, so its rendered prompt is unchanged apart from the SPDX comment that every template file carries. Closes #1492 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(setup-preflight): read the local lock under its own keys (#1523) The pre-flight parsed `.apache-magpie.local.lock` with the committed lock's parser, which accepts only `method`, `url`, `min_version`, `ref`, `commit` and `source`. The local lock that `install.md` and `upgrade.md` tell the agent to write uses the keys `locks.md` documents for it: `source_method`, `source_url`, `source_ref`, `fetched_commit` and `fetched_at`. Every snapshot install (git-branch, git-tag, svn-zip) therefore got `snapshot-unreadable` and stopped at `step-2`. `lockfile.parse_local` reads the local lock with that key set and still rejects unknown keys. The drift check compares each committed key with its local counterpart (`method`/`source_method`, `url`/`source_url`, `ref`/`source_ref`, `commit`/`fetched_commit`), as `upgrade.md` Step 1 does. Finding codes, facts keys and sections are unchanged, and so is the committed-lock parser. The tests wrote the local lock with the committed lock's keys, which hid the bug; they now write the documented format. The adoption-and-setup spec names the local-lock keys the drift check reads. Closes #1491 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(validator): check skill files reached through skills/ symlinks (#1522) * fix(validator): check skill files reached through skills/ symlinks Every skills/<name> entry is a symlink into plugins/magpie-<family>/skills/<alias>. Path.rglob() does not descend into symlinked directories before Python 3.13, so collect_files_to_check() returned only skills/.pytest_cache/README.md and the per-file checks in run_validation() skipped every skill. check-placeholders.sh had the same gap: grep -r skips symlinks it meets while recursing. collect_files_to_check() now walks skills/ with glob's "**", which follows the symlinks and skips dot-entries. Paths stay under skills/<name>/ and each real file is returned once. check-placeholders.sh scans with grep -R. Checking the skills again surfaced two HARD violations, fixed here: - pr-triage/backport-check.md linked an inline <a id="backports"> in the pr-management config template, which the validator's anchor check does not recognise. The link now targets the "Workflow choices" section that holds the backport_branches row, and the anchor, which had no other reference, is removed. - security-tracker-stats-dashboard/SKILL.md reads tracker issue titles and bodies but had no injection-guard callout. It now carries one. check-placeholders.sh finds no hardcoded references in the skill files it now scans. Signed-off-by: Davide Polato <dpol1@apache.org> * fix(validator): match pre-PR review delegation by the skills/ name PRE_PR_REVIEW_DELEGATED is keyed by the skills/<name> entry (security-model-prepare), but validate_pre_pr_review_block() iterated the resolved plugin directories, whose names are the plugin aliases (model-prepare). The delegated-skill entry never matched, so a delegating skill that lost its pre-PR review block would not be reported. The check now iterates the skills/<name> entries. Signed-off-by: Davide Polato <dpol1@apache.org> --------- Signed-off-by: Davide Polato <dpol1@apache.org> * docs(agents): open GitHub pages for the user with gh browse (#1529) The sandbox blocks macOS `open`, but `gh` already runs outside it and `gh browse` is allowed, so it opens a PR, issue, file or commit page with no prompt and no new sandbox exclusion. Generated-by: Claude Opus 5 * fix(pr-triage): check every --add-label value in the mark-ready guard (#1525) * fix(pr-triage): check every --add-label value in the mark-ready guard The mark-ready guard read the label with ctx.opt(), which returns only the first value of a flag. gh accepts --add-label more than once and parses each value as a CSV list, so these commands added the ready label without the Golden rule 1b check for runs awaiting approval: gh pr edit 5 --add-label triaged --add-label "ready for maintainer review" gh pr edit 5 --add-label "triaged,ready for maintainer review" gh pr edit 5 --add-label 'triaged,"ready for maintainer review"' Add GuardContext.opts(), which returns every value of a repeated flag in both the `--flag value` and `--flag=value` forms, and document it next to opt() in the agent-guard README. A token taken as a value is still scanned as a flag, so `--body --add-label --add-label X`, where gh reads the first --add-label as the body, still yields X. opt() now returns the first of these values; its result is unchanged. The guard drops CSV double quotes, splits each --add-label value on commas, and runs the check when any entry matches the ready label (trimmed, case-insensitive). Its fail-open paths are unchanged. Closes #1493 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(agent-guard): match gh:<group> triggers past global gh flags command_kinds() tagged a gh segment with argv[1], so `gh -R o/r pr edit` was tagged `gh:-R` and a contributed guard declaring TRIGGERS = ["gh:pr"] never ran for it. Resolve the group with gh_subcommand(), which skips global flags and their values, the way the git branch already uses git_subcommand_index(). When no group resolves (for example a bare `gh status`), the tag falls back to argv[1] as before. No shipped guard triggers on a gh:<group> tag today. Refs #1493 Signed-off-by: Davide Polato <dpol1@apache.org> --------- Signed-off-by: Davide Polato <dpol1@apache.org> * feat(pr-management-triage): opt-in pre-filter using typed_decision.choice() (#1403) * feat(cve-tool-vulnogram): get Vulnogram tokens through browser approval and allocate CVEs through the API (#1388) Generated-by: Claude Opus 5 * fix(dev): follow skills/ symlinks in check-placeholders on BSD grep too (#1531) #1522 switched the scan to `grep -R` so it follows the `skills/<name>` symlinks into `plugins/`. That holds for GNU grep, but BSD grep (the one macOS ships) only follows symlinks under `-R` when `-S` is also given, and GNU grep has no `-S`. On macOS the check therefore still skipped every skill, and the new test_reports_forbidden_pattern_in_symlinked_skill failed in the workspace pytest hook, so every local commit on macOS was rejected. Build the file list once with `find -L`, which follows the links on both, and grep that list with `-H` so each match keeps its `skills/<name>/...` path. Generated-by: Claude Opus 5 * fix(asf-nexus): resolve root-relative hrefs and always gate Step 6c in its own file Maintainer fixup on top of the review round: - crawl(): resolve a root-relative href (`/favicon.ico`, `/nexus/style.css`) against the host, so the `"$base"/*` guard drops it instead of it landing in the inventory as a staged file. Verified against a stubbed curl serving absolute, relative, root-relative and `../` links: the inventory is exactly the repository's files. - verify-rc SKILL.md: load nexus-staging.md whenever Step 6b ran and let its own gates (organization, resolvable id, id shape) report the explicit SKIP; when Step 6b did not run, report Step 6c as SKIP without loading it. Before, the file only loaded once an id resolved, so the "no id" SKIP it promises could never be emitted. - operations.md: drop the paragraph that duplicated "Collect every path…", one sentence per line in the new prose. - jvm-artefacts.md: link Step 6c (`nexus-staging.md`) directly instead of "landing via #1505", which goes stale on merge. Generated-by: Claude Opus 5 * feat(bitbucket): add guarded cloud PR merge (#1471) * feat(bitbucket): add guarded cloud PR merge * fix(bitbucket): align cloud merge with land contract * fix(bitbucket): harden cloud PR merge * fix(release-config): initialise skill before parsing it (#1514) CodeQL (py/uninitialized-local-variable) could not see that parser.error() exits, so it read `skill` as possibly unset on the error path. Initialise it first; behaviour is unchanged. Generated-by: Claude Opus 5 * feat(tools/mail-source): add Mailman 3 / Hyperkitty archive backend (#1474) * feat(tools/mail-source): add Mailman 3 / Hyperkitty archive backend Projects on Mailman 3 (Python, Fedora, GNU and many others) had no mail-source backend besides Gmail. Hyperkitty, the Mailman 3 archiver, serves its archive as a JSON API, so the adapter is a README of curl recipes rather than code: list_recent_threads, read_thread and thread_url, keyed by the root Message-ID like the IMAP and mbox adapters. Like PonyMail it only reads. A private archive needs a subscribed session the adapter does not wire, so it declines those and the resolution rule falls through to a subscriber-side backend. The endpoints, paging, thread keys and permission checks follow the Hyperkitty and mailman-web sources, and the Message-ID hash recipe is the computation of Hyperkitty's own get_message_id_hash. The contract's capability matrix and the other lists of mail-source backends now include it, and CONTRIBUTING no longer offers it as open work. Closes #306 Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Generated-by: Claude Code (Opus 5.5) * fix(tools/mail-source): probe a Hyperkitty thread before listing it Review on #1474 found that the read_thread fallback could never fire: thread/<hash>/emails/ is a filtered list, so Hyperkitty answers an unknown thread with 200 and no results instead of a 404. read_thread now fetches thread/<hash>/ first, which does 404, and the email/<hash>/ fallback rejoins at the emails step. The same review noted that a site with Basic authentication first in its API settings refuses anonymous private-list reads with 401 rather than 403, that date_active carries the server's UTC offset and has to be compared as a timezone-aware time, and that secure-setup adopters need their Hyperkitty host in sandbox.network.allowedDomains. The README now covers all three. Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Generated-by: Claude Code (Opus 5.5) --------- Signed-off-by: Andrea Cosentino <ancosen@gmail.com> * perf(release-management): wording pass on the release skills (#1517) The optimize-skill rewrite pass, with the style rules the maintainer approved on the security family, applied to all ten release skills and their step files: one sentence per line, three-line external-content paragraphs, and hard rules that repeated a golden rule now pointing at it. Headings, code blocks, emitted commands, tool invocations and eval-covered wording are unchanged. Each pass listed every removed sentence that carried a condition, exception or prohibition; each was reviewed and the rule found intact elsewhere. The skills were already lean after the extraction and split, so this saves little: SKILL.md tokens 67,379 -> 65,707 across the family. Fixes made along the way: - release-prepare: the manifest read no longer pipes `gh api` into base64 (it asks for the raw file), the planning issue body goes through a scratch file instead of a /tmp heredoc, and two references to "Step 2f" now name the archive review, Step 2e. - release-vote-draft: the planning-issue comment is posted with --body-file. - release-verify-rc: the Step 5 FAIL example now says there is nothing to diff, as the eval's expected answer does; after the reflow the model copied the shorter example literally and failed that case. - release-rc-cut: a hard rule cited a "Step 0 check 9" that no longer exists; it now points at release-config's reproducibility check. - release-vote-tally, keys-sync, archive-sweep: golden and hard rules now state the rules their scripts enforce (an ambiguous latest vote halts; secp256k1 refused; pre-releases never archived). Generated-by: Claude Opus 5 * perf(contributor-growth): trim activity-sweep skill routing metadata (#1483) * perf(release-management): shorter descriptions for five release skills (#1519) The descriptions every session loads, invoked or not. release-prepare, -verify-rc, -rc-cut, -keys-sync and -announce-draft carried whole paragraphs (long input lists, step numbers, every boundary). They now say what the skill does, its main boundary and its trigger phrases, in the style used for the security family; the detail stays in each body, read when the skill runs. description + when_to_use for the five: ~1,520 -> ~605 tokens. The family's advertised surface (name + description, as docs/setup/marketplace.md measures it) goes ~1.4k -> ~0.8k. Generated-by: Claude Opus 5 * fix(release-rc-cut): route its two GitHub calls through vetted operations (#1518) Golden rule 1 said the skill made no gh call, yet Step 0 read the RC tag with `gh api` and Step 4 posted the planning-issue comment with `gh issue comment`. The maintainer settled it: the skill still never runs a release command locally, and its only GitHub access goes through two existing vetted operations, `tags` (read) and `repo-issue-comment` (write, asks every time, after the RM confirms). The `tags` operation lists every tag under a prefix, so a check for rc1 also returns rc10: the tag exists only when a line is exactly refs/tags/<version>-<rcN>. A new eval case pins that. The vetted-ops README's caller example gains "release-rc-cut" = ["tags", "repo-issue-comment"]; adopters add the same grant to their policy. Without the secure setup the skill names the plain gh equivalents. Generated-by: Claude Opus 5 * fix(agent-guard): re-exec under Python 3.11+ when python3 is older (#1507) * fix(agent-guard): re-exec under Python 3.11+ when python3 is older Hooks invoke the guard engine as a bare `python3`, which resolves through the user's PATH. With an activated project virtualenv on Python 3.10 (a common adopter setup, e.g. Apache Airflow) the module-level `import tomllib` raised ModuleNotFoundError on every Bash call: a traceback in the UI each time, and the guard silently never ran. The engine now imports on 3.10 (tomllib is imported where it is used) and, when the interpreter is older than 3.11, re-runs itself under the newest `python3.N` (3.11+) on PATH. When none exists it exits 1 with one actionable line instead of a traceback. Every harness adapter benefits, since the check runs before `cli()` dispatches. Generated-by: Claude Code (Fable 5.1) * fix(agent-guard): clear the re-exec marker once on 3.11+ The marker stayed in the environment after the re-exec succeeded, so a guard run nested under `--exec` inherited it, skipped the interpreter search and exited with a false "no python3.11+ is on PATH". Drop it once the supported interpreter is running, give the already-re-exec'd case its own message, and replace the unknown comment tag. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * chore(vetted-ops): grant release-rc-cut its two operations in Magpie's policy (#1520) #1518 routed release-rc-cut's GitHub calls through vetted operations. Magpie self-adopts the framework, so its own policy needs the caller: "release-rc-cut" = ["tags", "repo-issue-comment"]. `tags` is a read; `repo-issue-comment` writes, so it runs through `vetted-op` and asks every time. Generated-by: Claude Opus 5 * chore(asf.yaml): require review threads to be resolved before merge (#1521) With the approval requirement lifted on main, an unresolved review thread is the only remaining signal that a reviewer's point is still open, and nothing stopped a PR from merging past it. Turn required_conversation_resolution back on so every thread is answered (fixed by the author, or resolved by the reviewer when a nit is left as-is) before merge. The bootstrap-phase note above it already says threads must be resolved; this makes that true again. Generated-by: Claude Opus 5 * feat(tools): add informational JVM checks 5-7 to maven-artifact-verify (#1506) * feat(tools): add informational JVM checks 5-7 to maven-artifact-verify The informational checks agreed on in #1173 (checks 5-7) close the issue's plan: cheap signals a reviewer currently derives by hand, deliberately never gates. Extend maven-artifact-verify with an `observations` section that never changes `status`: - Check 5: whether every file entry of a main jar shares one timestamp - consistent / not consistent with a reproducible configuration (project.build.outputTimestamp), never asserted as "reproducible"; empty or single-entry jars report INSUFFICIENT-DATA. Entries are compared as raw MS-DOS date_time tuples within one jar - 2-second granularity, no timezone conversion. - Check 6: whether the declared groupId sits under org.apache.* (informational even for ASF top-level projects - published coordinates cannot be renamed retroactively), and the proportion of class entries under the package path derived from the groupId plus the package roots actually found - a proportion and a list, never a boolean. META-INF/, module-info.class and multi-release overrides are excluded as legitimate divergences. - Check 7: whether -sources.jar carries .java/.scala/.kt sources and no .class files, and whether -javadoc.jar is non-empty. Placeholder companions are the Maven-Central-sanctioned pattern, reported as such, never failed; no Javadoc-specific structure is asserted (dokka/scaladoc output is equally valid). Opening a jar reads the zip central directory only (entry names and timestamps); no entry content is extracted. Surface the observations in release-verify-rc Step 6b's JSON contract (`observations`, graded as prose, never affecting the verdict), add two eval cases (observations-never-fail, namespace outside org.apache.*), sync the spec and spec-loop spec, and restamp the skill. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): link the asf-nexus reference by PR number until it lands The relative README link pointed at tools/asf-nexus, which does not exist on this branch yet (it ships with #1505); lychee correctly flagged it as a dead link. Reference the adapter as plain text with its PR number, and restore the relative link on the rebase after #1505 merges. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): keep a damaged jar from crashing the informational checks zipfile.BadZipFile escaped all three observation opens, so a zero-byte or truncated jar with a matching signature and checksum - which passes check 3 - aborted the whole run with a traceback and no JSON, taking the blocking report down with it. Each open now degrades to an unreadable observation, the aggregation comment says what actually keeps the observations out of the verdict, and the docs say insufficient-data in the case the tool emits. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * fix(tools): widen the observation guards and document unreadable The observations must never take the run down, but zipfile can escape with more than BadZipFile and OSError while parsing a damaged central directory: UnicodeDecodeError (real, reproduced - an entry name with the UTF-8 flag set over invalid bytes), plus NotImplementedError and the rest of ValueError. All three opens now catch the wider set and degrade to an unreadable observation. The parametrised damaged-jar test covers three variants: not-a-zip (BadZipFile), invalid-UTF-8-name-with-flag (UnicodeDecodeError), and the patched high version-needed bytes. Verified empirically: CPython does not validate that field at central-directory parse time, so that variant does not raise - the case pins that the report is emitted unchanged either way. The unreadable signal is documented where the RM meets it (tool README, jvm-artefacts.md, step-6b output-spec), and the asf-nexus / Step 6c references in the docstring and README are rephrased as pending (landing via #1505), since neither exists on main yet. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * chore(skills): re-apply the observations text on the reflowed sibling #1517 re-wrapped jvm-artefacts.md; re-apply the observations section, the observations field of the JSON contract and the asf-nexus pointer sentence on the new line breaks, with the unreadable signal documented. Refs #1173 Generated-by: ZCode (GLM-5.3-Flash) * test(maven-artifact-verify): patch the zip version byte to 12.9 The high-version case wrote 0x0C09 little-endian over the central-directory "version needed to extract" field, but that field is one byte; the result was version 0.9, which is valid, so the case never raised and asserted insufficient-data. Write 129 (12.9) instead: zipfile then raises NotImplementedError while parsing the central directory, the widened catch turns it into an unreadable observation, and the case now asserts that like the other damaged variants. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * perf(contributor-growth): trim contributor-to-committer body budget (#1487) * feat(tools/forgejo): add Forgejo/Gitea adapter bridge (part of #310) (#1469) * feat(tools/forgejo): add Forgejo/Gitea adapter bridge (part of #310) * docs(tools/forgejo): drop the @me assignee form and note JSON escaping `tea issues edit --add-assignees` (0.15.1) takes a comma-separated list of usernames and does not resolve `@me`, so the recipe would assign a literal "@me". Keep only the `<handle>` form. The body-edit and PR-create Write-tool payloads carry multi-line text, so say they must be properly escaped JSON, as the issue-create and comment recipes already do. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * ci(labeler): label PRs on workflow_run, from skills too, and pass labels to linked issues (#1527) * ci(labeler): label pull requests on workflow_run, from skills too, and pass labels to linked issues Many recent pull requests carried no labels. Four causes: - .github/labeler.yml only mapped tool directories to contract:* / substrate:*, so a change to skills, docs or workflows matched nothing, and family:* / skill capability:* were never applied automatically. - changed-files-labels-limit was 8, and actions/labeler applies no changed-files label at all once more than that match: a cliff, not a cap. - The hourly scheduled run labelled each pull request once, so a later push into another area was never relabelled. - Labels arrived up to an hour late. The generator now emits, from the repository's own declarations: - family:* and capability:* from every skill's frontmatter, on the skill's directory and its eval suite (found through the skills/ symlink); - the non-skill families: family:tools (tools/ without skill evals and specs, and tool-only plugins), family:ci (.github/, the pre-commit config, tools/dev/, root tooling files), family:docs (docs/, READMEs, root *.md), and family:setup (Magpie's own overrides and pin); - only labels docs/labels-and-capabilities.md defines. The limit goes to 20. The workflow follows magpie-site's privilege split: labeler-signal.yml is an unprivileged pull_request doorbell with no permissions, checkout or code, and labeler.yml runs on its workflow_run from the default branch. The labeler finds the pull request by its head SHA (checked to be hex) among the open ones and labels it with actions/labeler, then adds the same family/capability/contract/substrate labels to the issues the pull request closes or refers to, extracting only issue numbers and checking each is an issue. A daily run labels any open pull request still without a family label. Generated-by: Claude Opus 5 * ci(labeler): let only project members' or merged pull requests label issues A security review of the linked-issue step: the pull request's body chooses which issues get labels, so anyone opening a pull request could point the workflow's token at any issue. Labels are now passed on immediately only when the author is an OWNER, MEMBER or COLLABORATOR; an outside contributor's pull request passes them on once it is merged, which the doorbell now signals (`closed`), and the labeler finds the merged pull request through the commit's associated pull requests. Generated-by: Claude Opus 5 * ci(labeler): trust a PR body only from members, and check every label has a rule From a second security review of the linked-issue step: a PR's author can edit its body at any time, even after the merge, so "merged" did not make the body trustworthy, and the body was read at run time rather than at merge. The body is now read only for an OWNER, MEMBER or COLLABORATOR author. For anyone else it is never read: a merged PR labels only the issues whose recorded closer (the issue timeline's ClosedEvent) is that PR, which nobody can edit afterwards. A new check-labeler-coverage hook (generate-labeler-config.py --check-coverage) fails when a label docs/labels-and-capabilities.md defines has no labeler rule, unless UNMAPPED lists it with a reason, or when a rule names an undefined label. A label nobody can apply automatically is how pull requests ended up unlabelled. Generated-by: Claude Opus 5 * ci(labeler): count only explicit references when passing labels to issues (#1528) The first run of the new labeler (#1527) labelled #1173, #1347 and #1370, which #1527's description mentions only as test data: any #N in a project member's PR body counted as "refers to". Issues are now taken from GitHub's closing references plus those introduced with a reference phrase ("Part of #N", "Refs #N", "Related to #N", "Relates to #N", "Follow-up to #N", with #N or this repository's issue URL). A passing #N is not a reference. Outside contributors' PRs are unchanged: they label only the issues their merge closed. Generated-by: Claude Opus 5 * perf(contributor-growth): trim nomination body budget (#1489) * perf(contributor-growth): trim nomination body budget * perf(contributor-growth): keep the gaps and concerns in the nomination assessment The trim dropped two clauses from Step 4 that no companion file carries: the GitHub-breadth line no longer asked the brief to name areas that are thin or absent, only those with signal, and the community-interaction line lost "behaviour under feedback" and "any concerns". Gaps matter to a PMC weighing a nomination, so restore both clauses and re-stamp measured_tokens. Generated-by: Claude Opus 5 --------- Co-authored-by: Jarek Potiuk <potiuk@apache.org> * fix(bitbucket): report pull request state and source commit in Cloud pr status (#1526) On Bitbucket Cloud, `pr status` fetched only the pull request's /statuses endpoint. The normalizer reads the state from the pull request and the head commit from a `commit` field, so every Cloud run reported "state": "unknown" and "commit": null. Cloud get_pull_request_status() now fetches the pull request first and returns it under `pull_request`, with the source commit hash under `commit`, matching the Data Center payload. Build checks are still read from /statuses with pagination. test_cli_pr_status_cloud now fakes the HTTP transport with a realistic Cloud pull request and statuses page, covering the OPEN, MERGED and DECLINED states. Before, it mocked get_pull_request_status() with a shape the Cloud backend never returned. Closes #1495 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(skill-evals): give template-less eval steps a neutral user prompt (#1524) The runner's default user prompt, used by every step without a user-prompt-template.md, was the security-issue-import Step 2a template. It framed each case as an incoming report checked against a tracker corpus and a reporter roster, and asked the model to "apply the semantic sweep and reporter-identity check". 33 other steps (the release-* steps, reviewer-routing and non-asf-profile-smoke) received that framing, with an empty corpus and a "(none)" roster, next to a system prompt for an unrelated task. The default is now the case report followed by "Return JSON only.". security-issue-import/step-2a-semantic-sweep, the step the old default was written for, gets its own user-prompt-template.md with the old text, so its rendered prompt is unchanged apart from the SPDX comment that every template file carries. Closes #1492 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(setup-preflight): read the local lock under its own keys (#1523) The pre-flight parsed `.apache-magpie.local.lock` with the committed lock's parser, which accepts only `method`, `url`, `min_version`, `ref`, `commit` and `source`. The local lock that `install.md` and `upgrade.md` tell the agent to write uses the keys `locks.md` documents for it: `source_method`, `source_url`, `source_ref`, `fetched_commit` and `fetched_at`. Every snapshot install (git-branch, git-tag, svn-zip) therefore got `snapshot-unreadable` and stopped at `step-2`. `lockfile.parse_local` reads the local lock with that key set and still rejects unknown keys. The drift check compares each committed key with its local counterpart (`method`/`source_method`, `url`/`source_url`, `ref`/`source_ref`, `commit`/`fetched_commit`), as `upgrade.md` Step 1 does. Finding codes, facts keys and sections are unchanged, and so is the committed-lock parser. The tests wrote the local lock with the committed lock's keys, which hid the bug; they now write the documented format. The adoption-and-setup spec names the local-lock keys the drift check reads. Closes #1491 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(validator): check skill files reached through skills/ symlinks (#1522) * fix(validator): check skill files reached through skills/ symlinks Every skills/<name> entry is a symlink into plugins/magpie-<family>/skills/<alias>. Path.rglob() does not descend into symlinked directories before Python 3.13, so collect_files_to_check() returned only skills/.pytest_cache/README.md and the per-file checks in run_validation() skipped every skill. check-placeholders.sh had the same gap: grep -r skips symlinks it meets while recursing. collect_files_to_check() now walks skills/ with glob's "**", which follows the symlinks and skips dot-entries. Paths stay under skills/<name>/ and each real file is returned once. check-placeholders.sh scans with grep -R. Checking the skills again surfaced two HARD violations, fixed here: - pr-triage/backport-check.md linked an inline <a id="backports"> in the pr-management config template, which the validator's anchor check does not recognise. The link now targets the "Workflow choices" section that holds the backport_branches row, and the anchor, which had no other reference, is removed. - security-tracker-stats-dashboard/SKILL.md reads tracker issue titles and bodies but had no injection-guard callout. It now carries one. check-placeholders.sh finds no hardcoded references in the skill files it now scans. Signed-off-by: Davide Polato <dpol1@apache.org> * fix(validator): match pre-PR review delegation by the skills/ name PRE_PR_REVIEW_DELEGATED is keyed by the skills/<name> entry (security-model-prepare), but validate_pre_pr_review_block() iterated the resolved plugin directories, whose names are the plugin aliases (model-prepare). The delegated-skill entry never matched, so a delegating skill that lost its pre-PR review block would not be reported. The check now iterates the skills/<name> entries. Signed-off-by: Davide Polato <dpol1@apache.org> --------- Signed-off-by: Davide Polato <dpol1@apache.org> * docs(agents): open GitHub pages for the user with gh browse (#1529) The sandbox blocks macOS `open`, but `gh` already runs outside it and `gh browse` is allowed, so it opens a PR, issue, file or commit page with no prompt and no new sandbox exclusion. Generated-by: Claude Opus 5 * fix(pr-triage): check every --add-label value in the mark-ready guard (#1525) * fix(pr-triage): check every --add-label value in the mark-ready guard The mark-ready guard read the label with ctx.opt(), which returns only the first value of a flag. gh accepts --add-label more than once and parses each value as a CSV list, so these commands added the ready label without the Golden rule 1b check for runs awaiting approval: gh pr edit 5 --add-label triaged --add-label "ready for maintainer review" gh pr edit 5 --add-label "triaged,ready for maintainer review" gh pr edit 5 --add-label 'triaged,"ready for maintainer review"' Add GuardContext.opts(), which returns every value of a repeated flag in both the `--flag value` and `--flag=value` forms, and document it next to opt() in the agent-guard README. A token taken as a value is still scanned as a flag, so `--body --add-label --add-label X`, where gh reads the first --add-label as the body, still yields X. opt() now returns the first of these values; its result is unchanged. The guard drops CSV double quotes, splits each --add-label value on commas, and runs the check when any entry matches the ready label (trimmed, case-insensitive). Its fail-open paths are unchanged. Closes #1493 Signed-off-by: Davide Polato <dpol1@apache.org> * fix(agent-guard): match gh:<group> triggers past global gh flags command_kinds() tagged a gh segment with argv[1], so `gh -R o/r pr edit` was tagged `gh:-R` and a contributed guard declaring TRIGGERS = ["gh:pr"] never ran for it. Resolve the group with gh_subcommand(), which skips global flags and their values, the way the git branch already uses git_subcommand_index(). When no group resolves (for example a bare `gh status`), the tag falls back to argv[1] as before. No shipped guard triggers on a gh:<group> tag today. Refs #1493 Signed-off-by: Davide Polato <dpol1@apache.org> --------- Signed-off-by: Davide Polato <dpol1@apache.org> * feat(pr-management-triage): opt-in pre-filter using typed_decision.choice() (#1403) * feat(cve-tool-vulnogram): get Vulnogram tokens through browser approval and allocate CVEs through the API (#1388) Generated-by: Claude Opus 5 * fix(dev): follow skills/ symlinks in check-placeholders on BSD grep too (#1531) #1522 switched the scan to `grep -R` so it follows the `skills/<name>` symlinks into `plugins/`. That holds for GNU grep, but BSD grep (the one macOS ships) only follows symlinks under `-R` when `-S` is also given, and GNU grep has no `-S`. On macOS the check therefore still skipped every skill, and the new test_reports_forbidden_pattern_in_symlinked_skill failed in the workspace pytest hook, so every local commit on macOS was rejected. Build the file list once with `find -L`, which follows the links on both, and grep that list with `-H` so each match keeps its `skills/<name>/...` path. Generated-by: Claude Opus 5 * fix(bitbucket): harden the cloud merge pin, timeout and status reporting Maintainer fixup on top of the merge work: - Require `--expected-source-commit` to be 7-40 hex characters and check it before any request, so a one-character prefix cannot satisfy the pin by accident. - A timeout on the merge POST now says the outcome is unknown and points at `pr get <id>`, instead of a plain connection error that invites a retry while the merge may already be running. - `merge_status` keeps a fixed vocabulary (merged / submitted / failed); Bitbucket's task state is reported separately as `task_status`, and the Bitbucket strategy actually sent as `backend_strategy`. - Tests for the weak pin, a PR without a source commit hash, the timeout message, pass-through of other errors, and the reported strategy; the README row and the adapters spec describe the pin and the caller-run merge checks. Generated-by: Claude Opus 5 --------- Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Signed-off-by: Davide Polato <dpol1@apache.org> Co-authored-by: Jarek Potiuk <jarek@potiuk.com> Co-authored-by: Andrea Cosentino <ancosen@gmail.com> Co-authored-by: Vardhman Gupta <112063624+Kaap10@users.noreply.github.com> Co-authored-by: Shahar Epstein <60007259+shahar1@users.noreply.github.com> Co-authored-by: Jarek Potiuk <potiuk@apache.org> Co-authored-by: kuse <3133746534@qq.com> Co-authored-by: Davide Polato <dpol1@apache.org> Co-authored-by: Arnav <imarnavpurohit@gmail.com> * refactor(skills): generate the Adopter overrides section from a shared block (#1532) The "## Adopter overrides" preamble and its Hard rule were hand-copied into 50 SKILL.md files in about a dozen slightly different wordings. They now come from tools/dev/blocks/adopter-overrides.md, kept in sync by check-shared-blocks.py. check-shared-blocks.py gains an {override_name} placeholder, filled with the name of the repo-root skills/<name> symlink that points at the skill directory; a skill with no such symlink, or with several, is a hard error. Blocks without the placeholder are unchanged. write-skill scaffolds the empty block region for new skills. Generated-by: Claude Opus 5 * fix(release-vote-tally): bind a vote to its real sender address only (#1530) `normalise_address` took the first `<…@…>` anywhere in `from`, so a display name written as an address bound the vote to that person: `"<alice@apache.org>" <mallory@example.org>` counted as a binding vote from roster member alice, and because voter identity drives supersession, such a later vote replaced alice's real one. Parse `from` as an address header with `email.utils.getaddresses`, so only the actual mailbox counts. A `from` holding several addresses, or one the parser rejects, is non-binding and keeps an identity of its own, so it can never supersede another voter's vote. Bare addresses and bare handles are taken as written, as before. Generated-by: Claude Opus 5 * feat(config): keep install-only personal config in the git directory (#1533) A project that only installs Magpie families, without adopting Magpie (no committed .apache-magpie.lock), no longer gets anything in its working tree. Its personal configuration layer is now <git-common-dir>/apache-magpie/: never committed, needing no ignore entry, and shared by every worktree of the clone. Adopted projects keep .apache-magpie-local/ and .apache-magpie-overrides/ as before. The rule lives in setup_preflight/layers.py, which computes the git common directory by reading files rather than spawning git, and never creates the directory on a read. The tools that resolve configuration carry identical copies, kept in step by an AST test: release-config, adversarial-review, the privacy-llm checker, agent-guard, the status collector, container-gateway and sandbox-lint. - Pre-flight: a new legacy-local-dir finding offers, with confirmation, to move an old in-tree .apache-magpie-local/ of an unadopted repo into the git-directory home; until then it is still read. - privacy-llm checker: now reads the personal layer (it only ever read .apache-magpie/ and the overrides), and no longer looks in the framework snapshot. - container-gateway: an unadopted repo serves from <git-common-dir>/apache-magpie/run/<worktree-id>/, one per worktree, created level by level with mode 0700; serve refuses socket paths over the sun_path limit, and the run dir and personal layer are never accepted as bind sources (compared after resolving symlinks on both sides). A linked worktree's common directory is trusted only when it is owned by the user, not group/world-writable, holds HEAD and objects/, and its worktrees/<name>/gitdir links back to this worktree, so a rewritten .git file cannot choose where sockets are bound. sandbox-lint exempts exactly that path. - Specs and docs updated for the new locations. Generated-by: Claude Opus 5 * test(release-verify-rc): grade Step 6c paste recipes by their rules, not one reference text The Step 6c suite failed 4-6 of 9 cases on `paste_recipe` alone: the grader compared each candidate recipe with one reference recipe word for word, while the step only requires properties of it. The eval also never showed the model the adapter's recipes that Step 6c tells it to follow. - Replace the exact `paste_recipe` in every case with structural checks in a new `assertions.json`, encoding the output-spec rule: an existence check against a concrete repository URL, an inventory listing (the crawl, inlined or referenced), the `--netrc-file` state check only when credentials are available, no write verbs or request bodies, no inline `-u` credentials, and a comment for a `SKIP`. Every original reference recipe satisfies them. - Include `tools/asf-nexus/operations.md` in the step's `also_include`, as the step links it at runtime. - State two rules the fixtures relied on but the step never said: `staging_repos` is ordered by repository id, and a `SKIP` leaves `nexus_findings` empty except for a malformed id, which records the rejected value. The suite now passes 9/9 on two consecutive runs. Generated-by: Claude Opus 5 --------- Signed-off-by: Davide Polato <dpol1@apache.org> Signed-off-by: Andrea Cosentino <ancosen@gmail.com> Co-authored-by: Shahar Epstein <60007259+shahar1@users.noreply.github.com> Co-authored-by: Jarek Potiuk <potiuk@apache.org> Co-authored-by: Jarek Potiuk <jarek@potiuk.com> Co-authored-by: Vardhman Gupta <112063624+Kaap10@users.noreply.github.com> Co-authored-by: Davide Polato <dpol1@apache.org> Co-authored-by: Arnav <imarnavpurohit@gmail.com> Co-authored-by: Kavya Katal <KAVYAKATAL09@GMAIL.COM> Co-authored-by: Andrea Cosentino <ancosen@gmail.com>
Summary
nominationskill body budget from 5,610 down to 4,788 tokens (-14.7%, saving 822 tokens) to comply with the 5,000-token body ceiling (Optimize the contributor-growth skill family #1348).contributor-to-committer, identity map seeding limits, and merit-note triggers.Type of change
.claude/skills/<name>/) — eval fixtures updated belowtools/<system>/*.md)tools/*/withpyproject.toml)docs/,README.md,CONTRIBUTING.md)projects/_template/)prek, workflows, validators)Test plan
prek run --all-filespassesmeasured_tokensstamped to4788(surface_hash(sha256:ce38f115ea57c59b) reconciled and verifiedRFC-AI-0004 compliance
<PROJECT>,<tracker>,<upstream>,<security-list>) used in all skill / tool proseLinked issues
Part of #1348 (Phase 2, PR 5) — Refs #1342
Notes for reviewers (optional)
Addressed review feedback:
(required — do not skip)instruction.AGENTS.md.measured_tokens: 4788and verifiedsurface_hash: sha256:ce38f115ea57c59b.