Skip to content

perf(contributor-growth): split committer-onboarding, rewire activity-sweep, tighten wording - #1548

Merged
potiuk merged 5 commits into
apache:mainfrom
potiuk:perf/contributor-growth-opt
Oct 6, 2026
Merged

potiuk merged 5 commits into
apache:mainfrom
potiuk:perf/contributor-growth-opt

Conversation

@potiuk

@potiuk potiuk commented Oct 6, 2026

Copy link
Copy Markdown
Member

What

A second optimize-skill sweep of the contributor-growth family, after #1481–#1489.
Of the nine skills, only committer-onboarding was over both budgets, and activity-sweep still wrote its own GitHub queries.
One pass per commit:

Commit Pass Effect
activity-sweep → contributor-metrics rewire Four hand-written GraphQL searches (writing to /tmp) replaced by one cached contributor-metrics fetch + score. Body 3,569 → 3,361 tokens.
split committer-onboarding split Per governance model (asf-pmc / github-codeowners / maintainer-roster) and per IP-intake model (icla / dco / no-cla) branches moved byte-for-byte into detail/; only the configured model's file is read. 780 → 477 lines (clears the 500-line cap).
trim routing metadata metadata description + when_to_use ~205 → ~159 tokens, paid in every session; every trigger phrase kept, skip condition made explicit.
tighten wording rewrite Maintainer-approved style: restated detail cut in favour of a pointer, caveats merged with their action, imperative voice. Headings, code fences, templates, field names and gates unchanged.
fixes fix Two contradictions and an eval fixture (below).

committer-onboarding overall: body 8,042 → 4,722 tokens and 780 → 364 lines; a default asf-pmc + icla run loads about 6.7k tokens with its two model files, other models less.

activity-sweep counting

Counts now come from contributor-metrics, dated by the contributor's own activity like the rest of the family: reviews by review date (was PR creation), threads only when the contributor's own comment is in the window (was any thread updated in it), and an inclusive window start.
These are corrections, not regressions.
contributor-metrics gains fetch --since, adjustable substantive-review thresholds (activity-sweep keeps its own: body over 50 characters or at least 3 inline comments), and score --timeline-kinds, with tests and README.
The model no longer sees review titles, bodies or comment text — links, dates, counts and flags only.

Fixes

  • governance-asf-pmc.md said a new PMC member cannot self-subscribe to private@, contradicting karma-grant.md: they subscribe via Whimsy or private-subscribe@, and a moderator approves.
  • The asf-pmc completion example listed a manual GitHub org invite the checklist says the flow does not have; it now shows the automatic gitbox access.
  • The committer-onboarding step-1 eval graded fields from JSON but never asked for JSON, so 6–7 of its 9 cases errored on a prose answer on every run (also before this PR). Its prompt now asks for one JSON object.

Testing

  • Evals:
    • activity-sweep 12/12 before and after;
    • committer-onboarding steps 0, 3 and 4 pass in full;
    • step 1 went from 2–3/9 graded (the rest erroring) to 26/27 over three runs. The one miss, case-7 drafting the secretary request while the ICLA is unprocessed, passed in the other two runs.
  • tools/contributor-metrics pytest (54, 5 new), ruff, mypy.
  • Validator exit 0; committer-onboarding's skill-line-limit warning is gone, no new warnings.
  • prek run --all-files green.

Gen-AI disclosure: prepared with the assistance of Claude Code (Claude Opus 5); reviewed by the author.

🤖 Generated with Claude Code

potiuk added 5 commits October 6, 2026 16:29
…r-metrics

activity-sweep built its counts from four hand-written GitHub searches, each
writing its query to /tmp, and read raw review bodies into context to judge
substantive reviews. It now runs the family's cached contributor-metrics
fetch once and reads the raw counts, merge rate and timeline from score.
The card's rows, sections and wording are unchanged, and no titles, bodies
or comment text reach the skill any more.

To keep the card's numbers on its own definitions, the tool gains three
small options:

- fetch --since YYYY-MM-DD, for windows trimmed to a repository's creation
  date, which are not a whole number of months;
- fetch --substantive-body-chars / --substantive-line-comments, so the card
  keeps its "body over 50 characters or 3+ inline comments" rule while the
  nomination skills keep the 100 / 1 default (part of the cache key only
  when changed, so existing caches stay valid);
- score --timeline-kinds, so the card's timeline combines its four streams
  without the issues-triaged threads it does not show.

Generated-by: Claude Opus 5
…nd intake model

SKILL.md carried every branch for all three governance models
(asf-pmc, github-codeowners, maintainer-roster) and all three IP
intake models (icla, dco, no-cla), though a run only ever uses one
of each. Move each model's blocks, byte for byte, into one sibling
per model under detail/ -- the Step 0 vote bar, Step 1c account
request, Step 3 access checklist and Step 4 summary example for a
governance model; the Step 1a check for an intake model -- and leave
a pointer at each spot saying only the configured model's file is
read. Text shared by every model stays in SKILL.md.

SKILL.md drops from 780 to 477 lines and from 8,042 to 5,198
measured tokens, clearing the 500-line validator warning. The only
edits inside moved text are three relative link targets
(detail/email-templates.md, detail/karma-grant.md) rewritten to
resolve from detail/. The eval step-configs pull the new files in
via also_include so the extracted prompts keep the moved text.

The vendor-neutrality scorer reads SKILL.md only, so with the gh
recipes now under detail/ it rates the skill capability-pure;
docs/vendor-neutrality.md is regenerated to match.

Generated-by: Claude Opus 5
Fold the description into one sentence and drop the when_to_use
connective text ("Trigger phrases:", "Also appropriate immediately
after ... when the user asks"), keeping every quoted trigger phrase
and the contributor-nomination follow-on, and stating the skip
condition explicitly: not while the vote is still open.

description + when_to_use go from 819 to 635 frontmatter characters
(about 205 to 159 tokens at chars/4).

Generated-by: Claude Opus 5
Wording pass on committer-onboarding and its detail files, with the
maintainer approving the style: restated detail cut in favour of a
pointer to its source, caveats merged with their action, imperative
voice, semantic line breaks. Headings, code fences, templates, field
names, thresholds and every gate are unchanged.

Generated-by: Claude Opus 5
…ep-1 eval output

- governance-asf-pmc said a new PMC member cannot self-subscribe to
  private@, contradicting karma-grant: they subscribe via Whimsy or
  private-subscribe@ and a moderator approves.
- The asf-pmc completion example listed a manual GitHub org invite,
  which the checklist says the flow does not have; it now shows the
  automatic gitbox access.
- The step-1 eval graded fields from JSON but never asked for JSON, so
  most cases errored on a prose answer. Its prompt now asks for one
  JSON object, and its extraction includes the email templates.

Generated-by: Claude Opus 5
@github-actions github-actions Bot added family:tools tools/* family:docs Docs, MISSION.md, READMEs capability:triage Sweep + classify + propose disposition capability:resolve Close-out: invalidate, dedupe, CVE allocate, announcement capability:stats Read-only dashboards, metrics, governance evidence family:contributor-growth contributor-growth skills substrate:analytics Tool substrate: read-only metrics / dashboards / renderers labels Oct 6, 2026
@potiuk
potiuk merged commit d3f01ee into apache:main Oct 6, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

capability:resolve Close-out: invalidate, dedupe, CVE allocate, announcement capability:stats Read-only dashboards, metrics, governance evidence capability:triage Sweep + classify + propose disposition family:contributor-growth contributor-growth skills family:docs Docs, MISSION.md, READMEs family:tools tools/* substrate:analytics Tool substrate: read-only metrics / dashboards / renderers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant