feat(mcp): add Hub scenarios and opt-in agent evaluation - #1526
Conversation
|
🎨 Token Changes ReportTokens Changed (0)Original Branch: This comment was automatically generated by the token diff tool. 🤖 |
🧩 Component Schema Changes ReportComponent Schemas Changed (0 added, 0 deleted, 62 updated)Original Branch: ✅ No Breaking ChangesThis PR contains only non-breaking changes to component schemas. 🔄 Non-Breaking Updatesaccordion action-bar action-button action-group alert-dialog avatar-group avatar badge breadcrumbs button-group button calendar cards checkbox-group checkbox close-button color-area color-handle color-loupe color-slider color-wheel combo-box contextual-help date-picker divider drop-zone field-label help-text illustrated-message in-line-alert link list-view menu meter number-field picker popover progress-bar progress-circle radio-button radio-group search-field segmented-control select-box side-navigation slider standard-dialog status-light swatch-group swatch switch table tabs tag-group tag takeover-dialog text-area text-field thumbnail toast tooltip tree-view This diff was generated automatically and contains only backward-compatible changes. This comment was automatically generated by the component schema diff tool. 🤖 |
Run report for 1c636bddTotal time: 6m 4s | Comparison time: 16m 24s | Estimated savings: 10m 20s (63.0% faster)
Expanded report
Changed files |
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2ecc660 to
1c636bd
Compare
Description
Phase two of bead
spectrum-design-data-v3xu, stacked on #1525 (which builds on #1524).Merge in that order.
Adds the private
tools/design-data-mcp-evalpackage with 13 source-backed scenarioscovering copy, containers, typography exceptions, button behavior, unsupported
requests, conflicting guidance, and guideline/token lookup.
Offline AVA contracts run in normal CI and before publishing. A separate manually
triggered Actions workflow drives an agent through the real MCP tools using an
explicitly configured OpenAI-compatible endpoint. The workflow runs only on
mainand does not gate PRs or publishing. Weekly runs require explicit opt-in.
The runner exposes read tools only, checks packaged guidelines before spending
inference requests, bounds requests/context/turns, and retains answers, tool traces,
screen results, rubrics, model and artifact provenance. Execution/provider failures
are distinct from screen regressions. Pattern/citation screens still require human
review; they are not proof of semantic correctness.
The review fix loads isolation guards from outside the artifact instead of
copying them into the Moon-staged production bundle. Guards constrain both ESM
and CommonJS resolution to the server's artifact working directory. Regressions
check rejected workspace fallback, byte-for-byte staging invariance after a real
MCP connection, and exclusion of both helpers from a subsequently packed MCPB.
The corrected public corpus and contracts from #1524/#1525 are also propagated.
Related Issue
Bead
spectrum-design-data-v3xu. Depends on #1525 and #1524.Review follow-up:
spectrum-design-data-prix.Motivation and Context
The deterministic checks catch corpus and packaging regressions without a model.
The opt-in runner exercises actual tool selection and answer synthesis while keeping
provider costs and credentials out of ordinary PR CI.
Set repository variables
EVAL_ENDPOINTandEVAL_MODEL, plusEVAL_API_KEYif theapproved endpoint requires authentication. No provider or credentials have been
configured by this PR.
EVAL_SCHEDULE_ENABLED=trueenables weekly runs. There is noautomatic paid-provider fallback.
The spacing scenario queries
property=spacingand selectsraw.name.scaleIndex=300from returned data. The current filter grammar does not accept
scaleIndex, andproperty-only resolve cannot resolve a legacy name like
spacing-300. The fixturepreserves that actual API boundary rather than testing an unsupported query.
How Has This Been Tested?
Local macOS, Node 20.17.0, pnpm 10.17.1, Moon 2.3.3.
moon run :test: passed (32 tasks completed; 27 cached on the final review-fix run).moon run design-data-mcp-eval:test design-data-mcp:test design-data-agent-mcp:test spectrum-hub-fetcher:test s2-docs-to-document-blocks:test: passed; all 32 evaluation tests passed, including archive packing and both import-escape regressions.moon run design-data-mcp:bundle: passed after a default evaluation connection verified every current guideline without changing production staging. Inspection withunzip -Z1confirmed that the actual production archive contains neither test guard.node .github/scripts/check-ci-coverage.mjs: passed, zero uncovered tasks.node tools/changeset-linter/src/cli.js check --fail-on-warnings: passed.pnpm exec prettier --check tools/design-data-mcp-eval .github/workflows/hub-mcp-eval.yml .github/workflows/release.yml .github/ci-targets.json .moon/workspace.yml: passed after the commit hook.CLI integration tests use actual private MCP bundles and a local fixture inference
server. They verify tool-result feedback, report provenance, and exit codes for a
passing screen, regression, and provider outage. Further tests cover stale artifacts,
read-tool restrictions, budgets, citation/fact failures, and abstention-screen limits.
Full tests were also run separately on the updated #1524 and #1525 trees before
propagation, rather than validating only the stack tip.
No paid/live model was invoked. The new Actions workflow cannot run until the stack
is merged to
mainand provider configuration is supplied. Answer-quality calibrationremains a human-reviewed operational step, not a claimed validation result.
Screenshots (if appropriate):
Not applicable.
Types of changes
Checklist: