Skip to content

Agent metadata drift: CI Miri gating (PR #32) and qemu-boot skill's stale OVMF-fetch claim #34

Description

@chbaker0

Filed by an AI coding agent (Claude Code).

Routine audit of Claude/agent metadata (AGENTS.md, CLAUDE.md, .claude/**) against actual project state. Anchor commit was cab80240 (the last commit to touch any in-scope metadata file — the PR #28 merge, itself a metadata-drift fix, 2026-07-22). Everything reviewed is since then: PR #32 (81ef154/72f6b69, Miri CI optimization) and 077969b ("Add OVMF prebuilts"), plus a small edition tweak in Cargo.toml/shared/Cargo.toml. This is distinct from the previous audit's issue #27 — that drift was fixed by #28; the two items below are new drift introduced after it.

Drift found

  1. AGENTS.md's CI description predates PR CI: skip the Miri job when no Miri inputs changed #32's Miri gating. "Project status" enumerated the CI as "three jobs — cheap, expensive (Miri), smoke" and claimed CI "exercises everything ... on every push/PR". Since then, PR CI: skip the Miri job when no Miri inputs changed #32 made the expensive job path-gated: it runs Miri only when a Miri input changed (shared/**, Cargo.lock, Cargo.toml, .cargo/config.toml, rust-toolchain, or the workflow itself) and otherwise reports a fast green no-op, plus it added a daily schedule: run (cron 0 6 * * *) that forces the full Miri suite regardless of paths to catch nightly drift the path filter can't see. The "Verifying changes" section's "rely on that CI job instead" note was likewise incomplete — it didn't mention the job only runs Miri on Miri-input changes (which, since Miri only exercises shared, still covers every change Miri can catch) or the daily drift-catching run.

  2. .claude/skills/qemu-boot/SKILL.md step 2 falsely claimed make-image.sh fetches OVMF. It read "This fetches OVMF prebuilts (first run only) and builds the loader + kernel". But make-image.sh uses the vendored firmware under third_party/ovmf (committed to the repo, no network I/O) and errors out if the blobs are missing rather than downloading them. This became concretely wrong once the firmware blobs were actually committed post-anchor (077969b). Left as-is it would mislead the boot-verification skill into expecting a first-run download and misdiagnosing a "firmware not found" failure.

  3. Minor: "Project status" Last updated bumped 2026-07-21 → 2026-07-23.

Fix

Draft PR corrects all of the above in the in-scope metadata files only (no code changes): updates AGENTS.md's "Project status" and "Verifying changes" CI descriptions for PR #32's path-gating + daily scheduled Miri run, and rewrites the qemu-boot skill's step 2 to describe the vendored (no-fetch) OVMF firmware.

Flagged, not fixed here

No new fragile/manual workflow surfaced that isn't already covered by an existing skill/subagent (qemu-boot, source-grounded-explorer) or a prior issue (#7, #22, #27). Two standing gaps the audit re-surfaced, with external-tooling research (both subagent-verified against Anthropic's own docs/marketplace, not third-party mentions):

Gap A — Rust nightly toolchain drift resilience. rust-toolchain pins bare nightly (no date), so a fresh checkout can break on unstable-API drift (-Zbuild-std, target-spec fields, dependency trait impls); the new daily-Miri CI job exists precisely to catch this. Only one Anthropic-listed candidate survived confirmation, and it's an indirect fit:

  • rust-analyzer-lsp (official claude-plugins-official marketplace; https://code.claude.com/docs/en/discover-plugins) — wires the local rust-analyzer into Claude Code for post-edit diagnostics/navigation. Actively maintained (plugin touched Feb 2026; parent repo commits daily). It does not pin/monitor toolchains, but would surface a nightly break's compiler diagnostics inline, speeding the "suspect drift → diagnose" loop. Caveat: needs rust-analyzer installed and pointed at the custom target specs (x86_64-unknown-none.json etc.), untested for this no_std/custom-target project. No Anthropic-listed skill/plugin/MCP specifically targets toolchain-drift detection; the existing daily-Miri CI remains the most direct mitigation.

Gap B — nobody notices a red scheduled CI run. The new daily Miri job runs unattended; a failure (e.g. real nightly drift) can go unseen. Research:

  • The official GitHub MCP server (already in use) fully covers the read/triage mechanics — actions_list/actions_get/get_job_logs/get_check_run/actions_run_trigger list runs, fetch a run's status/logs, and re-trigger — but nothing in it watches on its own.
  • anthropics/claude-code-action (https://github.com/anthropics/claude-code-action; docs https://code.claude.com/docs/en/github-actions; very active, v1.0.181 released 2026-07-22) is the piece that would actually notice — it works with any GitHub event, so a workflow triggered on workflow_run (failure) or schedule could invoke Claude to pull the failing job's logs, summarize, and file/update an issue automatically. Concrete breadcrumb if the daily job's failures start slipping by.

If either gap becomes a repeat annoyance, the breadcrumbs above (a claude-code-action workflow_run-triggered auto-triage workflow; and/or the rust-analyzer-lsp plugin for faster local drift diagnosis) are the concrete next steps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions