Skip to content

downstream tmt test failure #2559

Description

@cgwalters

https://artifacts.dev.testing-farm.io/eef0709a-b4b6-42a5-9932-70045f05fb87/

We need to capture the journal (surprising tmt doesn't do this by default)

Activity

  1. cgwalters commented on Oct 8, 2026

    @cgwalters
    CollaboratorAuthor

    @cgwalters-bot take this at p0 - quick PR to add the logging ensure successful log of journal of each failed services w/synthetic failure.

    I think problem here is we are only testing the system-reinstall-bootc path in testing farm...probably most effective path here actually...can we setup a variant of our test suite that doesn't use bcvk but uses testcloud to boot and then does the in-place reprovision?

    If we do that even for the readonly test and can reproduce the porblem that'd be great. basically just tmt-readonly should run both provisioning paths.

    Look carefully at surrounding code and history create a plan don't make mistakes

  2. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    On it at P0, as two runs:

    • The journal capture for failed tests, with a synthetic failing unit as proof: run 37833703108. It comes back as a pull request for your approval.
    • The plan for a testcloud boot plus in-place reprovision variant, so just tmt-readonly covers both provisioning paths: run 37833802970, read-only, checking the code and history first. The plan will be posted here.

    Generated-by: https://github.com/cgwalters/#llms

  3. ckyrouac commented on Oct 8, 2026

    @ckyrouac
    Collaborator

    @cgwalters-bot ah I had an agent working on this during our meeting, here's the PR: #2560

  4. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    Seen, thanks: #2560 covers the journal output for failed units, so I've stopped my run for that half rather than open a second pull request for the same thing. The read-only planning run for the testcloud plus in-place reprovision variant continues and will post here.

    @cgwalters: you asked for proof with a synthetic failure. If you want it, I can run #2560 against a deliberately failing unit and report what the test output contains; say the word.

    Generated-by: https://github.com/cgwalters/#llms

  5. ckyrouac commented on Oct 8, 2026

    @ckyrouac
    Collaborator

    @cgwalters: you asked for proof with a synthetic failure. If you want it, I can run #2560 against a deliberately failing unit and report what the test output contains; say the word.

    @cgwalters-bot yes do that

  6. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    Started: run 37834711147 checks #2560 against synthetic failing units and will report here what the test output contains.

    Generated-by: https://github.com/cgwalters/#llms

  7. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    Review by an LLM (run 37834711147), requested by the coordinator. Not a human review.

    LLM-generated by praxis/gpt-6.1-sol

    Verdict: #2560 appears to implement the requested per-unit journal diagnostics for the failures this test selects, but I could not verify successful collection with synthetic failures. Level 1 is complete; level 2 was blocked provisioning a systemd container, so neither the failing/clean runtime cases nor the real tmt path were run.

    Source behavior (two sentences): When selected units have failed, the script prints their systemctl rows, runs journalctl --boot --unit $unit --no-pager separately for each, prints nonempty stdout/stderr and any nonzero journalctl status, then asserts zero failures. It prints journals only for failed unit names starting with systemd- or bootc-, excluding systemd-tpm2-setup.service when /sys/fs/selinux/enforce exists; plain synthetic-fail-a.service would be ignored.

    Commands actually run:

    git fetch https://github.com/bootc-dev/bootc pull/2560/head
    git worktree add --detach /home/runner-sandbox/bootc-pr2560 FETCH_HEAD
    git diff HEAD FETCH_HEAD -- tmt/tests/booted/readonly/014-test-no-failed-units.nu
    podman --version && podman info
    podman images --format '{{.Repository}}:{{.Tag}}' && command -v nu qemu-system-x86_64 virsh bcvk just; test -e /dev/kvm
    podman run -d --name bootc-pr2560-systemd --systemd=always --network=none quay.io/centos/centos:stream10 /sbin/init
    podman build -t localhost/bootc-pr2560-systemd -f /home/runner-sandbox/Containerfile.pr2560 /home/runner-sandbox
    

    The build was retried once with Python subprocess/monotonic timing: exit 1, 0.411 s. Scratch Containerfile: CentOS Stream 10 base; dnf -y install systemd curl tar && dnf clean all; fetch/extract Nushell 0.103.0 from the same release URL used by hack/provision-fetch.sh; /sbin/init as CMD. The build never reached the Nushell download.

    Relevant output:

    HEAD is now at abb4076 tmt: Print journals for failed units
    podman version 5.8.2
    rootless: true
    Error: crun: executable file `/sbin/init` not found: No such file or directory
    SSL certificate problem: unable to get local issuer certificate
    Error: Failed to download metadata for repo 'baseos': Cannot prepare internal mirrorlist
    Error: building at STEP "RUN dnf -y install systemd curl tar && dnf clean all": while running runtime: exit status 1
    

    Read the changed test, tap.nu (begin prints TAP version/description, ok prints ok), hack/provision-fetch.sh (CentOS 10 pins Nu 0.103.0), and the Justfile. No locally installed Nu, QEMU, virsh, bcvk, or just was found by that command, and /dev/kvm was absent. I did not disable TLS verification or change the host systemd manager. Scratch worktree/container were removed; the original checkout is unchanged.

    Not run: two selected synthetic failed units, verification of each unit's journal and a recognizable marker, the clean case, or just test-tmt readonly (the documented real path). Per the requested stop-at-first-unreachable-level rule, I stopped at level 2 after the systemd-container dependency installation failed TLS validation; there are no synthetic-test output/artifact claims.

    Edge cases below are source inspection, not runtime observations:

    • No journal lines: the heading still prints; any nonempty journalctl placeholder/stderr is printed, and the failed-unit assertion remains.
    • Escaped unit names: $unit is passed as a separate argument without shell interpolation; actual systemd escaping was not exercised.
    • Long journals: no line/byte limit; complete captures the entire current-boot journal per unit before printing it.
    • Failures before journald is available: no retry or special recovery; only available current-boot entries can be printed.
    • Discovery: the exact command is systemctl list-units --failed --no-legend --plain. --failed is the failed-state filter (equivalent in intent to --state=failed for list-units); this remains a loaded-unit query, not a scan of unit files or historic failures. No runtime comparison was made.

    Suggestions for the pull request's author:

    • Use systemd-synthetic-fail-a and bootc-synthetic-fail-b for the synthetic check so both pass the existing prefix filter; have the latter print a recognizable marker before exiting nonzero.
    • Consider an explicit journal size limit if unbounded diagnostic output is undesirable.
    • Exercise an escaped unit name and an unavailable/empty journal alongside the failing and clean cases on a working systemd/Nu 0.103.0 test system.

    Generated-by: https://github.com/cgwalters/#llms

  8. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    Review by an LLM (run 37833802970), requested by the coordinator. Not a human review.

    LLM-generated by Opus 5.5. This is an analysis and plan only; nothing was changed. Citations are to main at 7702842 unless a commit is named. Anything marked (inferred) was not verified by running it.

    Summary

    Your hypothesis holds. The failing check in the linked run is a unit that only exists on the system-reinstall-bootc path. Testing Farm runs bootc install to-existing-root --cleanup from a stock cloud image. That enables bootc-destructive-cleanup.service, and in the linked run that service is failed. The local flow (bcvk) and the GHA flow (also bcvk) never take that path. GHA does run system-reinstall-bootc, but it never reboots into the result. The cheapest reproducer is a second provisioning mode in cargo xtask run-tmt: testcloud boots a stock cloud image, then the image under test is reinstalled in place. It uses the same plans and selects the mode with a tmt context value.

    1. How the local tmt targets work today

    • There is no just tmt-readonly. The closest is just test-tmt readonly (Justfile:191-194). It runs build (podman build of the repo Dockerfile from the RPMs made by just package, giving localhost/bootc), then _build-upgrade-image, then test-tmt-nobuild. That last one is cargo xtask run-tmt ... --upgrade-image=localhost/bootc-upgrade localhost/bootc readonly (Justfile:298-299).
    • run_tmt (crates/xtask/src/tmt.rs:371 onward):
      • It requires bcvk, tmt, rsync and podman (tmt.rs:86-91).
      • It sets the contexts running_env=image_mode and distro=<ID-VERSION_ID of the image> (tmt.rs:384-394).
      • It rsyncs .fmf and tmt/ into target/tmt-workdir (a workaround for Still fix rsync to use .gitignore teemtee/tmt#4062) and lists plans with tmt plan ls --filter enabled:true (tmt.rs:445).
      • For each plan it boots a fresh VM with bcvk libvirt run ... [--bind-storage-ro] [--log-dir=journal,console=...] <image> (tmt.rs:615) and gets the SSH port and key from bcvk libvirt inspect.
      • It then runs tmt run --all -e TMT_SCRIPTS_DIR=... -e BCVK_EXPORT=1 provision --how=connect --guest=localhost ... plan --name <plan> (tmt.rs:743-752).
    • bcvk boots the bootc image directly, so the root is laid out by bootc install to-disk. (Inferred from bcvk's role and from the comment at tmt/tests/booted/readonly/052-test-bli-detection.nu:30-34: "In Packit/gating CI the system was installed via to-existing-root".)
    • Plans: cargo xtask update-generated writes everything between the BEGIN/END GENERATED PLANS markers in tmt/plans/integration.fmf from the metadata headers in tmt/tests/booted/*.nu (tmt.rs:1132); just validate checks it. The header above the markers is hand-written:
      • provision: how: virtual, image: $@{test_disk_image} (integration.fmf:1-4). Locally this is overridden by connect, and the context it reads is a leftover of the testcloud era (see section 4).
      • Three prepare phases, all when: running_env != image_mode (integration.fmf:26,34,41). Locally they are therefore off.
    • GHA (.github/workflows/ci.yml):
      • test-integration runs just test-tmt integration or just test-composefs ... (ci.yml:406-411).
      • test-upgrade boots the published base under bcvk and upgrades (Justfile:243-262).
      • tmt is pinned at 1.79.0 with [provision-virtual] (.github/actions/install-tmt/action.yml:18-19), so testcloud is already installed on the runners.

    2. What Testing Farm runs

    • Upstream Packit (.packit.yaml:84-127, label-gated since 44ce0b7):
      • Plans: tmt_plan: /tmt/plans/integration, which is the same generated plans as the local flow, with context running_env: packit.
      • ci/tier-1 runs centos-stream-10 and fedora-latest-stable on x86_64 and aarch64. ci/merge adds c9s and fedora-all.
    • Downstream (the linked run):
      • It is a Fedora CI run (initiator: fedora-ci) of Koji task 151003669, bootc-1.17.1-1.fc44, Fedora-44 x86_64.
      • Its plans come from plans/all.fmf in src.fedoraproject.org/forks/packit/rpms/bootc @ 06c498ef, not from this repo.
      • Those plans have the same prepare phases plus one extra line, cp os-image-map.json bootc/hack, and no running_env context. The phases still ran, so running_env != image_mode evaluates true when the dimension is undefined.
    • Provisioning:
      • Testing Farm brings its own machine. In this run tmt was invoked with provision --how connect --guest <EC2 address> --hard-reboot 'artemis-cli guest reboot ...' (per-plan log.txt). The plan's how: virtual is not used.
      • The in-place reprovision is done by the prepare phases, added in 704338d and adjusted in f1dec83. They:
        • install podman, bootc, system-reinstall-bootc, expect and ansible-core from the build under test;
        • unpack the SRPM and run hack/provision-packit.sh. That script builds localhost/bootc FROM the published base for the OS (os-image-map.json; provision-packit.sh:38,89). It uses hack/Containerfile.packit, which runs dnf -y update bootc from the test-artifacts repo and provision-derived.sh cloudinit (Containerfile.packit:29,32);
        • pull the LBI images (provision-packit.sh:92);
        • run hack/system-reinstall-bootc.exp, i.e. system-reinstall-bootc localhost/bootc. In the artifact this executed podman run ... localhost/bootc bootc install to-existing-root --acknowledge-destructive --skip-fetch-check --cleanup --root-ssh-authorized-keys ...;
        • reboot through the Ansible playbook hack/packit-reboot.yml, which is fetched by URL from GitHub main (integration.fmf:39-40), not from the tree under test.
    • What differs from the local path:
      1. to-existing-root instead of bcvk's disk install. The old OS stays under /sysroot.
      2. --cleanup (crates/system-reinstall-bootc/src/podman.rs:83-88) writes the etc/bootc-destructive-cleanup stamp (crates/lib/src/install.rs:224). The generator then enables bootc-destructive-cleanup.service (crates/lib/src/generator.rs:81-83), which runs contrib/scripts/fedora-bootc-destructive-cleanup: it rpm -es every package of the old root, then podman system prune --all -f.
      3. The image is the published base plus the RPM, not the repo Dockerfile image.
      4. There is no BCVK_EXPORT and no bind storage. Readonly tests already skip on that (017, 052).
      5. Cloud hardware (EC2 here), and aarch64 as well.
    • Nothing upstream boots a reinstalled system. The GHA install-tests job runs bootc-integration-tests system-reinstall on the Ubuntu host (ci.yml:227). That test only checks that the cleanup files are present (crates/tests-integration/src/system_reinstall.rs:92-99) and never reboots. So the cleanup unit has never run in upstream CI outside Testing Farm.

    3. The linked artifacts

    • 9 plans ran. Only plans/all/plan-01-readonly failed; the other 8 passed.
    • Prepare for readonly succeeded:
      • image built FROM quay.io/fedora/fedora-bootc:44 (23:03:35 to 23:09:21);
      • install finished at 23:12:47 ("Installation complete!");
      • Ansible reboot ok at 23:13:39.
    • Execute: subtests 000-013 pass. 014-test-no-failed-units.nu fails with:
      bootc-destructive-cleanup.service loaded failed failed Cleanup previous the installation after an alongside installation
      The AVC check has no matches. SELinux is enforcing, with selinux-policy-44.11-1.fc44.
    • 014 was added on 2026-09-18 in 41049cc. Before that, the readonly suite did not look at failed units, so a failing cleanup unit would have gone unreported. (Inferred: I can't tell from here since when it fails.)
    • The other plans also reinstalled with --cleanup, but none of them checks failed units. Whether the unit failed there too is unknown.
    • What is missing: the unit's journal (journalctl -b -u bootc-destructive-cleanup.service, which would include the script's set -x trace), its exit status (systemctl status), and the state of /sysroot afterwards. Nothing in the artifacts says which command in the script failed. That is part 1 of this issue (journal capture).

    4. testcloud

    • What tmt 1.79.0 needs: tmt[provision-virtual] requires testcloud>=0.11.7 (PyPI metadata; latest is 0.12.0). testcloud then needs libvirt and its Python bindings, plus qemu with /dev/kvm. Without KVM it falls back to TCG with multiplied timeouts (tmt/steps/provision/testcloud.py:833-835).

    • Defaults: qemu:///session (testcloud.py:176-185), 2048 MB of memory and a 40 GB disk (testcloud.py:176-185).

    • Images: an alias such as fedora-44 or centos-stream-10 (testcloud util.py:146-149), a URL, or an absolute qcow2 path (testcloud.py:875-905).

    • Networking and SSH: session mode uses user-mode networking with a host forward to guest port 22 (testcloud domain_configuration.py:283). tmt generates the SSH key and injects it with cloud-init, so the image needs cloud-init (stock cloud images have it).

    • Logs: console only (console.txt, testcloud.py:114). There is no journal capture.

    • What tmt does not expose: no virtiofs option, so nothing like bcvk's --bind-storage-ro.

    • UEFI: tmt 1.79.0 supports hardware: boot: method: uefi and tpm for testcloud (testcloud.py:218,473-578), and tests-install.fmf:8-12 already uses it. (Not exercised here.)

    • Compared with bcvk: bcvk boots the bootc image itself, returns its own SSH key, mounts host container storage over virtiofs, and captures the journal and console (bec34ca).

    • History: the repo used testcloud from 36be46a (2024-06-10) to f8ce015 (2025-11-04), but to boot a prebuilt bootc disk image, never to reinstall. tests/run-tmt.sh passed --context test_disk_image=target/bootc-integration-test.qcow2 (see git show f8ce0152^:tests/run-tmt.sh). It needed several workarounds:

      • pruning the testcloud cache (cc42685, testcloud#17);
      • a chcon shim on Ubuntu (testcloud#18, removed in f8ce015);
      • the tmt-workdir copy (tmt#4062).

      f8ce015 dropped testcloud because "testcloud+tmt doesn't support UEFI (Support configuring UEFI on x86_64 in testcloud provision teemtee/tmt#4203) ... a blocker for UKIs", and because of bcvk's ergonomics. tests-install.fmf:5-7 also records SELinux being disabled on Ubuntu (tmt#3364).

    • Machine for this analysis: it had /dev/kvm but no libvirt, tmt or bcvk, so I ran nothing.

    Plan

    Design (smallest that reuses every plan):

    • Selection: a flag cargo xtask run-tmt --provision=bcvk|reinstall, defaulting to bcvk so today's behaviour is unchanged. In reinstall mode, run_tmt:

      • sets --context running_env=reinstall instead of image_mode;
      • does not set BCVK_EXPORT;
      • skips bcvk;
      • runs tmt run --all ... provision --how virtual --image <alias> --memory 4096 plan --name <plan> per plan. The alias comes from the detected distro: centos-N becomes centos-stream-N, and fedora-N stays fedora-N.
    • Transport: before the loop, podman save the image under test into the tmt tree. This is the same trick the CoreOS path already uses (tmt.rs:422-431, loaded in tmt/tests/install/test-on-ostree/test-install-on-ostree.sh:30-33). Copy target/packages/*.rpm next to it, so the stock guest gets this build's system-reinstall-bootc. The repo Dockerfile already installs that RPM (Dockerfile:86-90).

    • Plans: add one hand-written prepare phase to the integration.fmf header, when: running_env == reinstall. It runs a new small script, hack/provision-reinstall-local.sh, which:

      • dnf installs the RPMs from the tree;
      • runs podman load;
      • pulls the LBIs (the same line as provision-packit.sh:92);
      • runs the existing hack/system-reinstall-bootc.exp;
      • reboots.

      The three Testing Farm phases change to when: running_env != image_mode and running_env != reinstall. Generated plans stay untouched. This does not affect downstream Fedora CI, which has its own copy of the header. (Inferred from the extra cp os-image-map.json line.)

    • Justfile:

      • test-tmt-reinstall *ARGS: build calls run-tmt with --provision=reinstall;
      • test-tmt-nobuild --provision=reinstall readonly already works, because ARGS are passed through;
      • a readonly target runs both paths (name: see decisions).
    • CI: Testing Farm stays as it is. Add a GHA leg for readonly only on the reinstall path, ostree only.

    Risks and unknowns:

    1. Reproduction: the failure was seen on Fedora 44 on EC2. The local default base is c10s, so the first try should use BOOTC_base=quay.io/fedora/fedora-bootc:44 with the fedora-44 cloud image. The failure may depend on the cloud image contents.
    2. The image differs from Testing Farm's (repo Dockerfile vs published base + RPM). This is deliberate, so that only the provisioning path changes. If the failure does not reproduce, a Containerfile.packit-style build is the fallback.
    3. Reboot in prepare: integration.fmf:35-36 says tmt-reboot does not work in prepare, but that was written in 2025-09. tmt 1.79.0's prepare/shell.py handles reboots (lines 139,314). (Unverified.) The fallback is a local copy of packit-reboot.yml, but ansible-playbook runs on the control node (tmt guest/init.py:2372), which means adding ansible-core to the install-tmt action.
    4. tmt when with an undefined dimension: the compound condition must still be true for downstream, which has no running_env. Check with tmt plan show both with no context and with running_env=packit.
    5. testcloud on Ubuntu GHA runners: the old SELinux/chcon problems may come back. Cloud images are fetched from mirrors (flake risk).
    6. Cost: each plan reinstalls. For readonly, Testing Farm went from the start of the in-guest build to the end of the reboot in about 10 min (23:03:35 to 23:13:39). About 6 min of that was the build, which the local path skips. Every plan also copies a tarball of about the image's size. This is why the reinstall path should be scoped to readonly at first.
    7. Arch and variants: testcloud runs only the native arch locally, so aarch64 remains Testing Farm only. composefs/UKI on the reinstall path is out of scope initially.
    8. Diagnosis: testcloud gives only the console log. The journal needs part 1, or a finish step that dumps journalctl for failed units on both paths.

    Steps (one PR each):

    1. Reproducer. Adds --provision=reinstall in xtask, the header phase, the script and just test-tmt-reinstall.
      Acceptance:
      • on a KVM host, BOOTC_base=quay.io/fedora/fedora-bootc:44 just test-tmt-reinstall readonly fails in 014 with bootc-destructive-cleanup.service failed, and the guest can be kept for journalctl;
      • just test-tmt readonly behaves as before;
      • cargo xtask update-generated direct --check passes;
      • tmt plan show still lists the Testing Farm phases with no context and with running_env=packit.
    2. Both paths in one target. A readonly target runs bcvk and then reinstall, with one summary. Acceptance: both runs appear and the exit status reflects both.
    3. Fix the cleanup failure from the journal obtained in step 1, keeping 014 as the regression check. Acceptance: step 1's command passes, and a ci/tier-1 Testing Farm run is green.
    4. CI. A GHA job runs reinstall readonly on centos-10 and fedora-latest, x86_64. Acceptance: green on a labelled PR, and red on a throwaway commit that makes the cleanup script exit 1.
    5. (Optional) Widen the reinstall path to more plans on a schedule, and/or add a Testing Farm-faithful image mode.

    Decisions for the maintainer

    • Target names. just tmt-readonly doesn't exist. I recommend test-tmt-readonly, running both paths, plus test-tmt-reinstall, and leaving test-tmt bcvk-only so that test-tmt integration in CI does not double in cost.
    • Image on the reinstall path. I recommend localhost/bootc, the same as bcvk, over a Testing Farm-style published-base + RPM build.
    • Transport. I recommend a tarball in the tmt tree over a local registry (which needs a registry service and insecure-registry configuration in the guest) or an in-guest build.
    • Reboot mechanism. I recommend trying tmt-reboot in prepare first, and using the local playbook plus ansible-core only if that fails.
    • Context value. I recommend a new running_env=reinstall rather than reusing packit, which would pull in the SRPM/test-artifacts phases.
    • Ordering. I recommend landing the fix (step 3) before making the GHA job (step 4) required. Otherwise run it with continue-on-error until the fix lands.

    Generated-by: https://github.com/cgwalters/#llms

  9. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    Review by an LLM (run 37834711147), requested by the coordinator. Not a human review.

    LLM-generated by praxis/gpt-6.1-sol

    Verdict: #2560 appears to implement the requested per-unit journal diagnostics for the failures this test selects, but I could not verify successful collection with synthetic failures. Level 1 is complete; level 2 was blocked provisioning a systemd container, so neither the failing/clean runtime cases nor the real tmt path were run.

    Source behavior (two sentences): When selected units have failed, the script prints their systemctl rows, runs journalctl --boot --unit $unit --no-pager separately for each, prints nonempty stdout/stderr and any nonzero journalctl status, then asserts zero failures. It prints journals only for failed unit names starting with systemd- or bootc-, excluding systemd-tpm2-setup.service when /sys/fs/selinux/enforce exists; plain synthetic-fail-a.service would be ignored.

    Commands actually run:

    git fetch https://github.com/bootc-dev/bootc pull/2560/head
    git worktree add --detach /home/runner-sandbox/bootc-pr2560 FETCH_HEAD
    git diff HEAD FETCH_HEAD -- tmt/tests/booted/readonly/014-test-no-failed-units.nu
    podman --version && podman info
    podman images --format '{{.Repository}}:{{.Tag}}' && command -v nu qemu-system-x86_64 virsh bcvk just; test -e /dev/kvm
    podman run -d --name bootc-pr2560-systemd --systemd=always --network=none quay.io/centos/centos:stream10 /sbin/init
    podman build -t localhost/bootc-pr2560-systemd -f /home/runner-sandbox/Containerfile.pr2560 /home/runner-sandbox
    

    The build was retried once with Python subprocess/monotonic timing: exit 1, 0.411 s. Scratch Containerfile: CentOS Stream 10 base; dnf -y install systemd curl tar && dnf clean all; fetch/extract Nushell 0.103.0 from the same release URL used by hack/provision-fetch.sh; /sbin/init as CMD. The build never reached the Nushell download.

    Relevant output:

    HEAD is now at abb4076 tmt: Print journals for failed units
    podman version 5.8.2
    rootless: true
    Error: crun: executable file `/sbin/init` not found: No such file or directory
    SSL certificate problem: unable to get local issuer certificate
    Error: Failed to download metadata for repo 'baseos': Cannot prepare internal mirrorlist
    Error: building at STEP "RUN dnf -y install systemd curl tar && dnf clean all": while running runtime: exit status 1
    

    Read the changed test, tap.nu (begin prints TAP version/description, ok prints ok), hack/provision-fetch.sh (CentOS 10 pins Nu 0.103.0), and the Justfile. No locally installed Nu, QEMU, virsh, bcvk, or just was found by that command, and /dev/kvm was absent. I did not disable TLS verification or change the host systemd manager. Scratch worktree/container were removed; the original checkout is unchanged.

    Not run: two selected synthetic failed units, verification of each unit's journal and a recognizable marker, the clean case, or just test-tmt readonly (the documented real path). Per the requested stop-at-first-unreachable-level rule, I stopped at level 2 after the systemd-container dependency installation failed TLS validation; there are no synthetic-test output/artifact claims.

    Edge cases below are source inspection, not runtime observations:

    • No journal lines: the heading still prints; any nonempty journalctl placeholder/stderr is printed, and the failed-unit assertion remains.
    • Escaped unit names: $unit is passed as a separate argument without shell interpolation; actual systemd escaping was not exercised.
    • Long journals: no line/byte limit; complete captures the entire current-boot journal per unit before printing it.
    • Failures before journald is available: no retry or special recovery; only available current-boot entries can be printed.
    • Discovery: the exact command is systemctl list-units --failed --no-legend --plain. --failed is the failed-state filter (equivalent in intent to --state=failed for list-units); this remains a loaded-unit query, not a scan of unit files or historic failures. No runtime comparison was made.

    Suggestions for the pull request's author:

    • Use systemd-synthetic-fail-a and bootc-synthetic-fail-b for the synthetic check so both pass the existing prefix filter; have the latter print a recognizable marker before exiting nonzero.
    • Consider an explicit journal size limit if unbounded diagnostic output is undesirable.
    • Exercise an escaped unit name and an unavailable/empty journal alongside the failing and clean cases on a working systemd/Nu 0.103.0 test system.

    Generated-by: https://github.com/cgwalters/#llms

  10. cgwalters-bot commented on Oct 8, 2026

    @cgwalters-bot
    Contributor

    Synthetic-failure check of #2560, run on a devspace (LLM-generated, Sonnet; it replaces the earlier report here, which could only read the source).

    bootc#2560 verification (head abb4076 "tmt: Print journals for failed units")

    Verdict: The change works as intended. With synthetic failing bootc-/systemd- units the test exits 1, names each unit and prints its journal, including each canary line. It ignores a non-prefixed failed unit and passes silently once the units are reset. The same output appears in a real just test-tmt readonly VM run, in tmt's log.txt, output.txt and failures.yaml, with the other readonly tests unaffected.

    Deviation: bot-devspace start accepts only 4/16/64 cores, so I used 16 (not 8), AMD EPYC 9V45, RHEL 10 devspace with systemd as PID 1. Devspace deleted at the end.

    Step 1: devspace host, nushell 0.103.0 (the version in hack/provision-fetch.sh), as root

    git fetch https://github.com/bootc-dev/bootc pull/2560/head; git checkout --detach FETCH_HEAD
    bot-devspace start --cores 16 --duration 120 verify-2560   # then pushed HEAD to src/bootc
    sudo systemd-run --unit=bootc-synthetic-fail-a sh -c 'echo SYNTHETIC-CANARY-A; exit 1'
    sudo systemd-run --unit=bootc-synthetic-fail-b sh -c 'echo SYNTHETIC-CANARY-B; echo SYNTHETIC-CANARY-B2 >&2; exit 1'
    sudo systemd-run --unit=systemd-synthetic-silent sh -c 'exec >/dev/null 2>&1; exit 1'
    sudo systemd-run --unit=synthetic-ignored sh -c 'echo SYNTHETIC-CANARY-IGNORED; exit 1'
    sudo ~/nu/nu tmt/tests/booted/readonly/014-test-no-failed-units.nu     # exit=1
    sudo systemctl reset-failed 'bootc-synthetic-*' 'systemd-synthetic-*' synthetic-ignored.service
    sudo ~/nu/nu .../014-test-no-failed-units.nu                            # exit=0
    

    Failing case: exit 1, both bootc-synthetic-fail-a/b.service listed, SYNTHETIC-CANARY-A, -B and -B2 (stderr) present, synthetic-ignored and its canary absent. Also listed: the silent systemd-synthetic-silent.service. Pre-existing failed units mcelog and temp-disk-dataloss-warning were correctly not selected.

    Journal for bootc-synthetic-fail-a.service:
    Oct 08 20:01:18 runnervm1fo5c sh[30293]: SYNTHETIC-CANARY-A
    Error:   x Expected zero failed systemd-*/bootc-* units
    

    Clean case: No failed systemd-*/bootc-* units, ok, exit 0, no journal output.

    Step 2: real path

    # scratch copy src/bootc-inject on the devspace only; added readonly/009-test-inject-failure.nu:
    #   ^systemd-run --unit=bootc-synthetic-fail-a sh -c 'echo SYNTHETIC-CANARY-A; exit 1'; ^sleep 2
    just build; just test-tmt readonly        # exit 1 (expected)
    

    In tmt's stream, the test's output.txt and failures.yaml under .../execute/data/guest/default-0/tmt/tests/tests/test-01-readonly-1/, and log.txt in the plan dir (all contain CANARY):

    Running booted/readonly/014-test-no-failed-units.nu...
    Journal for bootc-synthetic-fail-a.service:
    ... sh[1656]: SYNTHETIC-CANARY-A
    ... bootc-synthetic-fail-a.service: Failed with result 'exit-code'.
    Left : '1'   Right : '0'
    

    Edge cases

    • No journal lines: not reproducible for a unit that actually ran, as systemd always logs "Started"/"Failed with result" (silent unit above still shows 3 lines). If journalctl matched nothing, the script prints nothing for that unit, with no "no journal entries" note.
    • Unbounded length: no -n/--lines cap. A crash-looping unit (Restart=) could print thousands of lines into the TAP/tmt output. Not exercised.
    • Failing before journald: not run. journalctl would then return only what journald later received from the early stdout/kmsg, possibly empty. --boot is used, so it is limited to the current boot.
    • Filter limits (pre-existing): only systemd-* and bootc-* units are considered.

    Not run

    • Other tmt plans, sealed UKI/composefs variants (ostree/grub default only).
    • The "units failing before journald" and long-journal cases.
    • The tmt run used a scratch injection test file on the devspace (not in any repo).

    Suggestions for the PR author

    • Cap output, e.g. journalctl --boot --unit $unit --no-pager -n 100.
    • Print a "(no journal entries)" note when stdout is empty, so an empty journal is distinguishable from a skipped unit.
    • Consider -o short-precise or --output=cat if timestamps/hostnames are noise in tmt logs.
    • Consider also printing systemctl status --no-pager $unit for units that failed before journald was up.

    Generated-by: https://github.com/cgwalters/#llms

  11. jeckersb commented on Oct 8, 2026

    @jeckersb
    Collaborator

    I added some context in #2560, the tl;dr is that %preun for netcat fails because it tries to call alternatives to remove itself, but that fails. I think the "fix" here is to just pass --noscripts to the "remove the world" rpm -e invocation in the fedora-bootc-destructive-cleanup script.

  12. added a commit that references this issue on Oct 8, 2026
    22df875
  13. jeckersb commented on Oct 8, 2026

    @jeckersb
    Collaborator

    Ok so the cause here is actually a bit more interesting.

    This is on a plain old f44 cloud image that I got from testing-farm via testing-farm reserve --compose Fedora-44 --arch x86_64 --duration 60.

    The problem is that the system has /var in a separate btrfs subvolume:

    [root@ip-172-31-24-122 ~]# findmnt -no SOURCE,FSTYPE,OPTIONS /var
    /dev/nvme0n1p3[/var] btrfs rw,relatime,seclabel,compress=zstd:1,ssd,space_cache=v2,subvolid=259,subvol=/var
    [root@ip-172-31-24-122 ~]# btrfs subvolume list /
    ID 256 gen 59 top level 5 path root
    ID 257 gen 45 top level 5 path boot
    ID 258 gen 51 top level 5 path home
    ID 259 gen 61 top level 5 path var
    

    After we reboot into the reinstalled system, we mount the root subvolume at /sysroot, but we don't do anything with the var subvolume. Critically, that means we don't have /var/lib/alternatives from the old system, which is why the %preun script fails running alternatives; the expected state is missing completely.

  14. cgwalters commented on Oct 8, 2026

    @cgwalters
    CollaboratorAuthor

    Ah yeah that is interesting. And actually kind of complex to handle in the general case; we'd need to mount the volume only to have rpm delete the content on it.

    On a btrfs system like that we could also implement deletion by just creating new subvolumes and pruning the old ones and ignoring RPM e.g. entirely, but that would split our setup.

    Hmm, I think probably what we need to do in the general case here is discover the mounts needed during install and pass them through as e.g. /var/lib/bootc-destructive-cleanup-state.json or so and then our cleanup service mounts them (in a new mountns, not global) and uses that? This gets complicated fast and also a bit more dangerous.

    It'd be safer to special case known setups like Fedora Cloud, but also wouldn't solve the problem for any cloud images that might use LVM etc - plus the more general case of arbitrary OS installs (though those are much more likely to want to preserve state).

    I'm kind of thinking here that one thing that would be helpful here too is to push down into the OS streams this cleanup tooling; then images which are intended to support reprovisioning in this way can support their own "deprovisioning" scripts.

    I think the special case mount preservation code but only handling the Fedora cloud case to start would be pretty tractable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions