downstream tmt test failure #2559
Description
Activity
@cgwalters-bot take this at p0 - quick PR to add the logging ensure successful log of journal of each failed services w/synthetic failure.
I think problem here is we are only testing the system-reinstall-bootc path in testing farm...probably most effective path here actually...can we setup a variant of our test suite that doesn't use bcvk but uses testcloud to boot and then does the in-place reprovision?
If we do that even for the readonly test and can reproduce the porblem that'd be great. basically
just tmt-readonlyshould run both provisioning paths.Look carefully at surrounding code and history create a plan don't make mistakes
On it at P0, as two runs:
- The journal capture for failed tests, with a synthetic failing unit as proof: run 37833703108. It comes back as a pull request for your approval.
- The plan for a testcloud boot plus in-place reprovision variant, so
just tmt-readonlycovers both provisioning paths: run 37833802970, read-only, checking the code and history first. The plan will be posted here.
Generated-by: https://github.com/cgwalters/#llms
@cgwalters-bot ah I had an agent working on this during our meeting, here's the PR: #2560
Seen, thanks: #2560 covers the journal output for failed units, so I've stopped my run for that half rather than open a second pull request for the same thing. The read-only planning run for the testcloud plus in-place reprovision variant continues and will post here.
@cgwalters: you asked for proof with a synthetic failure. If you want it, I can run #2560 against a deliberately failing unit and report what the test output contains; say the word.
Generated-by: https://github.com/cgwalters/#llms
@cgwalters: you asked for proof with a synthetic failure. If you want it, I can run #2560 against a deliberately failing unit and report what the test output contains; say the word.
@cgwalters-bot yes do that
Started: run 37834711147 checks #2560 against synthetic failing units and will report here what the test output contains.
Generated-by: https://github.com/cgwalters/#llms
Review by an LLM (run 37834711147), requested by the coordinator. Not a human review.
LLM-generated by praxis/gpt-6.1-sol
Verdict: #2560 appears to implement the requested per-unit journal diagnostics for the failures this test selects, but I could not verify successful collection with synthetic failures. Level 1 is complete; level 2 was blocked provisioning a systemd container, so neither the failing/clean runtime cases nor the real tmt path were run.
Source behavior (two sentences): When selected units have failed, the script prints their systemctl rows, runs
journalctl --boot --unit $unit --no-pagerseparately for each, prints nonempty stdout/stderr and any nonzero journalctl status, then asserts zero failures. It prints journals only for failed unit names starting withsystemd-orbootc-, excludingsystemd-tpm2-setup.servicewhen/sys/fs/selinux/enforceexists; plainsynthetic-fail-a.servicewould be ignored.Commands actually run:
git fetch https://github.com/bootc-dev/bootc pull/2560/head git worktree add --detach /home/runner-sandbox/bootc-pr2560 FETCH_HEAD git diff HEAD FETCH_HEAD -- tmt/tests/booted/readonly/014-test-no-failed-units.nu podman --version && podman info podman images --format '{{.Repository}}:{{.Tag}}' && command -v nu qemu-system-x86_64 virsh bcvk just; test -e /dev/kvm podman run -d --name bootc-pr2560-systemd --systemd=always --network=none quay.io/centos/centos:stream10 /sbin/init podman build -t localhost/bootc-pr2560-systemd -f /home/runner-sandbox/Containerfile.pr2560 /home/runner-sandboxThe build was retried once with Python subprocess/monotonic timing: exit 1, 0.411 s. Scratch Containerfile: CentOS Stream 10 base;
dnf -y install systemd curl tar && dnf clean all; fetch/extract Nushell 0.103.0 from the same release URL used byhack/provision-fetch.sh;/sbin/initas CMD. The build never reached the Nushell download.Relevant output:
HEAD is now at abb4076 tmt: Print journals for failed units podman version 5.8.2 rootless: true Error: crun: executable file `/sbin/init` not found: No such file or directory SSL certificate problem: unable to get local issuer certificate Error: Failed to download metadata for repo 'baseos': Cannot prepare internal mirrorlist Error: building at STEP "RUN dnf -y install systemd curl tar && dnf clean all": while running runtime: exit status 1Read the changed test,
tap.nu(begin prints TAP version/description, ok printsok),hack/provision-fetch.sh(CentOS 10 pins Nu 0.103.0), and the Justfile. No locally installed Nu, QEMU, virsh, bcvk, or just was found by that command, and/dev/kvmwas absent. I did not disable TLS verification or change the host systemd manager. Scratch worktree/container were removed; the original checkout is unchanged.Not run: two selected synthetic failed units, verification of each unit's journal and a recognizable marker, the clean case, or
just test-tmt readonly(the documented real path). Per the requested stop-at-first-unreachable-level rule, I stopped at level 2 after the systemd-container dependency installation failed TLS validation; there are no synthetic-test output/artifact claims.Edge cases below are source inspection, not runtime observations:
- No journal lines: the heading still prints; any nonempty journalctl placeholder/stderr is printed, and the failed-unit assertion remains.
- Escaped unit names:
$unitis passed as a separate argument without shell interpolation; actual systemd escaping was not exercised. - Long journals: no line/byte limit;
completecaptures the entire current-boot journal per unit before printing it. - Failures before journald is available: no retry or special recovery; only available current-boot entries can be printed.
- Discovery: the exact command is
systemctl list-units --failed --no-legend --plain.--failedis the failed-state filter (equivalent in intent to--state=failedfor list-units); this remains a loaded-unit query, not a scan of unit files or historic failures. No runtime comparison was made.
Suggestions for the pull request's author:
- Use
systemd-synthetic-fail-aandbootc-synthetic-fail-bfor the synthetic check so both pass the existing prefix filter; have the latter print a recognizable marker before exiting nonzero. - Consider an explicit journal size limit if unbounded diagnostic output is undesirable.
- Exercise an escaped unit name and an unavailable/empty journal alongside the failing and clean cases on a working systemd/Nu 0.103.0 test system.
Generated-by: https://github.com/cgwalters/#llms
Review by an LLM (run 37833802970), requested by the coordinator. Not a human review.
LLM-generated by Opus 5.5. This is an analysis and plan only; nothing was changed. Citations are to
mainat 7702842 unless a commit is named. Anything marked (inferred) was not verified by running it.Summary
Your hypothesis holds. The failing check in the linked run is a unit that only exists on the system-reinstall-bootc path. Testing Farm runs
bootc install to-existing-root --cleanupfrom a stock cloud image. That enablesbootc-destructive-cleanup.service, and in the linked run that service isfailed. The local flow (bcvk) and the GHA flow (also bcvk) never take that path. GHA does run system-reinstall-bootc, but it never reboots into the result. The cheapest reproducer is a second provisioning mode incargo xtask run-tmt: testcloud boots a stock cloud image, then the image under test is reinstalled in place. It uses the same plans and selects the mode with a tmt context value.1. How the local tmt targets work today
- There is no
just tmt-readonly. The closest isjust test-tmt readonly(Justfile:191-194). It runsbuild(podman build of the repoDockerfilefrom the RPMs made byjust package, givinglocalhost/bootc), then_build-upgrade-image, thentest-tmt-nobuild. That last one iscargo xtask run-tmt ... --upgrade-image=localhost/bootc-upgrade localhost/bootc readonly(Justfile:298-299). run_tmt(crates/xtask/src/tmt.rs:371 onward):- It requires bcvk, tmt, rsync and podman (tmt.rs:86-91).
- It sets the contexts
running_env=image_modeanddistro=<ID-VERSION_ID of the image>(tmt.rs:384-394). - It rsyncs
.fmfandtmt/intotarget/tmt-workdir(a workaround for Still fixrsyncto use.gitignoreteemtee/tmt#4062) and lists plans withtmt plan ls --filter enabled:true(tmt.rs:445). - For each plan it boots a fresh VM with
bcvk libvirt run ... [--bind-storage-ro] [--log-dir=journal,console=...] <image>(tmt.rs:615) and gets the SSH port and key frombcvk libvirt inspect. - It then runs
tmt run --all -e TMT_SCRIPTS_DIR=... -e BCVK_EXPORT=1 provision --how=connect --guest=localhost ... plan --name <plan>(tmt.rs:743-752).
- bcvk boots the bootc image directly, so the root is laid out by
bootc install to-disk. (Inferred from bcvk's role and from the comment at tmt/tests/booted/readonly/052-test-bli-detection.nu:30-34: "In Packit/gating CI the system was installed via to-existing-root".) - Plans:
cargo xtask update-generatedwrites everything between theBEGIN/END GENERATED PLANSmarkers in tmt/plans/integration.fmf from the metadata headers in tmt/tests/booted/*.nu (tmt.rs:1132);just validatechecks it. The header above the markers is hand-written:provision: how: virtual, image: $@{test_disk_image}(integration.fmf:1-4). Locally this is overridden byconnect, and the context it reads is a leftover of the testcloud era (see section 4).- Three prepare phases, all
when: running_env != image_mode(integration.fmf:26,34,41). Locally they are therefore off.
- GHA (
.github/workflows/ci.yml):- test-integration runs
just test-tmt integrationorjust test-composefs ...(ci.yml:406-411). - test-upgrade boots the published base under bcvk and upgrades (Justfile:243-262).
- tmt is pinned at 1.79.0 with
[provision-virtual](.github/actions/install-tmt/action.yml:18-19), so testcloud is already installed on the runners.
- test-integration runs
2. What Testing Farm runs
- Upstream Packit (.packit.yaml:84-127, label-gated since 44ce0b7):
- Plans:
tmt_plan: /tmt/plans/integration, which is the same generated plans as the local flow, with contextrunning_env: packit. ci/tier-1runs centos-stream-10 and fedora-latest-stable on x86_64 and aarch64.ci/mergeadds c9s and fedora-all.
- Plans:
- Downstream (the linked run):
- It is a Fedora CI run (
initiator: fedora-ci) of Koji task 151003669, bootc-1.17.1-1.fc44, Fedora-44 x86_64. - Its plans come from
plans/all.fmfin src.fedoraproject.org/forks/packit/rpms/bootc @ 06c498ef, not from this repo. - Those plans have the same prepare phases plus one extra line,
cp os-image-map.json bootc/hack, and norunning_envcontext. The phases still ran, sorunning_env != image_modeevaluates true when the dimension is undefined.
- It is a Fedora CI run (
- Provisioning:
- Testing Farm brings its own machine. In this run tmt was invoked with
provision --how connect --guest <EC2 address> --hard-reboot 'artemis-cli guest reboot ...'(per-plan log.txt). The plan'show: virtualis not used. - The in-place reprovision is done by the prepare phases, added in 704338d and adjusted in f1dec83. They:
- install podman, bootc, system-reinstall-bootc, expect and ansible-core from the build under test;
- unpack the SRPM and run hack/provision-packit.sh. That script builds
localhost/bootcFROM the published base for the OS (os-image-map.json; provision-packit.sh:38,89). It uses hack/Containerfile.packit, which runsdnf -y update bootcfrom the test-artifacts repo andprovision-derived.sh cloudinit(Containerfile.packit:29,32); - pull the LBI images (provision-packit.sh:92);
- run hack/system-reinstall-bootc.exp, i.e.
system-reinstall-bootc localhost/bootc. In the artifact this executedpodman run ... localhost/bootc bootc install to-existing-root --acknowledge-destructive --skip-fetch-check --cleanup --root-ssh-authorized-keys ...; - reboot through the Ansible playbook hack/packit-reboot.yml, which is fetched by URL from GitHub
main(integration.fmf:39-40), not from the tree under test.
- Testing Farm brings its own machine. In this run tmt was invoked with
- What differs from the local path:
to-existing-rootinstead of bcvk's disk install. The old OS stays under/sysroot.--cleanup(crates/system-reinstall-bootc/src/podman.rs:83-88) writes theetc/bootc-destructive-cleanupstamp (crates/lib/src/install.rs:224). The generator then enablesbootc-destructive-cleanup.service(crates/lib/src/generator.rs:81-83), which runs contrib/scripts/fedora-bootc-destructive-cleanup: itrpm -es every package of the old root, thenpodman system prune --all -f.- The image is the published base plus the RPM, not the repo
Dockerfileimage. - There is no
BCVK_EXPORTand no bind storage. Readonly tests already skip on that (017, 052). - Cloud hardware (EC2 here), and aarch64 as well.
- Nothing upstream boots a reinstalled system. The GHA
install-testsjob runsbootc-integration-tests system-reinstallon the Ubuntu host (ci.yml:227). That test only checks that the cleanup files are present (crates/tests-integration/src/system_reinstall.rs:92-99) and never reboots. So the cleanup unit has never run in upstream CI outside Testing Farm.
3. The linked artifacts
- 9 plans ran. Only
plans/all/plan-01-readonlyfailed; the other 8 passed. - Prepare for readonly succeeded:
- image built FROM
quay.io/fedora/fedora-bootc:44(23:03:35 to 23:09:21); - install finished at 23:12:47 ("Installation complete!");
- Ansible reboot ok at 23:13:39.
- image built FROM
- Execute: subtests 000-013 pass.
014-test-no-failed-units.nufails with:
bootc-destructive-cleanup.service loaded failed failed Cleanup previous the installation after an alongside installation
The AVC check has no matches. SELinux is enforcing, with selinux-policy-44.11-1.fc44. - 014 was added on 2026-09-18 in 41049cc. Before that, the readonly suite did not look at failed units, so a failing cleanup unit would have gone unreported. (Inferred: I can't tell from here since when it fails.)
- The other plans also reinstalled with
--cleanup, but none of them checks failed units. Whether the unit failed there too is unknown. - What is missing: the unit's journal (
journalctl -b -u bootc-destructive-cleanup.service, which would include the script'sset -xtrace), its exit status (systemctl status), and the state of/sysrootafterwards. Nothing in the artifacts says which command in the script failed. That is part 1 of this issue (journal capture).
4. testcloud
-
What tmt 1.79.0 needs:
tmt[provision-virtual]requirestestcloud>=0.11.7(PyPI metadata; latest is 0.12.0). testcloud then needs libvirt and its Python bindings, plus qemu with/dev/kvm. Without KVM it falls back to TCG with multiplied timeouts (tmt/steps/provision/testcloud.py:833-835). -
Defaults:
qemu:///session(testcloud.py:176-185), 2048 MB of memory and a 40 GB disk (testcloud.py:176-185). -
Images: an alias such as
fedora-44orcentos-stream-10(testcloud util.py:146-149), a URL, or an absolute qcow2 path (testcloud.py:875-905). -
Networking and SSH: session mode uses user-mode networking with a host forward to guest port 22 (testcloud domain_configuration.py:283). tmt generates the SSH key and injects it with cloud-init, so the image needs cloud-init (stock cloud images have it).
-
Logs: console only (
console.txt, testcloud.py:114). There is no journal capture. -
What tmt does not expose: no virtiofs option, so nothing like bcvk's
--bind-storage-ro. -
UEFI: tmt 1.79.0 supports
hardware: boot: method: uefiandtpmfor testcloud (testcloud.py:218,473-578), and tests-install.fmf:8-12 already uses it. (Not exercised here.) -
Compared with bcvk: bcvk boots the bootc image itself, returns its own SSH key, mounts host container storage over virtiofs, and captures the journal and console (bec34ca).
-
History: the repo used testcloud from 36be46a (2024-06-10) to f8ce015 (2025-11-04), but to boot a prebuilt bootc disk image, never to reinstall. tests/run-tmt.sh passed
--context test_disk_image=target/bootc-integration-test.qcow2(seegit show f8ce0152^:tests/run-tmt.sh). It needed several workarounds:- pruning the testcloud cache (cc42685, testcloud#17);
- a
chconshim on Ubuntu (testcloud#18, removed in f8ce015); - the tmt-workdir copy (tmt#4062).
f8ce015 dropped testcloud because "testcloud+tmt doesn't support UEFI (Support configuring UEFI on x86_64 in
testcloudprovision teemtee/tmt#4203) ... a blocker for UKIs", and because of bcvk's ergonomics. tests-install.fmf:5-7 also records SELinux being disabled on Ubuntu (tmt#3364). -
Machine for this analysis: it had
/dev/kvmbut no libvirt, tmt or bcvk, so I ran nothing.
Plan
Design (smallest that reuses every plan):
-
Selection: a flag
cargo xtask run-tmt --provision=bcvk|reinstall, defaulting to bcvk so today's behaviour is unchanged. Inreinstallmode,run_tmt:- sets
--context running_env=reinstallinstead ofimage_mode; - does not set
BCVK_EXPORT; - skips bcvk;
- runs
tmt run --all ... provision --how virtual --image <alias> --memory 4096 plan --name <plan>per plan. The alias comes from the detected distro:centos-Nbecomescentos-stream-N, andfedora-Nstaysfedora-N.
- sets
-
Transport: before the loop,
podman savethe image under test into the tmt tree. This is the same trick the CoreOS path already uses (tmt.rs:422-431, loaded in tmt/tests/install/test-on-ostree/test-install-on-ostree.sh:30-33). Copytarget/packages/*.rpmnext to it, so the stock guest gets this build's system-reinstall-bootc. The repoDockerfilealready installs that RPM (Dockerfile:86-90). -
Plans: add one hand-written prepare phase to the integration.fmf header,
when: running_env == reinstall. It runs a new small script, hack/provision-reinstall-local.sh, which:- dnf installs the RPMs from the tree;
- runs
podman load; - pulls the LBIs (the same line as provision-packit.sh:92);
- runs the existing hack/system-reinstall-bootc.exp;
- reboots.
The three Testing Farm phases change to
when: running_env != image_mode and running_env != reinstall. Generated plans stay untouched. This does not affect downstream Fedora CI, which has its own copy of the header. (Inferred from the extracp os-image-map.jsonline.) -
Justfile:
test-tmt-reinstall *ARGS: buildcalls run-tmt with--provision=reinstall;test-tmt-nobuild --provision=reinstall readonlyalready works, because ARGS are passed through;- a readonly target runs both paths (name: see decisions).
-
CI: Testing Farm stays as it is. Add a GHA leg for
readonlyonly on the reinstall path, ostree only.
Risks and unknowns:
- Reproduction: the failure was seen on Fedora 44 on EC2. The local default base is c10s, so the first try should use
BOOTC_base=quay.io/fedora/fedora-bootc:44with thefedora-44cloud image. The failure may depend on the cloud image contents. - The image differs from Testing Farm's (repo
Dockerfilevs published base + RPM). This is deliberate, so that only the provisioning path changes. If the failure does not reproduce, a Containerfile.packit-style build is the fallback. - Reboot in prepare: integration.fmf:35-36 says
tmt-rebootdoes not work in prepare, but that was written in 2025-09. tmt 1.79.0's prepare/shell.py handles reboots (lines 139,314). (Unverified.) The fallback is a local copy of packit-reboot.yml, butansible-playbookruns on the control node (tmt guest/init.py:2372), which means adding ansible-core to the install-tmt action. - tmt
whenwith an undefined dimension: the compound condition must still be true for downstream, which has norunning_env. Check withtmt plan showboth with no context and withrunning_env=packit. - testcloud on Ubuntu GHA runners: the old SELinux/chcon problems may come back. Cloud images are fetched from mirrors (flake risk).
- Cost: each plan reinstalls. For readonly, Testing Farm went from the start of the in-guest build to the end of the reboot in about 10 min (23:03:35 to 23:13:39). About 6 min of that was the build, which the local path skips. Every plan also copies a tarball of about the image's size. This is why the reinstall path should be scoped to readonly at first.
- Arch and variants: testcloud runs only the native arch locally, so aarch64 remains Testing Farm only. composefs/UKI on the reinstall path is out of scope initially.
- Diagnosis: testcloud gives only the console log. The journal needs part 1, or a finish step that dumps
journalctlfor failed units on both paths.
Steps (one PR each):
- Reproducer. Adds
--provision=reinstallin xtask, the header phase, the script andjust test-tmt-reinstall.
Acceptance:- on a KVM host,
BOOTC_base=quay.io/fedora/fedora-bootc:44 just test-tmt-reinstall readonlyfails in 014 withbootc-destructive-cleanup.servicefailed, and the guest can be kept forjournalctl; just test-tmt readonlybehaves as before;cargo xtask update-generated direct --checkpasses;tmt plan showstill lists the Testing Farm phases with no context and withrunning_env=packit.
- on a KVM host,
- Both paths in one target. A readonly target runs bcvk and then reinstall, with one summary. Acceptance: both runs appear and the exit status reflects both.
- Fix the cleanup failure from the journal obtained in step 1, keeping 014 as the regression check. Acceptance: step 1's command passes, and a
ci/tier-1Testing Farm run is green. - CI. A GHA job runs reinstall readonly on centos-10 and fedora-latest, x86_64. Acceptance: green on a labelled PR, and red on a throwaway commit that makes the cleanup script
exit 1. - (Optional) Widen the reinstall path to more plans on a schedule, and/or add a Testing Farm-faithful image mode.
Decisions for the maintainer
- Target names.
just tmt-readonlydoesn't exist. I recommendtest-tmt-readonly, running both paths, plustest-tmt-reinstall, and leavingtest-tmtbcvk-only so thattest-tmt integrationin CI does not double in cost. - Image on the reinstall path. I recommend
localhost/bootc, the same as bcvk, over a Testing Farm-style published-base + RPM build. - Transport. I recommend a tarball in the tmt tree over a local registry (which needs a registry service and insecure-registry configuration in the guest) or an in-guest build.
- Reboot mechanism. I recommend trying
tmt-rebootin prepare first, and using the local playbook plus ansible-core only if that fails. - Context value. I recommend a new
running_env=reinstallrather than reusingpackit, which would pull in the SRPM/test-artifacts phases. - Ordering. I recommend landing the fix (step 3) before making the GHA job (step 4) required. Otherwise run it with
continue-on-erroruntil the fix lands.
Generated-by: https://github.com/cgwalters/#llms
- There is no
Review by an LLM (run 37834711147), requested by the coordinator. Not a human review.
LLM-generated by praxis/gpt-6.1-sol
Verdict: #2560 appears to implement the requested per-unit journal diagnostics for the failures this test selects, but I could not verify successful collection with synthetic failures. Level 1 is complete; level 2 was blocked provisioning a systemd container, so neither the failing/clean runtime cases nor the real tmt path were run.
Source behavior (two sentences): When selected units have failed, the script prints their systemctl rows, runs
journalctl --boot --unit $unit --no-pagerseparately for each, prints nonempty stdout/stderr and any nonzero journalctl status, then asserts zero failures. It prints journals only for failed unit names starting withsystemd-orbootc-, excludingsystemd-tpm2-setup.servicewhen/sys/fs/selinux/enforceexists; plainsynthetic-fail-a.servicewould be ignored.Commands actually run:
git fetch https://github.com/bootc-dev/bootc pull/2560/head git worktree add --detach /home/runner-sandbox/bootc-pr2560 FETCH_HEAD git diff HEAD FETCH_HEAD -- tmt/tests/booted/readonly/014-test-no-failed-units.nu podman --version && podman info podman images --format '{{.Repository}}:{{.Tag}}' && command -v nu qemu-system-x86_64 virsh bcvk just; test -e /dev/kvm podman run -d --name bootc-pr2560-systemd --systemd=always --network=none quay.io/centos/centos:stream10 /sbin/init podman build -t localhost/bootc-pr2560-systemd -f /home/runner-sandbox/Containerfile.pr2560 /home/runner-sandboxThe build was retried once with Python subprocess/monotonic timing: exit 1, 0.411 s. Scratch Containerfile: CentOS Stream 10 base;
dnf -y install systemd curl tar && dnf clean all; fetch/extract Nushell 0.103.0 from the same release URL used byhack/provision-fetch.sh;/sbin/initas CMD. The build never reached the Nushell download.Relevant output:
HEAD is now at abb4076 tmt: Print journals for failed units podman version 5.8.2 rootless: true Error: crun: executable file `/sbin/init` not found: No such file or directory SSL certificate problem: unable to get local issuer certificate Error: Failed to download metadata for repo 'baseos': Cannot prepare internal mirrorlist Error: building at STEP "RUN dnf -y install systemd curl tar && dnf clean all": while running runtime: exit status 1Read the changed test,
tap.nu(begin prints TAP version/description, ok printsok),hack/provision-fetch.sh(CentOS 10 pins Nu 0.103.0), and the Justfile. No locally installed Nu, QEMU, virsh, bcvk, or just was found by that command, and/dev/kvmwas absent. I did not disable TLS verification or change the host systemd manager. Scratch worktree/container were removed; the original checkout is unchanged.Not run: two selected synthetic failed units, verification of each unit's journal and a recognizable marker, the clean case, or
just test-tmt readonly(the documented real path). Per the requested stop-at-first-unreachable-level rule, I stopped at level 2 after the systemd-container dependency installation failed TLS validation; there are no synthetic-test output/artifact claims.Edge cases below are source inspection, not runtime observations:
- No journal lines: the heading still prints; any nonempty journalctl placeholder/stderr is printed, and the failed-unit assertion remains.
- Escaped unit names:
$unitis passed as a separate argument without shell interpolation; actual systemd escaping was not exercised. - Long journals: no line/byte limit;
completecaptures the entire current-boot journal per unit before printing it. - Failures before journald is available: no retry or special recovery; only available current-boot entries can be printed.
- Discovery: the exact command is
systemctl list-units --failed --no-legend --plain.--failedis the failed-state filter (equivalent in intent to--state=failedfor list-units); this remains a loaded-unit query, not a scan of unit files or historic failures. No runtime comparison was made.
Suggestions for the pull request's author:
- Use
systemd-synthetic-fail-aandbootc-synthetic-fail-bfor the synthetic check so both pass the existing prefix filter; have the latter print a recognizable marker before exiting nonzero. - Consider an explicit journal size limit if unbounded diagnostic output is undesirable.
- Exercise an escaped unit name and an unavailable/empty journal alongside the failing and clean cases on a working systemd/Nu 0.103.0 test system.
Generated-by: https://github.com/cgwalters/#llms
Synthetic-failure check of #2560, run on a devspace (LLM-generated, Sonnet; it replaces the earlier report here, which could only read the source).
bootc#2560 verification (head abb4076 "tmt: Print journals for failed units")
Verdict: The change works as intended. With synthetic failing
bootc-/systemd-units the test exits 1, names each unit and prints its journal, including each canary line. It ignores a non-prefixed failed unit and passes silently once the units are reset. The same output appears in a realjust test-tmt readonlyVM run, in tmt'slog.txt,output.txtandfailures.yaml, with the other readonly tests unaffected.Deviation:
bot-devspace startaccepts only 4/16/64 cores, so I used 16 (not 8), AMD EPYC 9V45, RHEL 10 devspace with systemd as PID 1. Devspace deleted at the end.Step 1: devspace host, nushell 0.103.0 (the version in hack/provision-fetch.sh), as root
git fetch https://github.com/bootc-dev/bootc pull/2560/head; git checkout --detach FETCH_HEAD bot-devspace start --cores 16 --duration 120 verify-2560 # then pushed HEAD to src/bootc sudo systemd-run --unit=bootc-synthetic-fail-a sh -c 'echo SYNTHETIC-CANARY-A; exit 1' sudo systemd-run --unit=bootc-synthetic-fail-b sh -c 'echo SYNTHETIC-CANARY-B; echo SYNTHETIC-CANARY-B2 >&2; exit 1' sudo systemd-run --unit=systemd-synthetic-silent sh -c 'exec >/dev/null 2>&1; exit 1' sudo systemd-run --unit=synthetic-ignored sh -c 'echo SYNTHETIC-CANARY-IGNORED; exit 1' sudo ~/nu/nu tmt/tests/booted/readonly/014-test-no-failed-units.nu # exit=1 sudo systemctl reset-failed 'bootc-synthetic-*' 'systemd-synthetic-*' synthetic-ignored.service sudo ~/nu/nu .../014-test-no-failed-units.nu # exit=0Failing case: exit 1, both
bootc-synthetic-fail-a/b.servicelisted,SYNTHETIC-CANARY-A,-Band-B2(stderr) present,synthetic-ignoredand its canary absent. Also listed: the silentsystemd-synthetic-silent.service. Pre-existing failed unitsmcelogandtemp-disk-dataloss-warningwere correctly not selected.Journal for bootc-synthetic-fail-a.service: Oct 08 20:01:18 runnervm1fo5c sh[30293]: SYNTHETIC-CANARY-A Error: x Expected zero failed systemd-*/bootc-* unitsClean case:
No failed systemd-*/bootc-* units,ok, exit 0, no journal output.Step 2: real path
# scratch copy src/bootc-inject on the devspace only; added readonly/009-test-inject-failure.nu: # ^systemd-run --unit=bootc-synthetic-fail-a sh -c 'echo SYNTHETIC-CANARY-A; exit 1'; ^sleep 2 just build; just test-tmt readonly # exit 1 (expected)In tmt's stream, the test's
output.txtandfailures.yamlunder.../execute/data/guest/default-0/tmt/tests/tests/test-01-readonly-1/, andlog.txtin the plan dir (all contain CANARY):Running booted/readonly/014-test-no-failed-units.nu... Journal for bootc-synthetic-fail-a.service: ... sh[1656]: SYNTHETIC-CANARY-A ... bootc-synthetic-fail-a.service: Failed with result 'exit-code'. Left : '1' Right : '0'Edge cases
- No journal lines: not reproducible for a unit that actually ran, as systemd always logs "Started"/"Failed with result" (silent unit above still shows 3 lines). If journalctl matched nothing, the script prints nothing for that unit, with no "no journal entries" note.
- Unbounded length: no
-n/--linescap. A crash-looping unit (Restart=) could print thousands of lines into the TAP/tmt output. Not exercised. - Failing before journald: not run. journalctl would then return only what journald later received from the early stdout/kmsg, possibly empty.
--bootis used, so it is limited to the current boot. - Filter limits (pre-existing): only
systemd-*andbootc-*units are considered.
Not run
- Other tmt plans, sealed UKI/composefs variants (ostree/grub default only).
- The "units failing before journald" and long-journal cases.
- The tmt run used a scratch injection test file on the devspace (not in any repo).
Suggestions for the PR author
- Cap output, e.g.
journalctl --boot --unit $unit --no-pager -n 100. - Print a "(no journal entries)" note when stdout is empty, so an empty journal is distinguishable from a skipped unit.
- Consider
-o short-preciseor--output=catif timestamps/hostnames are noise in tmt logs. - Consider also printing
systemctl status --no-pager $unitfor units that failed before journald was up.
Generated-by: https://github.com/cgwalters/#llms
I added some context in #2560, the tl;dr is that %preun for netcat fails because it tries to call
alternativesto remove itself, but that fails. I think the "fix" here is to just pass--noscriptsto the "remove the world"rpm -einvocation in thefedora-bootc-destructive-cleanupscript.- added a commit that references this issue
on Oct 8, 2026 Ok so the cause here is actually a bit more interesting.
This is on a plain old f44 cloud image that I got from testing-farm via
testing-farm reserve --compose Fedora-44 --arch x86_64 --duration 60.The problem is that the system has /var in a separate btrfs subvolume:
[root@ip-172-31-24-122 ~]# findmnt -no SOURCE,FSTYPE,OPTIONS /var /dev/nvme0n1p3[/var] btrfs rw,relatime,seclabel,compress=zstd:1,ssd,space_cache=v2,subvolid=259,subvol=/var [root@ip-172-31-24-122 ~]# btrfs subvolume list / ID 256 gen 59 top level 5 path root ID 257 gen 45 top level 5 path boot ID 258 gen 51 top level 5 path home ID 259 gen 61 top level 5 path varAfter we reboot into the reinstalled system, we mount the root subvolume at /sysroot, but we don't do anything with the var subvolume. Critically, that means we don't have /var/lib/alternatives from the old system, which is why the %preun script fails running
alternatives; the expected state is missing completely.Ah yeah that is interesting. And actually kind of complex to handle in the general case; we'd need to mount the volume only to have rpm delete the content on it.
On a btrfs system like that we could also implement deletion by just creating new subvolumes and pruning the old ones and ignoring RPM e.g. entirely, but that would split our setup.
Hmm, I think probably what we need to do in the general case here is discover the mounts needed during install and pass them through as e.g.
/var/lib/bootc-destructive-cleanup-state.jsonor so and then our cleanup service mounts them (in a new mountns, not global) and uses that? This gets complicated fast and also a bit more dangerous.It'd be safer to special case known setups like Fedora Cloud, but also wouldn't solve the problem for any cloud images that might use LVM etc - plus the more general case of arbitrary OS installs (though those are much more likely to want to preserve state).
I'm kind of thinking here that one thing that would be helpful here too is to push down into the OS streams this cleanup tooling; then images which are intended to support reprovisioning in this way can support their own "deprovisioning" scripts.
I think the special case mount preservation code but only handling the Fedora cloud case to start would be pretty tractable.
- added 3 commits that reference this issue
on Oct 10, 2026
https://artifacts.dev.testing-farm.io/eef0709a-b4b6-42a5-9932-70045f05fb87/
We need to capture the journal (surprising tmt doesn't do this by default)