Skip to content

docker: fix the cache directory, install security updates and only the vCenter bindings - #573

Merged
semx merged 8 commits into
developmentfrom
chore/slim-image
Oct 2, 2026
Merged

semx merged 8 commits into
developmentfrom
chore/slim-image

Conversation

@semx

@semx semx commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Changes to the image, all checked by running it against NetBox, not only by building it.

Cache directory. /app belongs to root and /app/cache does not exist in the image, so the service user cannot create it: every run logs Error writing to cache directory: /app/cache / NetBox caching DISABLED, even with a volume mounted on /app/cache (Docker creates that mount point as root). The code was also copied with the service user as owner, so the process could rewrite it. Now the code is copied as root and read-only for the service user, and /app/cache is created owned by the service user and group 0 with mode 0770 (group 0 covers platforms that run images with an arbitrary uid). The README example mounts a volume there.

Security updates. The final stage runs apt-get update && apt-get dist-upgrade -y, and pip is removed from the venv and from the base interpreter, since nothing uses it at runtime. Trivy on the current development image: 40 fixable OS findings (among them HIGH in openssl and pcre2 from trixie-security) and 6 in pip's vendored copies of urllib3, msgpack and setuptools. On this branch: 0 fixable.

Smaller, pinned SDK. vcf-sdk is a meta package that pulls every VMware Cloud Foundation binding (NSX, SDDC Manager, Operations, Fleet LCM, Installer, vSAN data protection). netbox-sync only uses create_vsphere_client and the tagging client from vmware-vcenter, which is now installed pinned to the version vcf-sdk resolves to today (9.1.1.0), so rebuilding a tag gives the same SDK. All 552 Python files under vmware/ and com/ are byte-identical to the current image; the 155 that are gone belong to SDDC Manager, the VCF installer and snapservice.

Image build trigger. Dockerfile and .dockerignore were not in the paths filter of docker-image.yml, so a change to them alone did not rebuild the development image.

current development image this PR
size on disk / download 339 MB / 61 MB 261 MB / 59 MB
Trivy, fixable findings 46 0
NetBox cache with a volume disabled every run 26 cache files after a run

Functional check under a rootless Docker daemon (Docker 29.1.3 on Ubuntu 26.04): NetBox 4.4.5, two vcsim instances replaying the captured vCenters from tests/fixtures/vcsim (vchvr 6.7 with 28 VMs, vc001 8.0.3 with 34 VMs) as two sources, netbox-sync run with --read-only --cap-drop ALL --security-opt no-new-privileges, settings mounted read-only, a volume on /app/cache, two passes each:

current image this PR
pass 1 exit 0, 0 errors, 1142 created, 253 updated the same
pass 2 exit 0, 0 errors, 0 created, 1 updated the same
NetBox after each pass 5 devices, 30 interfaces, 62 VMs, 96 VM interfaces, 2 clusters, 126 MAC addresses the same
warnings and errors identical line for line, apart from the two cache warnings

The docker run line from the README works as written. The arm64 build (QEMU, and native on Apple silicon) starts and imports the vSphere bindings. vCenter tag sync cannot be exercised against vcsim (it does not implement the vAPI endpoint the SDK talks to), which is why the SDK files are compared byte for byte instead.

Kubernetes, as Jobs on a kubeadm cluster with the security context from k8s-netbox-sync-cronjob.yaml: the image change makes no difference there except for an arbitrary uid in group 0 without a volume, see the comment below.

@semx
semx requested a review from bb-Ricardo as a code owner September 30, 2026 09:58
@bb-Ricardo

Copy link
Copy Markdown
Owner

Hi,

thank you for this improvement. I'm just curious if a cache directory is actually used inside the container and not as PVC. If the container runs as Job then the content of the cache folder will be purged after the Job finished.

And yes apt-get update by itself does nothing. But a apt-get update && apt-get dist-upgrade -y should install the latest security fixes for that distribution: https://pythonspeed.com/articles/security-updates-in-docker/

semx added 6 commits October 2, 2026 02:10
vcf-sdk is a meta package that pulls every VMware Cloud Foundation binding
(NSX, SDDC Manager, Operations, Fleet LCM, Installer, vSAN data protection).
netbox-sync only needs create_vsphere_client and the tagging client, which
come from vmware-vcenter and the vapi runtime and common client it depends
on. Install vmware-vcenter pinned to the version vcf-sdk resolves to today,
so a rebuild of the same tag gets the same SDK, and drop the apt-get line
that installed nothing.

Built from the same development commit: 339 MB -> 248 MB uncompressed,
58 MB -> 52 MB gzipped. All 552 Python files under vmware/ and com/ in the
new image are byte-identical to the current one; the 155 files that are no
longer installed belong to SDDC Manager, the VCF installer and snapservice.
pyvmomi, vmware-vapi-runtime, vmware-vapi-common-client and vmware-vcenter
stay at 9.1.1.0, --help works and the process still runs as uid 1000.
The application files were copied with the service user as owner, so the
running process could rewrite its own code, while /app itself stayed owned
by root, so the default cache directory (/app/cache) could not be created
and every container run logged "NetBox caching DISABLED" and fetched all
objects from NetBox again.

Copy the code as root, read-only for the service user, and create /app/cache
owned by the service user and group 0 with mode 0770, which also covers
platforms that run the image with an arbitrary uid in group 0.

Checked with --read-only --cap-drop ALL --security-opt no-new-privileges and
a volume on /app/cache: the cache is writable, the code is not, --help works,
and an arbitrary uid in group 0 can write the cache. Before this change the
same checks failed on the cache and succeeded on writing the code.
@semx
semx force-pushed the chore/slim-image branch from 7ae0b07 to 15b968e Compare October 1, 2026 22:29
@semx semx changed the title docker: fix the cache directory, keep the code read-only and install only the vCenter bindings docker: fix the cache directory, install security updates and only the vCenter bindings Oct 1, 2026
@semx

semx commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

I ran both images (the current development image and this branch) as Jobs on a kubeadm cluster with the security context from k8s-netbox-sync-cronjob.yaml (runAsUser: 1000, readOnlyRootFilesystem: true), against NetBox 4.4.5 and the two captured vCenters:

pod setup current image this PR
example manifest as is, no volume caching disabled caching disabled
emptyDir on /app/cache cache written, gone with the pod the same
PVC, volume root-owned 0755 caching disabled caching disabled
PVC writable for gid 1000, second Job cache read for 17 object types the same
arbitrary uid in group 0, no volume caching disabled cache works

So you're right: in Kubernetes this PR makes no difference except for the arbitrary uid case, and the cache only outlives the Job on a PVC the uid can write to (fsGroup: 1000 on most storage). Where it does matter is Docker: /app/cache doesn't exist in the current image, so a named volume mounted there is created root-owned and uid 1000 can't write to it. With this PR a new volume takes netbox-sync:0 from the image, and so does a volume left over from the current image, since it is still empty (checked).

On this small inventory the warm-cache run took 22 s against 23 s cold. With a cache netbox-sync still requests a brief id list of every object type and fetches full data only for objects changed since the last run, so the gain is on large NetBox instances.

dist-upgrade: added, and I also removed pip from the venv and from the base interpreter. Trivy on the current development image (built 2026-10-01) reports 40 fixable OS findings, HIGH among them in openssl and pcre2 from trixie-security, plus 6 in pip's vendored urllib3, msgpack and setuptools, which netbox-sync doesn't use. On this branch it reports 0 fixable. The updates only reach users when the image is rebuilt, so a scheduled rebuild of the latest release would keep it current; I can send that as a separate PR if you want it.

Also in the branch: Dockerfile and .dockerignore added to the paths filter of docker-image.yml (merging this PR alone would not have rebuilt the development image), and the README docker run example mounts a volume on /app/cache. Rebased on development, description updated with the new numbers.

Signed-off-by: Ricardo Bartels <ricardo.bartels@telekom.de>
@bb-Ricardo

Copy link
Copy Markdown
Owner

@semx

I worked the Dockerfile a bit. Removed the python3 -m pip uninstall -y pip as it has no effect on the container size because it is part of underlying layers. consolidated RUN layers into one to remove amount of layers and possibly the size.

If this is fine with you, then you can merge this MR and we can release the new version.

@semx

semx commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Sorry, I missed your comment and read the change as an accident. You're right about the size: removing pip doesn't shrink the image, it adds 1.8 MB because the files stay in the base layer. The reason I had it there is the scan: with pip in /usr/local, Trivy reports 6 findings in its vendored msgpack, setuptools and urllib3, and without it 0. If that isn't worth it to you, I'll drop a211401 and merge your version as is. Otherwise I merge with it. Either way, the rest looks good to me, thanks for the cleanup.

@bb-Ricardo

Copy link
Copy Markdown
Owner

Yes, That is also a good reason, didn't think about that. Then let's keep it how it is and remove pip during build.

@semx
semx merged commit b6dda13 into development Oct 2, 2026
1 check passed
@bb-Ricardo
bb-Ricardo deleted the chore/slim-image branch October 3, 2026 19:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants