docker: fix the cache directory, install security updates and only the vCenter bindings - #573
Conversation
|
Hi, thank you for this improvement. I'm just curious if a cache directory is actually used inside the container and not as PVC. If the container runs as Job then the content of the cache folder will be purged after the Job finished. And yes |
vcf-sdk is a meta package that pulls every VMware Cloud Foundation binding (NSX, SDDC Manager, Operations, Fleet LCM, Installer, vSAN data protection). netbox-sync only needs create_vsphere_client and the tagging client, which come from vmware-vcenter and the vapi runtime and common client it depends on. Install vmware-vcenter pinned to the version vcf-sdk resolves to today, so a rebuild of the same tag gets the same SDK, and drop the apt-get line that installed nothing. Built from the same development commit: 339 MB -> 248 MB uncompressed, 58 MB -> 52 MB gzipped. All 552 Python files under vmware/ and com/ in the new image are byte-identical to the current one; the 155 files that are no longer installed belong to SDDC Manager, the VCF installer and snapservice. pyvmomi, vmware-vapi-runtime, vmware-vapi-common-client and vmware-vcenter stay at 9.1.1.0, --help works and the process still runs as uid 1000.
The application files were copied with the service user as owner, so the running process could rewrite its own code, while /app itself stayed owned by root, so the default cache directory (/app/cache) could not be created and every container run logged "NetBox caching DISABLED" and fetched all objects from NetBox again. Copy the code as root, read-only for the service user, and create /app/cache owned by the service user and group 0 with mode 0770, which also covers platforms that run the image with an arbitrary uid in group 0. Checked with --read-only --cap-drop ALL --security-opt no-new-privileges and a volume on /app/cache: the cache is writable, the code is not, --help works, and an arbitrary uid in group 0 can write the cache. Before this change the same checks failed on the cache and succeeded on writing the code.
|
I ran both images (the current
So you're right: in Kubernetes this PR makes no difference except for the arbitrary uid case, and the cache only outlives the Job on a PVC the uid can write to ( On this small inventory the warm-cache run took 22 s against 23 s cold. With a cache netbox-sync still requests a brief id list of every object type and fetches full data only for objects changed since the last run, so the gain is on large NetBox instances.
Also in the branch: |
Signed-off-by: Ricardo Bartels <ricardo.bartels@telekom.de>
|
I worked the Dockerfile a bit. Removed the If this is fine with you, then you can merge this MR and we can release the new version. |
|
Sorry, I missed your comment and read the change as an accident. You're right about the size: removing pip doesn't shrink the image, it adds 1.8 MB because the files stay in the base layer. The reason I had it there is the scan: with pip in |
|
Yes, That is also a good reason, didn't think about that. Then let's keep it how it is and remove pip during build. |
Changes to the image, all checked by running it against NetBox, not only by building it.
Cache directory.
/appbelongs to root and/app/cachedoes not exist in the image, so the service user cannot create it: every run logsError writing to cache directory: /app/cache/NetBox caching DISABLED, even with a volume mounted on/app/cache(Docker creates that mount point as root). The code was also copied with the service user as owner, so the process could rewrite it. Now the code is copied as root and read-only for the service user, and/app/cacheis created owned by the service user and group 0 with mode 0770 (group 0 covers platforms that run images with an arbitrary uid). The README example mounts a volume there.Security updates. The final stage runs
apt-get update && apt-get dist-upgrade -y, and pip is removed from the venv and from the base interpreter, since nothing uses it at runtime. Trivy on the currentdevelopmentimage: 40 fixable OS findings (among them HIGH in openssl and pcre2 from trixie-security) and 6 in pip's vendored copies of urllib3, msgpack and setuptools. On this branch: 0 fixable.Smaller, pinned SDK.
vcf-sdkis a meta package that pulls every VMware Cloud Foundation binding (NSX, SDDC Manager, Operations, Fleet LCM, Installer, vSAN data protection). netbox-sync only usescreate_vsphere_clientand the tagging client fromvmware-vcenter, which is now installed pinned to the versionvcf-sdkresolves to today (9.1.1.0), so rebuilding a tag gives the same SDK. All 552 Python files undervmware/andcom/are byte-identical to the current image; the 155 that are gone belong to SDDC Manager, the VCF installer and snapservice.Image build trigger.
Dockerfileand.dockerignorewere not in thepathsfilter ofdocker-image.yml, so a change to them alone did not rebuild thedevelopmentimage.developmentimageFunctional check under a rootless Docker daemon (Docker 29.1.3 on Ubuntu 26.04): NetBox 4.4.5, two vcsim instances replaying the captured vCenters from
tests/fixtures/vcsim(vchvr 6.7 with 28 VMs, vc001 8.0.3 with 34 VMs) as two sources, netbox-sync run with--read-only --cap-drop ALL --security-opt no-new-privileges, settings mounted read-only, a volume on/app/cache, two passes each:The
docker runline from the README works as written. The arm64 build (QEMU, and native on Apple silicon) starts and imports the vSphere bindings. vCenter tag sync cannot be exercised against vcsim (it does not implement the vAPI endpoint the SDK talks to), which is why the SDK files are compared byte for byte instead.Kubernetes, as Jobs on a kubeadm cluster with the security context from
k8s-netbox-sync-cronjob.yaml: the image change makes no difference there except for an arbitrary uid in group 0 without a volume, see the comment below.