hubsingest accepts data from people outside the project and hands it to admins for review. This page lists the trust boundaries, what each credential grants, and how to report a problem.
Contributor (untrusted)
│ S3 over HTTPS, one key pair per endpoint
▼
Endpoint namespace <username>-ns: gateway, volume, Secret, Ingress
│ admin scales the gateway down and mounts the same volume
▼
ClamAV scan, then RStudio review ← admin works on contributor data
│ rclone, with the admin's storage credentials
▼
Hubs storage (trusted)
The tooling does not enforce the order scan, review, promote. An admin can launch RStudio on data that was never scanned.
Grants full S3 access to one endpoint: create and delete buckets, read, write and delete objects. It is the gateway's root key and reaches nothing outside its namespace.
- Generate it with
openssl rand -hex 32, one per contribution. - The access key is the username and is public; only the secret key is secret.
- Send it to the contributor over a channel other than email, and never in an issue or pull request.
- It exists in the repository secret, the
versitygw-credentialsSecret and the contributor's machine. Deleting the endpoint removes the Secret; delete the repository secret at the same time.
A leaked key lets anyone write to the endpoint, and an admin later scans and opens what was written. Rotate it as in operations.md or delete the endpoint.
The RStudio password of every session that admin launches, for any
contributor. RStudio runs as rstudio with the contribution mounted and rclone
installed. The launch workflow uses the secret of the admin who dispatches it,
and the run log records who launched which session. Anyone with write access
to the repository can read every ADMINPASS_* value, through a modified
workflow or from a running rstudio Deployment.
- Use a long, unique password from a password manager.
hubsingest_launch_rstudio.shwrites the password into therstudioDeployment as a plain environment value, readable by anyone who can read Deployments in the namespace.- Delete the secret when the admin leaves the team.
Grants whatever its service account can do. Everyone who can run workflows in the repository can use it, so repository write access is cluster access.
The tooling needs the following permissions. Bind this ClusterRole instead of
cluster-admin:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: hubsingest-ci
rules:
- apiGroups: [""]
resources: ["namespaces"]
verbs: ["get", "list", "create", "patch", "delete"]
- apiGroups: [""]
resources: ["secrets", "services", "persistentvolumeclaims"]
verbs: ["get", "list", "watch", "create", "patch", "delete"]
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["pods/exec"] # kubectl cp of the scan report
verbs: ["create"]
- apiGroups: ["apps"]
resources: ["deployments", "deployments/scale"]
verbs: ["get", "list", "watch", "create", "patch", "update", "delete"]
- apiGroups: ["batch"]
resources: ["jobs"]
verbs: ["get", "list", "watch", "create", "patch", "delete"]
- apiGroups: ["networking.k8s.io"]
resources: ["ingresses"]
verbs: ["get", "list", "watch", "create", "patch", "delete"]
- apiGroups: ["cert-manager.io"]
resources: ["certificates", "certificaterequests"]
verbs: ["get", "list", "watch", "delete"]
- apiGroups: ["acme.cert-manager.io"]
resources: ["orders"]
verbs: ["get", "list", "delete"]This role can still read every Secret and start workloads in every namespace, which amounts to control of the cluster. Kubernetes RBAC cannot limit a ClusterRole to namespaces that do not exist yet.
The API server must accept connections from GitHub-hosted runners, so it is reachable from the internet. A self-hosted runner inside the cluster network would allow closing it.
Public, behind TLS, authenticated by one key pair. There is no rate limit or address allow-list; the ingress body size limit is the only request limit. Options, in order of effort:
- An ingress-nginx rate limit annotation (
nginx.ingress.kubernetes.io/limit-rps). - An address allow-list annotation (
nginx.ingress.kubernetes.io/whitelist-source-range) when the contributor's network is known. - Deleting endpoints that are not in use; see operations.md.
The gateway runs with --debug, which writes every request's URL, headers
and query arguments, and request and response bodies other than object data,
to the pod log.
Anyone who can read pod logs in the namespace sees bucket and object names.
Public, behind TLS, protected by one password with no second factor, and running where admin credentials meet contributor data. It runs until the endpoint is deleted. Delete the endpoint when the review is finished.
Rancher (rancher.cloudman.hubsingest.bioconductor.org) and CloudMan
(cloudman.hubsingest.bioconductor.org) are served by the same ingress
controller as the endpoints, and both manage the cluster. Restrict them to
known addresses or put them behind a VPN, and keep Rancher on a release line
that receives security fixes; see updating.md.
ClamAV detects known malware. It does not detect R data built to run code.
Objects read from .rds, .rda and .RData files can contain functions,
environments and formulas that run when used, and R before 4.4.0 could run
code while reading such a file (CVE-2024-27322).
- Do not
source()contributor scripts. - Read contributor R objects only in a session that holds no production credentials; configure rclone afterwards, in a new RStudio launch if R objects were loaded.
- Give rclone a credential for the destination bucket only.
Deletes every namespace ending in -ns, without confirmation and without
undo, for anyone with repository write access. The workflow lists the
namespaces before deleting them.
- Several steps interpolate
inputs.usernamedirectly into shell commands, so a crafted username runs shell code. Anyone who can dispatch a workflow already holds the cluster credential through it. Pass inputs and secrets to new steps throughenv:. - Each workflow limits
GITHUB_TOKENtocontents: read; Build RStudio Image also haspackages: write. actions/checkoutis pinned to a major version tag, which its publisher can move. The Docker actions are pinned to commit SHAs, and the create workflow'samazon/aws-clicontainer to an image digest.- The workflows check the downloaded
kubectlagainst the SHA-256 checksum published with it.
- At rest: depends on the StorageClass. The example class sets
encrypted: "true"; check the driver used in production. - In transit: TLS from the client to the ingress controller, plain HTTP from the controller to the gateway inside the cluster.
- Retention: data stays until the namespace is deleted. With
reclaimPolicy: Deletethe cloud volume is deleted too; withRetainit remains until removed by hand. Deleting a volume is not a secure erase. - Backups: none. See operations.md.
Not applied in the templates or the scripts:
- A NetworkPolicy in each endpoint namespace that admits traffic only from the ingress controller. Inside the cluster, the gateway and RStudio serve plain HTTP to any pod.
- An egress policy for the
rstudiopod that blocks the Kubernetes API and the instance metadata address169.254.169.254. The pod holds contributor data and the admin's rclone credentials. Whether policies are enforced depends on the cluster's network plugin. automountServiceAccountToken: falseon the gateway, scan and RStudio pods. None of them calls the Kubernetes API.- A restrictive security context where the image allows it, such as the
gateway and the scan's
holdercontainer:allowPrivilegeEscalation: false,capabilities: {drop: [ALL]}andrunAsNonRoot: true. - Resource requests and limits; see operations.md.
Do not open a public issue. Email the Bioconductor core team at bioconductorcoreteam@gmail.com, the address listed at https://bioconductor.org/about/.