Skip to content

Add an opt-in open URL cache at /url/ - #399

Open
benoit-nexthop wants to merge 1 commit into
git-pkgs:mainfrom
nexthop-ai:benoit.url-proxy
Open

benoit-nexthop wants to merge 1 commit into
git-pkgs:mainfrom
nexthop-ai:benoit.url-proxy

Conversation

@benoit-nexthop

Copy link
Copy Markdown

Why

Some build systems don't fetch from a package registry. They download pinned source archives straight from wherever each project publishes them: GitHub release assets and /archive/ tarballs, GNU and kernel.org mirrors, project download pages, and so on. One build can pull from a few dozen hosts, and the list changes whenever a dependency moves. Usually the build already knows each archive's sha256.

We'd like to cache those downloads with the proxy we already run for PyPI, npm and crates.io. /generic/ isn't a good fit for this:

  • Every host must be a named upstream. That's a config change each time a dependency moves to a new host, and the list is never quite complete.
  • Large archives are buffered in memory. Paths other than GitHub release assets go through the metadata cache, which holds the whole body in memory up to metadata_max_size (100MB by default) and revalidates it after metadata_ttl. Archives that are both large and immutable are better served by the artifact cache.

This PR adds an opt-in route for that case. We realise it is an open proxy for whoever can reach it, unlike /generic/, whose docs promise it isn't one. So it is off by default, and the trade-offs are documented below and in docs/configuration.md. We run it in a staging and a production deployment on an internal network.

What it adds

GET|HEAD /url/{host}/{path}?{query}, switched on with url_proxy.enabled.

  • It fetches https://{host}/{path}?{query}. Only https on port 443 is fetched; ports, userinfo and IPv6 literals are rejected.
  • The file is treated as immutable. It streams into the artifact cache through GetOrFetchArtifactFromURL[WithDigest], so it gets the existing coalescing, per-fetch storage paths, eviction and the denylist, with no size limit.
  • Later requests are served from the cache without revalidation, including while the host is down.

/url/sha256/{hex}/{host}/{path} also checks the download against that digest.

  • A mismatching download isn't cached and returns 502.
  • A cached copy with another digest is fetched again.
  • This is the recommended form: without a digest, the first copy fetched is kept until it's evicted.

Cache identity:

Field Value
ecosystem url
name the host
version a 128-bit hash of path and query
filename the sanitized last path segment

Storage paths therefore look like url/{host}/{hash}/{fetch id}/{file}.

Redirects (url_proxy.direct_serve), for this route only. GET requests get a 302 to the stored object, both on a hit and right after a miss is stored.

  • The target is a presigned URL, as with storage.direct_serve.
  • With the new storage.direct_serve_public_url, the target is a plain URL under a bucket location that serves objects anonymously, e.g. http://rgw:7480/bucket/prefix, so it never expires. Global storage.direct_serve uses this too when it's set, through a small directServeURL helper shared with checkCache.
  • HEAD is answered from the cache record and is never redirected, because a presigned GET URL rejects HEAD.

Objects that disappear. Before redirecting, the handler checks the object still exists. A record whose object was removed outside the proxy, for example by a bucket lifecycle rule, is refetched once, and a second failure returns 502. Each case is counted in the new proxy_cache_missing_objects_total{ecosystem}. The streaming path already recovered like this through the shared fetch's recheck.

Config: url_proxy.enabled, url_proxy.direct_serve, url_proxy.fetch_timeout (default 9m, a whole-body bound), and storage.direct_serve_public_url, each with a PROXY_* environment variable. url_proxy.enabled is rejected together with storage.cache_artifacts: false, like direct_serve and mirror_api. The route is labelled url in request metrics and in the dashboard's ecosystem list.

Security considerations

The route will fetch and store any public https file for any client that can reach it. Here's what limits that, and what's left to operators.

Built in:

  • Off by default, like mirror_api.
  • Its own upstream client (newURLProxy in server.go):
    • It keeps the safehttp dial gate, so loopback, RFC 1918, CGNAT and link-local targets (including 169.254.169.254) are refused on every redirect hop. upstream.allow_private_hosts and upstream.allow_loopback still apply.
    • It is not wrapped in the auth transport, so upstream.auth credentials and the OCI bearer-challenge exchange never apply to arbitrary hosts.
    • It clears Transport.Proxy, so an HTTPS_PROXY in the environment can't be used to reach targets the dial gate would refuse.
    • It has its own whole-request timeout.
  • No inbound credentials or other client headers are forwarded upstream.
  • Paths: traversal, backslashes, ports and userinfo are rejected, and cached filenames are limited to [A-Za-z0-9._+-].

Left to operators, as documented:

  • Expose the route only to trusted networks, since there's no inbound auth (as for the rest of the proxy).
  • Set storage.max_size to bound how much it can store.
  • direct_serve_public_url requires the bucket, or at least its url/ prefix, to allow anonymous reads. That suits caches of public downloads only.
  • Without a digest, a URL whose content changes keeps serving its first copy.

Not addressed: bandwidth and storage abuse by trusted-network clients beyond max_size eviction, rate limiting, and a host allowlist or denylist. A per-route host allowlist would be easy to add if you'd prefer the route to support one.

Testing

  • go test -race ./... and golangci-lint run ./... (v2.13.1, as in CI) pass.
  • New tests in internal/handler/urlproxy_test.go:
    • path parsing, including the digest form and the rejected inputs above;
    • caching and serving while the host is down;
    • HEAD from the cache record;
    • a digest mismatch and a changed digest;
    • redirects to a public URL, a presigned URL, and the streaming fallback;
    • missing-object refetch, with redirects on and off, plus the failure case;
    • a 120MB download against file storage.
  • internal/server/urlproxy_test.go checks that the client refuses loopback targets and sends no configured credentials. There are also config and PublicObjectURL tests.
  • Running on our side behind an internal ingress, with Ceph RGW storage and direct_serve_public_url. Cold and cached downloads of a 124MB GitHub release asset, digest rejection, refusal of 169.254.169.254, and recovery from a deleted object all behaved as described.

Open questions

  • Are you happy to take an open route like this upstream, given the opt-in and the documentation? If not, we'll keep it in our fork. Feedback on the safeguards is welcome either way.
  • Naming: /url/, url_proxy and the url ecosystem are placeholders.
  • storage.direct_serve_public_url is useful without /url/. I'm happy to split it into its own PR.

🤖 Generated with Claude Code

GET /url/{host}/{path}?{query} fetches https://{host}/{path}?{query} from
any host and caches it as an immutable artifact; /url/sha256/{hex}/... also
verifies the download against that digest, refetching a cached copy with
another one. Files stream into the artifact cache, so unlike /generic/'s
metadata path there is no size cap.

With url_proxy.direct_serve, GETs are redirected to the stored object, on
hits and right after a miss is stored: to storage.direct_serve_public_url
when the bucket serves objects anonymously, else to a presigned URL. HEAD
is answered from the cache record. Before redirecting, the object's
existence is checked; one removed behind the proxy's back (a bucket
lifecycle rule) is refetched once and counted in
proxy_cache_missing_objects_total.

The route fetches through its own client: the safehttp dial gate on every
redirect hop, no upstream.auth credentials or OCI token exchange, no
environment HTTP proxy, and url_proxy.fetch_timeout (default 9m). It is
off by default since it is an open proxy for its clients.
@benoit-nexthop

Copy link
Copy Markdown
Author

BTW we're finding this proxy super useful for our builds, I know this is a fairly recent project started earlier this year, so thanks for open sourcing it!

@andrew

andrew commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thanks for the thorough writeup; I want to cover this case, and the client isolation in newURLProxy is the right approach.

Two pieces would be better as their own PRs:

  • storage.direct_serve_public_url, with a note in the storage docs that it replaces presigning for every ecosystem when global direct_serve is on.
  • The missing-object refetch and metric, moved into the shared cache-hit path since every ecosystem can hit it under direct_serve.

For the route itself, I'd like it folded into /generic/ as a new entry type, roughly:

upstream:
  generic:
    github: "https://github.com"       # existing shorthand
    sources:
      hosts: ["*"]                      # or ["*.gnu.org", "github.com"]
      immutable: true
      direct_serve: true
      fetch_timeout: "9m"

/generic/sources/{host}/{path}, with an optional sha256/{hex}/ segment after the entry name. A few constraints:

  • hosts entries always use the isolated client, with upstream.allow_private_hosts / allow_loopback forced off regardless of global config.
  • The allowlist checks the request host only; redirect targets are left to safehttp. Worth a line in the docs.
  • When hosts is ["*"], the sha256/ form is required.

On that last point: how does the digest reach the proxy from your builds? Bazel's --downloader_config rewrite only has the URL to work with, so a plain rewrite can't supply it. If it comes via a wrapper or a header that might change where it goes.

If folding into /generic/ turns out messier than a separate route, say so; the constraints above apply either way. The two split-out PRs are useful on their own if the full rework is more than you want to take on.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants