Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,7 @@ mux.Handle("/newregistry/", http.StripPrefix("/newregistry", newHandler.Routes()
- Keep functions short and focused
- Write tests for new functionality
- Document exported types and functions
- A new Prometheus metric needs a tile on `/ui/analytics` and an entry in `metricSurface` (`internal/server/analytics_coverage_test.go`). The page is meant to be a complete view of `/metrics`, so `TestEveryMetricIsSurfaced` fails until both exist.

## Testing

Expand Down
56 changes: 55 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -1111,6 +1111,7 @@ Response:
The proxy serves a web UI under `/ui`. No separate frontend build is needed -- templates and assets are embedded in the binary. `GET /` redirects to `/ui/`. The UI is mounted under its own prefix so a reverse proxy can apply different access rules to it than to the package endpoints (for example, requiring auth for `PathPrefix(/ui)` while leaving `/npm`, `/pypi` etc. open to build machines).

- **Dashboard** (`/ui/`) -- cache stats, popular packages, recently cached artifacts, and vulnerability overview.
- **Analytics** (`/ui/analytics`) -- accumulated download size as a ring broken down by ecosystem with the total in the middle, the cache size, artifact, package and version counts, a per-ecosystem table, the vulnerability overview, and a Runtime card mirroring every counter `/metrics` exposes. See [Analytics](#analytics).
- **Install guide** (`/ui/install`) -- per-ecosystem configuration instructions, so you don't have to look them up here.
- **Package browser** (`/ui/packages`) -- browse all cached packages with filtering by ecosystem and sorting by hits, size, name, or vulnerability count.
- **Search** (`/ui/search?q=...`) -- search cached packages by name.
Expand Down Expand Up @@ -1139,15 +1140,68 @@ The proxy exposes Prometheus metrics at `GET /metrics`. All metric names are pre
| `proxy_health_probe_failures_total` | counter | `step` | Storage health probe failures by failing step (`write`, `size`, `read`, `verify`, `delete`). |
| `proxy_circuit_breaker_state` | gauge | `registry` | Artifact-fetch circuit breaker state per upstream registry (0 closed, 2 open). Published once that registry's breaker has tripped. |
| `proxy_circuit_breaker_trips_total` | counter | `registry` | Circuit breaker trips per upstream registry. |
| `proxy_ecosystem_downloaded_bytes` | gauge | `ecosystem` | Accumulated bytes served from cache: cache hits multiplied by the artifact size they served. |
| `proxy_ecosystem_artifact_downloads` | gauge | `ecosystem` | Accumulated artifact downloads served from cache. |
| `proxy_ecosystem_cache_size_bytes` | gauge | `ecosystem` | Size of cached artifacts per ecosystem. |
| `proxy_ecosystem_cached_artifacts` | gauge | `ecosystem` | Number of cached artifacts per ecosystem. |
| `proxy_ecosystem_packages` | gauge | `ecosystem` | Known packages per ecosystem. |
| `proxy_ecosystem_versions` | gauge | `ecosystem` | Known package versions per ecosystem. |
| `proxy_response_bytes_total` | counter | `ecosystem` | Response body bytes written to clients. Route-labelled, see the label caveat below. |

Cache size and artifact count are refreshed every 60 seconds. Circuit breaker state is read from the fetcher on each scrape of `/metrics` and each `/health` request, so `proxy_circuit_breaker_trips_total` counts the trips visible between those reads — a breaker that opens and recovers entirely between two scrapes is not counted. The remaining metrics update on each request.
Cache size, artifact count and the per-ecosystem gauges are refreshed every 60 seconds, from a single pass over the database. Circuit breaker state is read from the fetcher on each scrape of `/metrics` and each `/health` request, so `proxy_circuit_breaker_trips_total` counts the trips visible between those reads — a breaker that opens and recovers entirely between two scrapes is not counted. The remaining metrics update on each request.

The breaker metrics carry one series per upstream host, but only for hosts whose breaker has tripped at least once since startup. A breaker is created per host the proxy fetches artifacts from, and for some ecosystems that host comes from upstream metadata rather than from configuration (composer takes it from a package's `dist.url`, helm from the chart URLs in `index.yaml`), so publishing every host would let upstream content grow the series count for the lifetime of the process. Once a host has tripped it keeps reporting, so a recovery still shows up as a transition to 0 rather than as a series that vanishes. `/health` is not a persistent time series and lists every breaker, tripped or not.

The `registry` label is the host of the URL the artifact was fetched from. Because that URL can come from upstream metadata, it is not always one a host can be read off — a signed `dist.url` that fails to parse, for instance — and such a breaker is labelled `hostless-url-<digest>` instead, where the digest is keyed by a value drawn fresh at startup. Neither `/metrics` nor `/health` requires authentication, so a fetch URL is never published as a label or a key; the digest identifies the breaker for as long as the process runs without revealing the URL behind it or letting a chosen URL be matched against it.

Alert on `proxy_circuit_breaker_state == 2` sustained for more than a few minutes: while a breaker is open, artifact downloads for that upstream fail with HTTP 502 on every cache miss, and only a single probe request per backoff interval reaches the upstream. Cached artifacts keep serving, and so does metadata for the same ecosystem (metadata does not go through the circuit breaker), so installs fail in a way that looks like a partial upstream outage.

#### Accumulated download size

`proxy_ecosystem_downloaded_bytes` is, for every cached artifact, the number of times it was served multiplied by its size. It answers "how much traffic has this proxy actually carried", which is the number that matters when sizing egress or justifying the cache. Two properties are worth knowing before alerting on it.

**It counts cache hits, not upstream fetches.** The request that first pulls an artifact through the proxy is a miss and is not counted; only later hits are. So the accumulated total is also the upstream bandwidth the cache has saved, not the total bytes the proxy has ever sent.

**Eviction removes history.** Evicting an artifact clears its size, so its past hits drop out of the total. That is why these are gauges rather than counters, and why the figure can step downwards. Chart them with `max_over_time` rather than `increase`, and read a drop after an eviction sweep as expected rather than as data loss.

Nothing at the schema level ties `artifacts.version_purl` to a version row, so a cached artifact can end up with no ecosystem to attribute it to. Those are reported under the ecosystem `unattributed` rather than dropped, which keeps the per-ecosystem figures adding up to `proxy_cache_size_bytes` and `proxy_cached_artifacts_total`. A non-zero `unattributed` means the database holds artifact rows whose version or package rows have gone missing.

The `ecosystem` label on these six is taken from the package record and normalized, so aliases collapse: a database carrying both `gem` and `rubygems` rows -- the proxy writes the former, git-pkgs the latter -- reports one `rubygems` series with the two summed.

### Analytics

`/ui/analytics` reports the accumulated download size as a ring broken down by ecosystem, the cache figures from the dashboard, a per-ecosystem table, the vulnerability overview, and a **Runtime** card covering every remaining metric `/metrics` exposes.

The ring shows at most six slices, because part-to-whole stops being readable past that. When more ecosystems are active the smallest are folded into a single "Other" slice; the table below lists every one of them, so nothing is hidden, only summarised.

#### No history is kept

The proxy stores no time series. The page reads the database and the in-process metric registry at request time and reports current state; there is nowhere for it to read yesterday's figures from, and nothing is written for tomorrow. That splits the figures in two, and the page says which is which.

**Database-derived figures survive a restart.** Download volume, cache size and the package, version and artifact counts come from the `artifacts`, `packages` and `versions` tables, so they are as durable as the database.

**Registry-derived figures do not.** Everything in the Runtime card -- request counts and latencies, cache hit rate, upstream and storage errors, circuit breaker state, scan results -- lives only in this process's Prometheus registry and starts from zero on restart. A small number there next to a large one above just means the proxy started recently.

For history, trends and alerting, scrape `/metrics` with Prometheus. That is the intended split: the UI answers "what is true now", Prometheus answers "what happened".

The same figures are available as JSON from `GET /stats`, which reports `downloaded_bytes`, `downloads` and an `ecosystems` array carrying the per-ecosystem breakdown, served from the same 60-second snapshot the page and the gauges read. When the aggregation fails with no snapshot to fall back on, the response carries `stats_unavailable: true` rather than passing zeros off as a count -- the endpoint keeps answering with the artifact count and cache size either way.

#### Three ecosystem label sets

`ecosystem` means three slightly different things across `/metrics`, and queries that join across them need to know which.

**From the package record, normalized.** The six `proxy_ecosystem_*` gauges, `proxy_cache_hits_total`, `proxy_cache_misses_total`, `proxy_integrity_failures_total` and the scan metrics. Aliases collapse here: `gem` reads as `rubygems`, `composer` as `packagist`, `go` as `golang`.

**From the request path.** `proxy_requests_total`, `proxy_request_duration_seconds` and `proxy_response_bytes_total`. The names mostly coincide with the normalized ones -- these also report `rubygems`, `packagist` and `golang` -- but the Debian route reports `debian` where the package record says `deb`, and any path that is not a package endpoint reports `other`, which corresponds to no ecosystem at all.

**From the handler's own name.** `proxy_upstream_fetch_duration_seconds` and `proxy_upstream_errors_total`, which report `composer`, `gem` and `go` where the other two sets report `packagist`, `rubygems` and `golang`. These are published series and are deliberately left as they are; renaming them would break existing queries and alerts.

### Grafana dashboard

A ready-made dashboard lives at [`deploy/grafana/git-pkgs-proxy.json`](deploy/grafana/git-pkgs-proxy.json). Import it via **Dashboards -> New -> Import** and pick your Prometheus data source when prompted; it has no hardcoded data source UID.

It carries three ecosystem filters rather than one, because `ecosystem` means three different things across `/metrics` -- see the label sets above. **Ecosystem** filters the database-derived gauges, **Route** the request-path counters, and **Upstream** the two upstream fetch metrics. All three are query variables, so they populate from whatever labels your proxy is actually reporting; a panel is on the one its metric belongs to, and the panel descriptions say which.

### Health Check

`/health` returns a structured JSON report of subsystem health. HTTP 200 if all checks pass; 503 if any fail.
Expand Down
Loading
Loading