From 4ebd805ffcc8d0afd2b5b757cd34910810ce2c12 Mon Sep 17 00:00:00 2001 From: Wicliff Wolda Date: Thu, 1 Oct 2026 14:32:20 +0200 Subject: [PATCH 1/6] Per-ecosystem cache and download statistics - Add GetEcosystemStats, aggregating packages, versions, artifacts, cache size, downloads and downloaded bytes per ecosystem - Publish six proxy_ecosystem_* gauges from it on the existing one-minute cache-stats tick - Set and selectively delete rather than Reset, so no scrape lands on a half-populated vector - Report artifacts with no package row under "unattributed" rather than dropping them, so the figures still add up Part 2 of 6 splitting #381 up. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 20 +- internal/database/analytics_test.go | 375 ++++++++++++++++++++++++++++ internal/database/queries.go | 201 +++++++++++++++ internal/metrics/analytics_test.go | 99 ++++++++ internal/metrics/metrics.go | 131 ++++++++++ internal/server/server.go | 25 ++ 6 files changed, 850 insertions(+), 1 deletion(-) create mode 100644 internal/database/analytics_test.go create mode 100644 internal/metrics/analytics_test.go diff --git a/README.md b/README.md index eb265244..24285ca5 100644 --- a/README.md +++ b/README.md @@ -1139,8 +1139,14 @@ The proxy exposes Prometheus metrics at `GET /metrics`. All metric names are pre | `proxy_health_probe_failures_total` | counter | `step` | Storage health probe failures by failing step (`write`, `size`, `read`, `verify`, `delete`). | | `proxy_circuit_breaker_state` | gauge | `registry` | Artifact-fetch circuit breaker state per upstream registry (0 closed, 2 open). Published once that registry's breaker has tripped. | | `proxy_circuit_breaker_trips_total` | counter | `registry` | Circuit breaker trips per upstream registry. | +| `proxy_ecosystem_downloaded_bytes` | gauge | `ecosystem` | Accumulated bytes served from cache: cache hits multiplied by the artifact size they served. | +| `proxy_ecosystem_artifact_downloads` | gauge | `ecosystem` | Accumulated artifact downloads served from cache. | +| `proxy_ecosystem_cache_size_bytes` | gauge | `ecosystem` | Size of cached artifacts per ecosystem. | +| `proxy_ecosystem_cached_artifacts` | gauge | `ecosystem` | Number of cached artifacts per ecosystem. | +| `proxy_ecosystem_packages` | gauge | `ecosystem` | Known packages per ecosystem. | +| `proxy_ecosystem_versions` | gauge | `ecosystem` | Known package versions per ecosystem. | -Cache size and artifact count are refreshed every 60 seconds. Circuit breaker state is read from the fetcher on each scrape of `/metrics` and each `/health` request, so `proxy_circuit_breaker_trips_total` counts the trips visible between those reads — a breaker that opens and recovers entirely between two scrapes is not counted. The remaining metrics update on each request. +Cache size, artifact count and the per-ecosystem gauges are refreshed every 60 seconds, from a single pass over the database. Circuit breaker state is read from the fetcher on each scrape of `/metrics` and each `/health` request, so `proxy_circuit_breaker_trips_total` counts the trips visible between those reads — a breaker that opens and recovers entirely between two scrapes is not counted. The remaining metrics update on each request. The breaker metrics carry one series per upstream host, but only for hosts whose breaker has tripped at least once since startup. A breaker is created per host the proxy fetches artifacts from, and for some ecosystems that host comes from upstream metadata rather than from configuration (composer takes it from a package's `dist.url`, helm from the chart URLs in `index.yaml`), so publishing every host would let upstream content grow the series count for the lifetime of the process. Once a host has tripped it keeps reporting, so a recovery still shows up as a transition to 0 rather than as a series that vanishes. `/health` is not a persistent time series and lists every breaker, tripped or not. @@ -1148,6 +1154,18 @@ The `registry` label is the host of the URL the artifact was fetched from. Becau Alert on `proxy_circuit_breaker_state == 2` sustained for more than a few minutes: while a breaker is open, artifact downloads for that upstream fail with HTTP 502 on every cache miss, and only a single probe request per backoff interval reaches the upstream. Cached artifacts keep serving, and so does metadata for the same ecosystem (metadata does not go through the circuit breaker), so installs fail in a way that looks like a partial upstream outage. +#### Accumulated download size + +`proxy_ecosystem_downloaded_bytes` is, for every cached artifact, the number of times it was served multiplied by its size. It answers "how much traffic has this proxy actually carried", which is the number that matters when sizing egress or justifying the cache. Two properties are worth knowing before alerting on it. + +**It counts cache hits, not upstream fetches.** The request that first pulls an artifact through the proxy is a miss and is not counted; only later hits are. So the accumulated total is also the upstream bandwidth the cache has saved, not the total bytes the proxy has ever sent. + +**Eviction removes history.** Evicting an artifact clears its size, so its past hits drop out of the total. That is why these are gauges rather than counters, and why the figure can step downwards. Chart them with `max_over_time` rather than `increase`, and read a drop after an eviction sweep as expected rather than as data loss. + +Nothing at the schema level ties `artifacts.version_purl` to a version row, so a cached artifact can end up with no ecosystem to attribute it to. Those are reported under the ecosystem `unattributed` rather than dropped, which keeps the per-ecosystem figures adding up to `proxy_cache_size_bytes` and `proxy_cached_artifacts_total`. A non-zero `unattributed` means the database holds artifact rows whose version or package rows have gone missing. + +The `ecosystem` label on these six is taken from the package record and normalized, so aliases collapse: a database carrying both `gem` and `rubygems` rows -- the proxy writes the former, git-pkgs the latter -- reports one `rubygems` series with the two summed. + ### Health Check `/health` returns a structured JSON report of subsystem health. HTTP 200 if all checks pass; 503 if any fail. diff --git a/internal/database/analytics_test.go b/internal/database/analytics_test.go new file mode 100644 index 00000000..269b9d2f --- /dev/null +++ b/internal/database/analytics_test.go @@ -0,0 +1,375 @@ +package database + +import ( + "database/sql" + "path/filepath" + "testing" + "time" + + "github.com/git-pkgs/purl" +) + +// seedArtifact inserts a package, a version and a cached artifact so that the +// ecosystem aggregation has a complete join path to walk. +func seedArtifact(t *testing.T, db *DB, ecosystem, name, version string, size, hits int64) { + t.Helper() + + pkgPURL := "pkg:" + ecosystem + "/" + name + versionPURL := pkgPURL + "@" + version + + if err := db.UpsertPackage(&Package{PURL: pkgPURL, Ecosystem: ecosystem, Name: name}); err != nil { + t.Fatalf("UpsertPackage(%s): %v", pkgPURL, err) + } + if err := db.UpsertVersion(&Version{PURL: versionPURL, PackagePURL: pkgPURL}); err != nil { + t.Fatalf("UpsertVersion(%s): %v", versionPURL, err) + } + if err := db.UpsertArtifact(&Artifact{ + VersionPURL: versionPURL, + Filename: name + "-" + version + ".tgz", + UpstreamURL: "https://example.test/" + name, + StoragePath: sql.NullString{String: "objects/" + name, Valid: true}, + Size: sql.NullInt64{Int64: size, Valid: true}, + FetchedAt: sql.NullTime{Time: time.Now(), Valid: true}, + HitCount: hits, + }); err != nil { + t.Fatalf("UpsertArtifact(%s): %v", versionPURL, err) + } +} + +func newTestDB(t *testing.T) *DB { + t.Helper() + db, err := Create(filepath.Join(t.TempDir(), "analytics.db")) + if err != nil { + t.Fatalf("Create: %v", err) + } + t.Cleanup(func() { _ = db.Close() }) + return db +} + +func TestGetEcosystemStats(t *testing.T) { + db := newTestDB(t) + + // npm: 2 artifacts across 2 packages, 1000*3 + 500*1 = 3500 bytes served. + seedArtifact(t, db, "npm", "lodash", "4.17.21", 1000, 3) + seedArtifact(t, db, "npm", "express", "4.18.2", 500, 1) + // cargo: a single heavily-hit artifact, 200*10 = 2000 bytes served. + seedArtifact(t, db, "cargo", "serde", "1.0.0", 200, 10) + + stats, err := db.GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + if len(stats) != 2 { + t.Fatalf("expected 2 ecosystems, got %d: %+v", len(stats), stats) + } + + // Ordered by accumulated download volume, so npm (3500) precedes cargo (2000). + npm, cargo := stats[0], stats[1] + if npm.Ecosystem != "npm" || cargo.Ecosystem != "cargo" { + t.Fatalf("expected npm then cargo, got %q then %q", npm.Ecosystem, cargo.Ecosystem) + } + + if npm.DownloadedBytes != 3500 { + t.Errorf("npm DownloadedBytes = %d, want 3500", npm.DownloadedBytes) + } + if npm.Downloads != 4 { + t.Errorf("npm Downloads = %d, want 4", npm.Downloads) + } + if npm.CacheSize != 1500 { + t.Errorf("npm CacheSize = %d, want 1500", npm.CacheSize) + } + if npm.Artifacts != 2 { + t.Errorf("npm Artifacts = %d, want 2", npm.Artifacts) + } + if npm.Packages != 2 { + t.Errorf("npm Packages = %d, want 2", npm.Packages) + } + if npm.Versions != 2 { + t.Errorf("npm Versions = %d, want 2", npm.Versions) + } + + if cargo.DownloadedBytes != 2000 { + t.Errorf("cargo DownloadedBytes = %d, want 2000", cargo.DownloadedBytes) + } + if cargo.Downloads != 10 { + t.Errorf("cargo Downloads = %d, want 10", cargo.Downloads) + } +} + +func TestGetEcosystemStatsEmptyDatabase(t *testing.T) { + stats, err := newTestDB(t).GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + if len(stats) != 0 { + t.Errorf("expected no ecosystems, got %+v", stats) + } +} + +// An ecosystem known from metadata but with nothing cached should still appear, +// so the analytics table can show it as idle rather than omitting it. +func TestGetEcosystemStatsIncludesEcosystemsWithoutArtifacts(t *testing.T) { + db := newTestDB(t) + seedArtifact(t, db, "npm", "lodash", "4.17.21", 1000, 2) + + if err := db.UpsertPackage(&Package{PURL: "pkg:gem/rails", Ecosystem: "gem", Name: "rails"}); err != nil { + t.Fatalf("UpsertPackage: %v", err) + } + + stats, err := db.GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + if len(stats) != 2 { + t.Fatalf("expected 2 ecosystems, got %d: %+v", len(stats), stats) + } + + // gem has no download volume, so it sorts last. + gem := stats[1] + if gem.Ecosystem != "gem" { + t.Fatalf("expected gem last, got %q", gem.Ecosystem) + } + if gem.Packages != 1 { + t.Errorf("gem Packages = %d, want 1", gem.Packages) + } + if gem.Artifacts != 0 || gem.CacheSize != 0 || gem.DownloadedBytes != 0 { + t.Errorf("expected gem to report no cached artifacts, got %+v", gem) + } +} + +// Eviction clears storage_path and size, which drops the artifact's historical +// hits out of the accumulated total. Pinning that here so the documented +// behaviour of GetEcosystemStats does not drift. +func TestGetEcosystemStatsExcludesEvictedArtifacts(t *testing.T) { + db := newTestDB(t) + seedArtifact(t, db, "npm", "lodash", "4.17.21", 1000, 3) + seedArtifact(t, db, "npm", "express", "4.18.2", 500, 2) + + cleared, err := db.ClearArtifactCache("pkg:npm/lodash@4.17.21", "lodash-4.17.21.tgz", "objects/lodash") + if err != nil { + t.Fatalf("ClearArtifactCache: %v", err) + } + if !cleared { + t.Fatal("ClearArtifactCache cleared nothing") + } + + stats, err := db.GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + if len(stats) != 1 { + t.Fatalf("expected 1 ecosystem, got %d: %+v", len(stats), stats) + } + + // Only express remains cached: 500 * 2 = 1000. + if got := stats[0].DownloadedBytes; got != 1000 { + t.Errorf("DownloadedBytes = %d, want 1000 (evicted artifact must not count)", got) + } + if got := stats[0].Artifacts; got != 1 { + t.Errorf("Artifacts = %d, want 1", got) + } + // The package row survives eviction, so both packages still count. + if got := stats[0].Packages; got != 2 { + t.Errorf("Packages = %d, want 2", got) + } +} + +func TestSortEcosystemStatsTieBreaks(t *testing.T) { + stats := []EcosystemStats{ + {Ecosystem: "zzz"}, + {Ecosystem: "aaa"}, + {Ecosystem: "mid", CacheSize: 10}, + {Ecosystem: "top", DownloadedBytes: 5}, + } + sortEcosystemStats(stats) + + want := []string{"top", "mid", "aaa", "zzz"} + for i, name := range want { + if stats[i].Ecosystem != name { + t.Errorf("position %d = %q, want %q (full order: %+v)", i, stats[i].Ecosystem, name, stats) + } + } +} + +// A cached artifact whose version row is missing must still be counted, so the +// per-ecosystem totals reconcile with GetTotalCacheSize rather than quietly +// coming up short. Nothing at the schema level enforces the link. +func TestGetEcosystemStatsCountsUnattributedArtifacts(t *testing.T) { + db := newTestDB(t) + seedArtifact(t, db, "npm", "lodash", "4.17.21", 1000, 2) + + // An artifact pointing at a version that was never recorded. + if err := db.UpsertArtifact(&Artifact{ + VersionPURL: "pkg:npm/orphan@9.9.9", + Filename: "orphan-9.9.9.tgz", + UpstreamURL: "https://example.test/orphan", + StoragePath: sql.NullString{String: "objects/orphan", Valid: true}, + Size: sql.NullInt64{Int64: 500, Valid: true}, + FetchedAt: sql.NullTime{Time: time.Now(), Valid: true}, + HitCount: 4, + }); err != nil { + t.Fatalf("UpsertArtifact: %v", err) + } + + stats, err := db.GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + + var cacheSize, artifacts, downloaded int64 + var sawUnattributed bool + for _, e := range stats { + cacheSize += e.CacheSize + artifacts += e.Artifacts + downloaded += e.DownloadedBytes + if e.Ecosystem == unattributedEcosystem { + sawUnattributed = true + } + } + + if !sawUnattributed { + t.Errorf("the orphaned artifact was dropped; got %+v", stats) + } + + // These must match what the unjoined queries report, or the UI shows two + // totals that do not add up. + wantSize, err := db.GetTotalCacheSize() + if err != nil { + t.Fatalf("GetTotalCacheSize: %v", err) + } + if cacheSize != wantSize { + t.Errorf("per-ecosystem cache size sums to %d, but GetTotalCacheSize reports %d", cacheSize, wantSize) + } + + wantCount, err := db.GetCachedArtifactCount() + if err != nil { + t.Fatalf("GetCachedArtifactCount: %v", err) + } + if artifacts != wantCount { + t.Errorf("per-ecosystem artifacts sum to %d, but GetCachedArtifactCount reports %d", artifacts, wantCount) + } + + // 1000*2 from lodash plus 500*4 from the orphan. + if downloaded != 4000 { + t.Errorf("downloaded bytes = %d, want 4000", downloaded) + } +} + +// The proxy writes "gem" and git-pkgs writes "rubygems". A database that has +// seen both must report one ecosystem, not two rows splitting its share. +func TestGetEcosystemStatsMergesAliasedEcosystems(t *testing.T) { + db := newTestDB(t) + seedArtifact(t, db, "gem", "colorize", "1.1.0", 1000, 3) + seedArtifact(t, db, "rubygems", "rails", "7.1.0", 4000, 2) + seedArtifact(t, db, "npm", "lodash", "4.17.21", 500, 1) + + stats, err := db.GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + if len(stats) != 2 { + t.Fatalf("expected gem and rubygems merged into one row alongside npm, got %+v", stats) + } + + var ruby *EcosystemStats + for i := range stats { + if purl.NormalizeEcosystem(stats[i].Ecosystem) == "rubygems" { + ruby = &stats[i] + } + } + if ruby == nil { + t.Fatal("no row for the rubygems ecosystem") + } + + // 1000*3 + 4000*2 = 11000 + if ruby.DownloadedBytes != 11000 { + t.Errorf("DownloadedBytes = %d, want 11000", ruby.DownloadedBytes) + } + if ruby.CacheSize != 5000 { + t.Errorf("CacheSize = %d, want 5000", ruby.CacheSize) + } + if ruby.Artifacts != 2 || ruby.Packages != 2 { + t.Errorf("Artifacts=%d Packages=%d, want 2 and 2", ruby.Artifacts, ruby.Packages) + } + + // The surviving row keeps the spelling with the most cached bytes, because + // the UI filters and links by it and the packages table stores it raw. + if ruby.Ecosystem != "rubygems" { + t.Errorf("Ecosystem = %q, want the dominant raw spelling %q", ruby.Ecosystem, "rubygems") + } +} + +// A version whose package row is missing must be bucketed like an orphaned +// artifact, so the per-ecosystem figures reconcile with GetCacheStats. +func TestGetEcosystemStatsCountsUnattributedVersions(t *testing.T) { + db := newTestDB(t) + seedArtifact(t, db, "npm", "lodash", "4.17.21", 1000, 1) + if err := db.UpsertVersion(&Version{PURL: "pkg:npm/ghost@1.0.0", PackagePURL: "pkg:npm/ghost"}); err != nil { + t.Fatalf("UpsertVersion: %v", err) + } + + stats, err := db.GetEcosystemStats() + if err != nil { + t.Fatalf("GetEcosystemStats: %v", err) + } + + var versions int64 + for _, e := range stats { + versions += e.Versions + } + + cacheStats, err := db.GetCacheStats() + if err != nil { + t.Fatalf("GetCacheStats: %v", err) + } + if versions != cacheStats.TotalVersions { + t.Errorf("per-ecosystem versions sum to %d, but GetCacheStats reports %d", + versions, cacheStats.TotalVersions) + } +} + +// The surviving spelling must be the one with the most cached bytes whatever +// order the rows arrive in. GetEcosystemStats builds its input by ranging a +// map, so an implementation that compared against the running total instead of +// each row's own size would pick a different name between refreshes — flipping +// the analytics row's badge and its /ui/packages filter link. +func TestMergeAliasedEcosystemsPicksLargestWhateverTheOrder(t *testing.T) { + rows := []EcosystemStats{ + {Ecosystem: "gem", CacheSize: 5, Artifacts: 1}, + {Ecosystem: "rubygems", CacheSize: 4, Artifacts: 1}, + {Ecosystem: "RubyGems", CacheSize: 6, Artifacts: 1}, + } + + for _, order := range [][]int{ + {0, 1, 2}, {0, 2, 1}, {1, 0, 2}, {1, 2, 0}, {2, 0, 1}, {2, 1, 0}, + } { + in := make([]EcosystemStats, 0, len(order)) + for _, i := range order { + in = append(in, rows[i]) + } + + got := mergeAliasedEcosystems(in) + if len(got) != 1 { + t.Fatalf("order %v: got %d rows, want 1", order, len(got)) + } + if got[0].Ecosystem != "RubyGems" { + t.Errorf("order %v: Ecosystem = %q, want the largest spelling %q", + order, got[0].Ecosystem, "RubyGems") + } + if got[0].CacheSize != 15 { + t.Errorf("order %v: CacheSize = %d, want 15", order, got[0].CacheSize) + } + } +} + +// Equal sizes still have to resolve to one answer, for the same reason. +func TestMergeAliasedEcosystemsBreaksSizeTiesByName(t *testing.T) { + a := []EcosystemStats{{Ecosystem: "gem", CacheSize: 5}, {Ecosystem: "rubygems", CacheSize: 5}} + b := []EcosystemStats{{Ecosystem: "rubygems", CacheSize: 5}, {Ecosystem: "gem", CacheSize: 5}} + + first := mergeAliasedEcosystems(a)[0].Ecosystem + second := mergeAliasedEcosystems(b)[0].Ecosystem + if first != second { + t.Errorf("input order changed the surviving spelling: %q vs %q", first, second) + } +} diff --git a/internal/database/queries.go b/internal/database/queries.go index 12d076e2..ae6351b3 100644 --- a/internal/database/queries.go +++ b/internal/database/queries.go @@ -4,9 +4,11 @@ import ( "database/sql" "errors" "fmt" + "sort" "time" "github.com/git-pkgs/artifacts" + "github.com/git-pkgs/purl" "github.com/opencontainers/go-digest" ) @@ -1128,3 +1130,202 @@ func (db *DB) UpsertMetadataCache(entry *MetadataCacheEntry) error { } return nil } + +// Analytics queries + +// EcosystemStats aggregates cache and download activity for one ecosystem. +// +// DownloadedBytes is the accumulated download volume: every cache hit on an +// artifact served its full size, so the sum of hit_count * size is the number +// of bytes the proxy has handed to clients from cache for this ecosystem. +type EcosystemStats struct { + Ecosystem string `db:"ecosystem"` + Packages int64 `db:"packages"` + Versions int64 `db:"versions"` + Artifacts int64 `db:"artifacts"` + CacheSize int64 `db:"cache_size"` + Downloads int64 `db:"downloads"` + DownloadedBytes int64 `db:"downloaded_bytes"` +} + +// GetEcosystemStats returns per-ecosystem cache and download totals, ordered by +// accumulated download volume descending. Ecosystems with rows in packages but +// nothing cached are included with zeroed artifact counters. +// +// Artifacts evicted from the cache no longer contribute: eviction clears the +// size column, so their historical hits drop out of the accumulated total. +func (db *DB) GetEcosystemStats() ([]EcosystemStats, error) { + byEcosystem := make(map[string]*EcosystemStats) + + get := func(ecosystem string) *EcosystemStats { + if s, ok := byEcosystem[ecosystem]; ok { + return s + } + s := &EcosystemStats{Ecosystem: ecosystem} + byEcosystem[ecosystem] = s + return s + } + + if err := db.eachCount(`SELECT ecosystem, COUNT(*) FROM packages GROUP BY ecosystem`, + func(ecosystem string, n int64) { get(ecosystem).Packages = n }); err != nil { + return nil, err + } + + // Left joined and bucketed for the same reason as artifacts below: a version + // whose package row is missing would otherwise vanish here while still + // counting in GetCacheStats' COUNT(*), leaving the two unable to reconcile. + if err := db.eachCount(` + SELECT COALESCE(p.ecosystem, '`+unattributedEcosystem+`'), COUNT(*) + FROM versions v + LEFT JOIN packages p ON p.purl = v.package_purl + GROUP BY COALESCE(p.ecosystem, '`+unattributedEcosystem+`') + `, func(ecosystem string, n int64) { get(ecosystem).Versions = n }); err != nil { + return nil, err + } + + // The artifacts table is proxy-specific: a database inherited from + // git-pkgs carries packages and versions without it. + hasArtifacts, err := db.HasTable("artifacts") + if err != nil { + return nil, err + } + if hasArtifacts { + if err := db.eachArtifactStat(get); err != nil { + return nil, err + } + } + + stats := make([]EcosystemStats, 0, len(byEcosystem)) + for _, s := range byEcosystem { + stats = append(stats, *s) + } + stats = mergeAliasedEcosystems(stats) + sortEcosystemStats(stats) + return stats, nil +} + +// mergeAliasedEcosystems combines rows whose ecosystem names normalize to the +// same canonical name. The proxy writes "gem" and git-pkgs writes "rubygems", +// so a database that has seen both carries two rows for one ecosystem; left +// split they would render as two table rows and two chart slices, each with +// half the real share. +// +// The surviving row keeps the raw spelling of whichever input held the most +// cached bytes, because that string is what the UI filters and links by: the +// packages table stores the raw value, so substituting the canonical name would +// produce links that match nothing. That comparison is against each input's own +// size, not the running total, which would otherwise let the first spelling win +// simply by being merged into first. +func mergeAliasedEcosystems(stats []EcosystemStats) []EcosystemStats { + merged := make(map[string]*EcosystemStats, len(stats)) + largest := make(map[string]int64, len(stats)) + order := make([]string, 0, len(stats)) + + for i := range stats { + key := purl.NormalizeEcosystem(stats[i].Ecosystem) + into, ok := merged[key] + if !ok { + row := stats[i] + merged[key] = &row + largest[key] = stats[i].CacheSize + order = append(order, key) + continue + } + + // The name tiebreak matters: GetEcosystemStats builds its input by + // ranging a map, so without it two spellings of equal size would swap + // between refreshes and flip the row's badge and filter link. + if stats[i].CacheSize > largest[key] || + (stats[i].CacheSize == largest[key] && stats[i].Ecosystem < into.Ecosystem) { + largest[key] = stats[i].CacheSize + into.Ecosystem = stats[i].Ecosystem + } + into.Packages += stats[i].Packages + into.Versions += stats[i].Versions + into.Artifacts += stats[i].Artifacts + into.CacheSize += stats[i].CacheSize + into.Downloads += stats[i].Downloads + into.DownloadedBytes += stats[i].DownloadedBytes + } + + out := make([]EcosystemStats, 0, len(order)) + for _, key := range order { + out = append(out, *merged[key]) + } + return out +} + +// unattributedEcosystem collects cached artifacts whose version or package row +// is missing. Nothing enforces that link at the schema level, so an inner join +// would silently drop such rows and leave the per-ecosystem totals short of +// GetTotalCacheSize — two numbers that sit side by side in the UI and in +// Grafana. Bucketing them keeps the two reconcilable. +const unattributedEcosystem = "unattributed" + +func (db *DB) eachArtifactStat(get func(string) *EcosystemStats) error { + rows, err := db.Query(` + SELECT COALESCE(p.ecosystem, '` + unattributedEcosystem + `'), + COUNT(*), + COALESCE(SUM(a.size), 0), + COALESCE(SUM(a.hit_count), 0), + COALESCE(SUM(a.hit_count * a.size), 0) + FROM artifacts a + LEFT JOIN versions v ON v.purl = a.version_purl + LEFT JOIN packages p ON p.purl = v.package_purl + WHERE a.storage_path IS NOT NULL + GROUP BY COALESCE(p.ecosystem, '` + unattributedEcosystem + `') + `) + if err != nil { + return err + } + defer func() { _ = rows.Close() }() + + for rows.Next() { + var ecosystem string + var artifacts, cacheSize, downloads, downloadedBytes int64 + if err := rows.Scan(&ecosystem, &artifacts, &cacheSize, &downloads, &downloadedBytes); err != nil { + return err + } + s := get(ecosystem) + s.Artifacts = artifacts + s.CacheSize = cacheSize + s.Downloads = downloads + s.DownloadedBytes = downloadedBytes + } + return rows.Err() +} + +// eachCount runs a two-column "group by" query and hands each (key, count) pair to fn. +func (db *DB) eachCount(query string, fn func(key string, n int64)) error { + rows, err := db.Query(query) + if err != nil { + return err + } + defer func() { _ = rows.Close() }() + + for rows.Next() { + var key string + var n int64 + if err := rows.Scan(&key, &n); err != nil { + return err + } + fn(key, n) + } + return rows.Err() +} + +// sortEcosystemStats orders by accumulated download volume, then by cache size, +// then by name, so that ecosystems with no traffic yet still sort predictably. +func sortEcosystemStats(stats []EcosystemStats) { + sort.Slice(stats, func(i, j int) bool { + a, b := stats[i], stats[j] + switch { + case a.DownloadedBytes != b.DownloadedBytes: + return a.DownloadedBytes > b.DownloadedBytes + case a.CacheSize != b.CacheSize: + return a.CacheSize > b.CacheSize + default: + return a.Ecosystem < b.Ecosystem + } + }) +} diff --git a/internal/metrics/analytics_test.go b/internal/metrics/analytics_test.go new file mode 100644 index 00000000..6d391bc6 --- /dev/null +++ b/internal/metrics/analytics_test.go @@ -0,0 +1,99 @@ +package metrics + +import ( + "testing" + + "github.com/prometheus/client_golang/prometheus" + "github.com/prometheus/client_golang/prometheus/testutil" +) + +func TestUpdateEcosystemStats(t *testing.T) { + UpdateEcosystemStats([]EcosystemStats{ + {Ecosystem: "npm", Packages: 10, Versions: 25, Artifacts: 40, CacheSize: 1500, Downloads: 4, DownloadedBytes: 3500}, + {Ecosystem: "cargo", Packages: 2, Versions: 3, Artifacts: 3, CacheSize: 200, Downloads: 10, DownloadedBytes: 2000}, + }) + + if got := testutil.ToFloat64(EcosystemDownloadedBytes.WithLabelValues("npm")); got != 3500 { + t.Errorf("npm downloaded bytes = %v, want 3500", got) + } + if got := testutil.ToFloat64(EcosystemDownloads.WithLabelValues("npm")); got != 4 { + t.Errorf("npm downloads = %v, want 4", got) + } + if got := testutil.ToFloat64(EcosystemCacheSize.WithLabelValues("cargo")); got != 200 { + t.Errorf("cargo cache size = %v, want 200", got) + } + if got := testutil.ToFloat64(EcosystemCachedArtifacts.WithLabelValues("npm")); got != 40 { + t.Errorf("npm cached artifacts = %v, want 40", got) + } + if got := testutil.ToFloat64(EcosystemPackages.WithLabelValues("npm")); got != 10 { + t.Errorf("npm packages = %v, want 10", got) + } + if got := testutil.ToFloat64(EcosystemVersions.WithLabelValues("npm")); got != 25 { + t.Errorf("npm versions = %v, want 25", got) + } +} + +// A refresh that no longer mentions an ecosystem must drop its series rather +// than leave it frozen at the last observed value. +func TestUpdateEcosystemStatsDropsStaleSeries(t *testing.T) { + UpdateEcosystemStats([]EcosystemStats{ + {Ecosystem: "npm", DownloadedBytes: 3500}, + {Ecosystem: "gem", DownloadedBytes: 900}, + }) + if got := testutil.CollectAndCount(EcosystemDownloadedBytes); got != 2 { + t.Fatalf("expected 2 series after first refresh, got %d", got) + } + + UpdateEcosystemStats([]EcosystemStats{{Ecosystem: "npm", DownloadedBytes: 4000}}) + + if got := testutil.CollectAndCount(EcosystemDownloadedBytes); got != 1 { + t.Errorf("expected 1 series after gem disappeared, got %d", got) + } + if got := testutil.ToFloat64(EcosystemDownloadedBytes.WithLabelValues("npm")); got != 4000 { + t.Errorf("npm downloaded bytes = %v, want 4000", got) + } +} + +// Ecosystem labels are normalized so these gauges join against the other +// metrics that take their ecosystem from a package record. +func TestUpdateEcosystemStatsNormalizesLabels(t *testing.T) { + UpdateEcosystemStats([]EcosystemStats{{Ecosystem: "NPM", DownloadedBytes: 12}}) + + if got := testutil.ToFloat64(EcosystemDownloadedBytes.WithLabelValues("npm")); got != 12 { + t.Errorf("normalized npm gauge = %v, want 12", got) + } +} + +// GetEcosystemStats groups by the raw packages.ecosystem column, so a database +// carrying both spellings of an aliased ecosystem yields two rows that +// normalize to one label. They must sum rather than overwrite each other. +func TestUpdateEcosystemStatsSumsAliasedRows(t *testing.T) { + reg := prometheus.NewRegistry() + reg.MustRegister(EcosystemDownloadedBytes, EcosystemCacheSize, EcosystemPackages) + + UpdateEcosystemStats([]EcosystemStats{ + {Ecosystem: "gem", DownloadedBytes: 100, CacheSize: 10, Packages: 1}, + {Ecosystem: "rubygems", DownloadedBytes: 200, CacheSize: 20, Packages: 2}, + {Ecosystem: "npm", DownloadedBytes: 50, CacheSize: 5, Packages: 3}, + }) + + if got := testutil.ToFloat64(EcosystemDownloadedBytes.WithLabelValues("rubygems")); got != 300 { + t.Errorf("rubygems downloaded bytes = %v, want 300", got) + } + if got := testutil.ToFloat64(EcosystemCacheSize.WithLabelValues("rubygems")); got != 30 { + t.Errorf("rubygems cache size = %v, want 30", got) + } + if got := testutil.ToFloat64(EcosystemPackages.WithLabelValues("rubygems")); got != 3 { + t.Errorf("rubygems packages = %v, want 3", got) + } + // An unaliased ecosystem is unaffected. + if got := testutil.ToFloat64(EcosystemDownloadedBytes.WithLabelValues("npm")); got != 50 { + t.Errorf("npm downloaded bytes = %v, want 50", got) + } + + // A later refresh replaces the value rather than adding to it. + UpdateEcosystemStats([]EcosystemStats{{Ecosystem: "npm", DownloadedBytes: 50}}) + if got := testutil.ToFloat64(EcosystemDownloadedBytes.WithLabelValues("npm")); got != 50 { + t.Errorf("npm downloaded bytes after a second refresh = %v, want 50", got) + } +} diff --git a/internal/metrics/metrics.go b/internal/metrics/metrics.go index aff57682..2f20fb5a 100644 --- a/internal/metrics/metrics.go +++ b/internal/metrics/metrics.go @@ -4,6 +4,7 @@ package metrics import ( "net/http" "strconv" + "sync" "time" "github.com/git-pkgs/purl" @@ -138,6 +139,57 @@ var ( []string{"step"}, ) + // Per-ecosystem gauges, derived from the database rather than incremented + // in the request path, and refreshed on the same tick as the cache gauges + // above. + EcosystemDownloadedBytes = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "proxy_ecosystem_downloaded_bytes", + Help: "Accumulated bytes served from cache per ecosystem (cache hits x artifact size)", + }, + []string{"ecosystem"}, + ) + + EcosystemDownloads = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "proxy_ecosystem_artifact_downloads", + Help: "Accumulated artifact downloads served from cache per ecosystem", + }, + []string{"ecosystem"}, + ) + + EcosystemCacheSize = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "proxy_ecosystem_cache_size_bytes", + Help: "Size of cached artifacts per ecosystem in bytes", + }, + []string{"ecosystem"}, + ) + + EcosystemCachedArtifacts = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "proxy_ecosystem_cached_artifacts", + Help: "Number of cached artifacts per ecosystem", + }, + []string{"ecosystem"}, + ) + + EcosystemPackages = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "proxy_ecosystem_packages", + Help: "Number of known packages per ecosystem", + }, + []string{"ecosystem"}, + ) + + EcosystemVersions = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "proxy_ecosystem_versions", + Help: "Number of known package versions per ecosystem", + }, + []string{"ecosystem"}, + ) + // Scanning metrics ScanDuration = prometheus.NewHistogramVec( prometheus.HistogramOpts{ @@ -183,6 +235,12 @@ func init() { ActiveRequests, IntegrityFailures, HealthProbeFailures, + EcosystemDownloadedBytes, + EcosystemDownloads, + EcosystemCacheSize, + EcosystemCachedArtifacts, + EcosystemPackages, + EcosystemVersions, ScanDuration, ScanBlocked, ScanErrors, @@ -263,6 +321,79 @@ func UpdateCacheStats(sizeBytes, artifactCount int64) { CachedArtifacts.Set(float64(artifactCount)) } +// EcosystemStats is one ecosystem's row of the snapshot published as gauges. +type EcosystemStats struct { + Ecosystem string + Packages int64 + Versions int64 + Artifacts int64 + CacheSize int64 + Downloads int64 + DownloadedBytes int64 +} + +// publishedEcosystems tracks which labels the per-ecosystem gauges currently +// carry, so a label that disappears can be deleted individually. +var ( + publishedMu sync.Mutex + publishedEcosystems = map[string]bool{} +) + +// UpdateEcosystemStats republishes the per-ecosystem gauges from a fresh snapshot. +// +// Rows are summed by label before anything is published, because normalizing +// collapses aliases: a database holding both "gem" and "rubygems" rows -- the +// proxy writes the former, git-pkgs the latter -- arrives as two rows belonging +// to one label, and publishing them one at a time would leave only the last. +// +// The vectors are not Reset() first. Reset followed by a repopulating loop +// leaves a window in which a scrape sees the families empty or half filled, +// which renders as a spurious gap on any panel built from them. Each series is +// Set instead, and only labels that have actually disappeared are deleted. +func UpdateEcosystemStats(stats []EcosystemStats) { + totals := make(map[string]EcosystemStats, len(stats)) + for _, s := range stats { + ecosystem := purl.NormalizeEcosystem(s.Ecosystem) + t := totals[ecosystem] + t.Packages += s.Packages + t.Versions += s.Versions + t.Artifacts += s.Artifacts + t.CacheSize += s.CacheSize + t.Downloads += s.Downloads + t.DownloadedBytes += s.DownloadedBytes + totals[ecosystem] = t + } + + for ecosystem, t := range totals { + EcosystemDownloadedBytes.WithLabelValues(ecosystem).Set(float64(t.DownloadedBytes)) + EcosystemDownloads.WithLabelValues(ecosystem).Set(float64(t.Downloads)) + EcosystemCacheSize.WithLabelValues(ecosystem).Set(float64(t.CacheSize)) + EcosystemCachedArtifacts.WithLabelValues(ecosystem).Set(float64(t.Artifacts)) + EcosystemPackages.WithLabelValues(ecosystem).Set(float64(t.Packages)) + EcosystemVersions.WithLabelValues(ecosystem).Set(float64(t.Versions)) + } + + publishedMu.Lock() + defer publishedMu.Unlock() + + for ecosystem := range publishedEcosystems { + if _, still := totals[ecosystem]; still { + continue + } + EcosystemDownloadedBytes.DeleteLabelValues(ecosystem) + EcosystemDownloads.DeleteLabelValues(ecosystem) + EcosystemCacheSize.DeleteLabelValues(ecosystem) + EcosystemCachedArtifacts.DeleteLabelValues(ecosystem) + EcosystemPackages.DeleteLabelValues(ecosystem) + EcosystemVersions.DeleteLabelValues(ecosystem) + } + + publishedEcosystems = make(map[string]bool, len(totals)) + for ecosystem := range totals { + publishedEcosystems[ecosystem] = true + } +} + // UpdateCircuitBreakerState updates circuit breaker state gauge. // state: 0=closed, 1=half-open, 2=open func UpdateCircuitBreakerState(registry string, state int) { diff --git a/internal/server/server.go b/internal/server/server.go index 35066486..e7fc5ddf 100644 --- a/internal/server/server.go +++ b/internal/server/server.go @@ -514,6 +514,31 @@ func (s *Server) updateCacheStats() { return } metrics.UpdateCacheStats(stats.TotalSize, stats.TotalArtifacts) + + ecosystems, err := s.db.GetEcosystemStats() + if err != nil { + s.logger.Warn("failed to get ecosystem stats for metrics", "error", err) + return + } + metrics.UpdateEcosystemStats(ecosystemMetrics(ecosystems)) +} + +// ecosystemMetrics converts database rows into the metrics package's own +// snapshot type, so that package keeps no dependency on the database schema. +func ecosystemMetrics(stats []database.EcosystemStats) []metrics.EcosystemStats { + out := make([]metrics.EcosystemStats, 0, len(stats)) + for _, e := range stats { + out = append(out, metrics.EcosystemStats{ + Ecosystem: e.Ecosystem, + Packages: e.Packages, + Versions: e.Versions, + Artifacts: e.Artifacts, + CacheSize: e.CacheSize, + Downloads: e.Downloads, + DownloadedBytes: e.DownloadedBytes, + }) + } + return out } // Shutdown gracefully shuts down the server. From 05015344cb5476b56ac2f1dd08265d063cb7f01c Mon Sep 17 00:00:00 2001 From: Wicliff Wolda Date: Thu, 1 Oct 2026 14:32:20 +0200 Subject: [PATCH 2/6] Add the /ui/analytics page - Ring of download size by ecosystem, cache figures, per-ecosystem table, vulnerability overview and a Runtime card - Add metrics.Gather so the page can render counters that were never in the database - Add proxy_response_bytes_total and response-writer byte counting - Serve a retained snapshot behind a staleness banner when the aggregation fails, rather than rendering old figures as current - Extract the security overview into a shared component Part 3 of 6 splitting #381 up. Co-Authored-By: Claude Opus 5 (1M context) --- CONTRIBUTING.md | 1 + README.md | 28 ++ internal/metrics/metrics.go | 22 ++ internal/metrics/snapshot.go | 179 +++++++++++ internal/metrics/snapshot_test.go | 140 ++++++++ internal/server/analytics.go | 221 +++++++++++++ internal/server/analytics_coverage_test.go | 107 +++++++ internal/server/analytics_donut.go | 208 ++++++++++++ internal/server/analytics_donut_test.go | 211 ++++++++++++ internal/server/analytics_runtime.go | 235 ++++++++++++++ internal/server/analytics_runtime_test.go | 184 +++++++++++ internal/server/analytics_test.go | 277 ++++++++++++++++ internal/server/ecosystem_cache.go | 94 ++++++ internal/server/ecosystem_cache_test.go | 199 ++++++++++++ internal/server/middleware.go | 7 +- internal/server/server.go | 26 +- internal/server/server_test.go | 1 + internal/server/templates.go | 23 ++ .../templates/components/runtime_metrics.html | 162 ++++++++++ .../components/security_overview.html | 30 ++ internal/server/templates/layout/header.html | 1 + .../server/templates/pages/analytics.html | 302 ++++++++++++++++++ .../server/templates/pages/dashboard.html | 30 +- internal/server/templates_test.go | 40 +++ 24 files changed, 2686 insertions(+), 42 deletions(-) create mode 100644 internal/metrics/snapshot.go create mode 100644 internal/metrics/snapshot_test.go create mode 100644 internal/server/analytics.go create mode 100644 internal/server/analytics_coverage_test.go create mode 100644 internal/server/analytics_donut.go create mode 100644 internal/server/analytics_donut_test.go create mode 100644 internal/server/analytics_runtime.go create mode 100644 internal/server/analytics_runtime_test.go create mode 100644 internal/server/analytics_test.go create mode 100644 internal/server/ecosystem_cache.go create mode 100644 internal/server/ecosystem_cache_test.go create mode 100644 internal/server/templates/components/runtime_metrics.html create mode 100644 internal/server/templates/components/security_overview.html create mode 100644 internal/server/templates/pages/analytics.html diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 68a6acf9..2c697283 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -160,6 +160,7 @@ mux.Handle("/newregistry/", http.StripPrefix("/newregistry", newHandler.Routes() - Keep functions short and focused - Write tests for new functionality - Document exported types and functions +- A new Prometheus metric needs a tile on `/ui/analytics` and an entry in `metricSurface` (`internal/server/analytics_coverage_test.go`). The page is meant to be a complete view of `/metrics`, so `TestEveryMetricIsSurfaced` fails until both exist. ## Testing diff --git a/README.md b/README.md index 24285ca5..3390bd96 100644 --- a/README.md +++ b/README.md @@ -1111,6 +1111,7 @@ Response: The proxy serves a web UI under `/ui`. No separate frontend build is needed -- templates and assets are embedded in the binary. `GET /` redirects to `/ui/`. The UI is mounted under its own prefix so a reverse proxy can apply different access rules to it than to the package endpoints (for example, requiring auth for `PathPrefix(/ui)` while leaving `/npm`, `/pypi` etc. open to build machines). - **Dashboard** (`/ui/`) -- cache stats, popular packages, recently cached artifacts, and vulnerability overview. +- **Analytics** (`/ui/analytics`) -- accumulated download size as a ring broken down by ecosystem with the total in the middle, the cache size, artifact, package and version counts, a per-ecosystem table, the vulnerability overview, and a Runtime card mirroring every counter `/metrics` exposes. See [Analytics](#analytics). - **Install guide** (`/ui/install`) -- per-ecosystem configuration instructions, so you don't have to look them up here. - **Package browser** (`/ui/packages`) -- browse all cached packages with filtering by ecosystem and sorting by hits, size, name, or vulnerability count. - **Search** (`/ui/search?q=...`) -- search cached packages by name. @@ -1145,6 +1146,7 @@ The proxy exposes Prometheus metrics at `GET /metrics`. All metric names are pre | `proxy_ecosystem_cached_artifacts` | gauge | `ecosystem` | Number of cached artifacts per ecosystem. | | `proxy_ecosystem_packages` | gauge | `ecosystem` | Known packages per ecosystem. | | `proxy_ecosystem_versions` | gauge | `ecosystem` | Known package versions per ecosystem. | +| `proxy_response_bytes_total` | counter | `ecosystem` | Response body bytes written to clients. Route-labelled, see the label caveat below. | Cache size, artifact count and the per-ecosystem gauges are refreshed every 60 seconds, from a single pass over the database. Circuit breaker state is read from the fetcher on each scrape of `/metrics` and each `/health` request, so `proxy_circuit_breaker_trips_total` counts the trips visible between those reads — a breaker that opens and recovers entirely between two scrapes is not counted. The remaining metrics update on each request. @@ -1166,6 +1168,32 @@ Nothing at the schema level ties `artifacts.version_purl` to a version row, so a The `ecosystem` label on these six is taken from the package record and normalized, so aliases collapse: a database carrying both `gem` and `rubygems` rows -- the proxy writes the former, git-pkgs the latter -- reports one `rubygems` series with the two summed. +### Analytics + +`/ui/analytics` reports the accumulated download size as a ring broken down by ecosystem, the cache figures from the dashboard, a per-ecosystem table, the vulnerability overview, and a **Runtime** card covering every remaining metric `/metrics` exposes. + +The ring shows at most six slices, because part-to-whole stops being readable past that. When more ecosystems are active the smallest are folded into a single "Other" slice; the table below lists every one of them, so nothing is hidden, only summarised. + +#### No history is kept + +The proxy stores no time series. The page reads the database and the in-process metric registry at request time and reports current state; there is nowhere for it to read yesterday's figures from, and nothing is written for tomorrow. That splits the figures in two, and the page says which is which. + +**Database-derived figures survive a restart.** Download volume, cache size and the package, version and artifact counts come from the `artifacts`, `packages` and `versions` tables, so they are as durable as the database. + +**Registry-derived figures do not.** Everything in the Runtime card -- request counts and latencies, cache hit rate, upstream and storage errors, circuit breaker state, scan results -- lives only in this process's Prometheus registry and starts from zero on restart. A small number there next to a large one above just means the proxy started recently. + +For history, trends and alerting, scrape `/metrics` with Prometheus. That is the intended split: the UI answers "what is true now", Prometheus answers "what happened". + +#### Three ecosystem label sets + +`ecosystem` means three slightly different things across `/metrics`, and queries that join across them need to know which. + +**From the package record, normalized.** The six `proxy_ecosystem_*` gauges, `proxy_cache_hits_total`, `proxy_cache_misses_total`, `proxy_integrity_failures_total` and the scan metrics. Aliases collapse here: `gem` reads as `rubygems`, `composer` as `packagist`, `go` as `golang`. + +**From the request path.** `proxy_requests_total`, `proxy_request_duration_seconds` and `proxy_response_bytes_total`. The names mostly coincide with the normalized ones -- these also report `rubygems`, `packagist` and `golang` -- but the Debian route reports `debian` where the package record says `deb`, and any path that is not a package endpoint reports `other`, which corresponds to no ecosystem at all. + +**From the handler's own name.** `proxy_upstream_fetch_duration_seconds` and `proxy_upstream_errors_total`, which report `composer`, `gem` and `go` where the other two sets report `packagist`, `rubygems` and `golang`. These are published series and are deliberately left as they are; renaming them would break existing queries and alerts. + ### Health Check `/health` returns a structured JSON report of subsystem health. HTTP 200 if all checks pass; 503 if any fail. diff --git a/internal/metrics/metrics.go b/internal/metrics/metrics.go index 2f20fb5a..4021099d 100644 --- a/internal/metrics/metrics.go +++ b/internal/metrics/metrics.go @@ -190,6 +190,14 @@ var ( []string{"ecosystem"}, ) + ResponseBytes = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "proxy_response_bytes_total", + Help: "Total response body bytes written to clients, by ecosystem", + }, + []string{"ecosystem"}, + ) + // Scanning metrics ScanDuration = prometheus.NewHistogramVec( prometheus.HistogramOpts{ @@ -241,6 +249,7 @@ func init() { EcosystemCachedArtifacts, EcosystemPackages, EcosystemVersions, + ResponseBytes, ScanDuration, ScanBlocked, ScanErrors, @@ -259,6 +268,19 @@ func RecordRequest(ecosystem string, status int, duration time.Duration) { RequestDuration.WithLabelValues(ecosystem, statusStr).Observe(duration.Seconds()) } +// RecordResponse tracks what a client actually downloaded. +// +// Distinct from proxy_ecosystem_downloaded_bytes: that gauge is derived from +// the database and counts cache hits multiplied by artifact size, while this +// counts body bytes as they are written, including metadata responses and +// cache misses. +func RecordResponse(ecosystem string, bytes int64) { + if bytes <= 0 { + return + } + ResponseBytes.WithLabelValues(ecosystem).Add(float64(bytes)) +} + // RecordCacheHit increments cache hit counter. func RecordCacheHit(ecosystem string) { CacheHits.WithLabelValues(purl.NormalizeEcosystem(ecosystem)).Inc() diff --git a/internal/metrics/snapshot.go b/internal/metrics/snapshot.go new file mode 100644 index 00000000..6ac66fba --- /dev/null +++ b/internal/metrics/snapshot.go @@ -0,0 +1,179 @@ +package metrics + +import ( + "fmt" + "sort" + + "github.com/prometheus/client_golang/prometheus" + dto "github.com/prometheus/client_model/go" +) + +// Sample is one time series read out of the registry at a point in time. +// +// Counters and gauges carry their value in Value. Histograms carry their +// observation count and total in Count and Sum, and leave Value at zero — +// there is no single "value" for a histogram, and the UI shows a mean derived +// from Sum/Count rather than pretending otherwise. +type Sample struct { + Labels map[string]string + Value float64 + Count uint64 + Sum float64 +} + +// Label returns the value of one label, or "" if absent. +func (s Sample) Label(name string) string { + return s.Labels[name] +} + +// Mean returns the average observation of a histogram sample, or 0 when +// nothing has been observed yet. +func (s Sample) Mean() float64 { + if s.Count == 0 { + return 0 + } + return s.Sum / float64(s.Count) +} + +// Snapshot is the value of every registered proxy metric at one instant. +// +// These are process-lifetime figures held in the Prometheus registry, not +// database state: they start at zero when the proxy starts and are lost on +// restart. Anything that needs to survive a restart, or needs history, has to +// come from a Prometheus server scraping /metrics. +type Snapshot struct { + families map[string][]Sample +} + +// Gather reads the default Prometheus registry. It is the same data /metrics +// serves, shaped for rendering rather than for scraping. +func Gather() (*Snapshot, error) { + return GatherFrom(prometheus.DefaultGatherer) +} + +// GatherFrom reads an explicit gatherer, so tests can supply their own registry. +func GatherFrom(g prometheus.Gatherer) (*Snapshot, error) { + families, err := g.Gather() + if err != nil { + return nil, fmt.Errorf("gathering metrics: %w", err) + } + + snap := &Snapshot{families: make(map[string][]Sample, len(families))} + for _, mf := range families { + samples := make([]Sample, 0, len(mf.GetMetric())) + for _, m := range mf.GetMetric() { + samples = append(samples, sampleOf(m)) + } + snap.families[mf.GetName()] = samples + } + return snap, nil +} + +func sampleOf(m *dto.Metric) Sample { + s := Sample{Labels: make(map[string]string, len(m.GetLabel()))} + for _, lp := range m.GetLabel() { + s.Labels[lp.GetName()] = lp.GetValue() + } + + switch { + case m.GetCounter() != nil: + s.Value = m.GetCounter().GetValue() + case m.GetGauge() != nil: + s.Value = m.GetGauge().GetValue() + case m.GetHistogram() != nil: + h := m.GetHistogram() + s.Count = h.GetSampleCount() + s.Sum = h.GetSampleSum() + case m.GetSummary() != nil: + sm := m.GetSummary() + s.Count = sm.GetSampleCount() + s.Sum = sm.GetSampleSum() + case m.GetUntyped() != nil: + s.Value = m.GetUntyped().GetValue() + } + return s +} + +// Names returns the name of every metric family in the snapshot, sorted. +func (s *Snapshot) Names() []string { + if s == nil { + return nil + } + + names := make([]string, 0, len(s.families)) + for name := range s.families { + names = append(names, name) + } + sort.Strings(names) + return names +} + +// Samples returns every series of a metric, or nil when it has never been +// observed. Series are ordered by their label values so that rendering is +// stable between scrapes. +func (s *Snapshot) Samples(name string) []Sample { + if s == nil { + return nil + } + out := append([]Sample(nil), s.families[name]...) + sort.Slice(out, func(i, j int) bool { + return labelKey(out[i].Labels) < labelKey(out[j].Labels) + }) + return out +} + +func labelKey(labels map[string]string) string { + names := make([]string, 0, len(labels)) + for n := range labels { + names = append(names, n) + } + sort.Strings(names) + + key := "" + for _, n := range names { + key += n + "=" + labels[n] + "," + } + return key +} + +// Sum totals every series of a counter or gauge. +func (s *Snapshot) Sum(name string) float64 { + var total float64 + for _, sample := range s.Samples(name) { + total += sample.Value + } + return total +} + +// Count adds up the observation counts of a histogram across every series. +func (s *Snapshot) Count(name string) uint64 { + var total uint64 + for _, sample := range s.Samples(name) { + total += sample.Count + } + return total +} + +// Mean returns the average observation of a histogram across every series, or +// 0 when nothing has been observed. +func (s *Snapshot) Mean(name string) float64 { + var count uint64 + var sum float64 + for _, sample := range s.Samples(name) { + count += sample.Count + sum += sample.Sum + } + if count == 0 { + return 0 + } + return sum / float64(count) +} + +// SumBy groups a metric by one label and totals each group. +func (s *Snapshot) SumBy(name, label string) map[string]float64 { + out := make(map[string]float64) + for _, sample := range s.Samples(name) { + out[sample.Label(label)] += sample.Value + } + return out +} diff --git a/internal/metrics/snapshot_test.go b/internal/metrics/snapshot_test.go new file mode 100644 index 00000000..dd695f72 --- /dev/null +++ b/internal/metrics/snapshot_test.go @@ -0,0 +1,140 @@ +package metrics + +import ( + "testing" + + "github.com/prometheus/client_golang/prometheus" +) + +// newTestRegistry builds an isolated registry so these tests are not affected +// by counters other tests in this package have already incremented. +func newTestRegistry(t *testing.T) (*prometheus.Registry, *prometheus.CounterVec, *prometheus.HistogramVec, *prometheus.GaugeVec) { + t.Helper() + + reg := prometheus.NewRegistry() + counter := prometheus.NewCounterVec( + prometheus.CounterOpts{Name: "test_errors_total", Help: "t"}, + []string{"ecosystem", "kind"}, + ) + hist := prometheus.NewHistogramVec( + prometheus.HistogramOpts{Name: "test_duration_seconds", Help: "t"}, + []string{"op"}, + ) + gauge := prometheus.NewGaugeVec(prometheus.GaugeOpts{Name: "test_active", Help: "t"}, []string{}) + reg.MustRegister(counter, hist, gauge) + return reg, counter, hist, gauge +} + +func TestSnapshotSumAndSumBy(t *testing.T) { + reg, counter, _, gauge := newTestRegistry(t) + counter.WithLabelValues("npm", "timeout").Add(3) + counter.WithLabelValues("npm", "refused").Add(2) + counter.WithLabelValues("pypi", "timeout").Add(5) + gauge.WithLabelValues().Set(7) + + snap, err := GatherFrom(reg) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + + if got := snap.Sum("test_errors_total"); got != 10 { + t.Errorf("Sum = %v, want 10", got) + } + if got := snap.Sum("test_active"); got != 7 { + t.Errorf("gauge Sum = %v, want 7", got) + } + + byEco := snap.SumBy("test_errors_total", "ecosystem") + if byEco["npm"] != 5 || byEco["pypi"] != 5 { + t.Errorf("SumBy(ecosystem) = %v, want npm=5 pypi=5", byEco) + } +} + +func TestSnapshotHistogram(t *testing.T) { + reg, _, hist, _ := newTestRegistry(t) + hist.WithLabelValues("get").Observe(0.010) + hist.WithLabelValues("get").Observe(0.030) + hist.WithLabelValues("put").Observe(0.100) + + snap, err := GatherFrom(reg) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + + if got := snap.Count("test_duration_seconds"); got != 3 { + t.Errorf("Count = %d, want 3", got) + } + // (0.010 + 0.030 + 0.100) / 3 + if got := snap.Mean("test_duration_seconds"); got < 0.0466 || got > 0.0467 { + t.Errorf("Mean = %v, want ~0.04667", got) + } + + // A histogram carries no single value, so Value stays zero. + for _, s := range snap.Samples("test_duration_seconds") { + if s.Value != 0 { + t.Errorf("histogram sample carries Value %v, want 0", s.Value) + } + if s.Label("op") == "get" && s.Count != 2 { + t.Errorf("get count = %d, want 2", s.Count) + } + } +} + +func TestSnapshotSamplesAreOrdered(t *testing.T) { + reg, counter, _, _ := newTestRegistry(t) + for _, eco := range []string{"pypi", "npm", "cargo"} { + counter.WithLabelValues(eco, "timeout").Inc() + } + + snap, err := GatherFrom(reg) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + + var got []string + for _, s := range snap.Samples("test_errors_total") { + got = append(got, s.Label("ecosystem")) + } + want := []string{"cargo", "npm", "pypi"} + for i := range want { + if got[i] != want[i] { + t.Fatalf("sample order = %v, want %v (stable rendering depends on it)", got, want) + } + } +} + +func TestSnapshotMissingMetric(t *testing.T) { + reg, _, _, _ := newTestRegistry(t) + snap, err := GatherFrom(reg) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + + if got := snap.Samples("nope_total"); got != nil { + t.Errorf("Samples of an unknown metric = %v, want nil", got) + } + if got := snap.Sum("nope_total"); got != 0 { + t.Errorf("Sum of an unknown metric = %v, want 0", got) + } + if got := snap.Mean("nope_total"); got != 0 { + t.Errorf("Mean of an unknown metric = %v, want 0", got) + } +} + +// Gather reads the real registry, so every metric this package registers must +// come back. This is the check that a newly added metric is reachable by the UI. +func TestGatherSeesRegisteredMetrics(t *testing.T) { + RecordCacheHit("npm") + RecordScanError("npm", "clamav", "timeout") + + snap, err := Gather() + if err != nil { + t.Fatalf("Gather: %v", err) + } + + for _, name := range []string{"proxy_cache_hits_total", "proxy_scan_errors_total"} { + if len(snap.Samples(name)) == 0 { + t.Errorf("%s is registered but absent from the snapshot", name) + } + } +} diff --git a/internal/server/analytics.go b/internal/server/analytics.go new file mode 100644 index 00000000..b734edf8 --- /dev/null +++ b/internal/server/analytics.go @@ -0,0 +1,221 @@ +package server + +import ( + "fmt" + "net/http" + "strconv" + "strings" + + "github.com/git-pkgs/proxy/internal/database" + "github.com/git-pkgs/proxy/internal/metrics" +) + +// AnalyticsData contains data for rendering the analytics dashboard. +type AnalyticsData struct { + Layout + Totals AnalyticsTotals + Donut DonutView + EnrichmentStats EnrichmentStatsView + Ecosystems []EcosystemRow + Runtime RuntimeView + // StatsFailed records that the per-ecosystem query errored, so the page + // can say the figures are unavailable instead of claiming an empty cache. + StatsFailed bool + // StatsStale is set when the query errored but a previous snapshot was + // retained and is being shown. Without it a database that has been down + // for an hour renders hour-old figures as current, with the failure + // visible only in the logs. + StatsStale bool + // StatsAge is when the retained snapshot was read, phrased for display. + StatsAge string +} + +// AnalyticsTotals holds the headline figures across every ecosystem. +type AnalyticsTotals struct { + DownloadedBytes int64 + Downloaded string + Downloads string + CacheSize string + CachedArtifacts string + Packages string + Versions string + Ecosystems int + ActiveEcosystems int + // Amplification is accumulated download volume divided by the bytes + // currently held in cache: how many times over the cache has served + // what it stores. Empty when nothing is cached. + Amplification string +} + +// EcosystemRow is one ecosystem's row in the analytics table. +type EcosystemRow struct { + Ecosystem string + DownloadedBytes int64 + Downloaded string + Downloads string + CacheSize string + AvgArtifactSize string + Artifacts string + Packages string + Versions string + // SharePct is this ecosystem's share of the accumulated download total. + SharePct string +} + +func formatCount(n int64) string { + s := strconv.FormatInt(n, 10) + neg := strings.HasPrefix(s, "-") + if neg { + s = s[1:] + } + + var b strings.Builder + for i, digit := range s { + if i > 0 && (len(s)-i)%3 == 0 { + b.WriteByte(',') + } + b.WriteRune(digit) + } + if neg { + return "-" + b.String() + } + return b.String() +} + +// percentOf renders part/whole as a percentage with one decimal place, always +// as a bare number. formatPercent wraps it to catch the case where a real +// share rounds away to zero. +func percentOf(part, whole int64) string { + if whole <= 0 { + return "0.0" + } + return strconv.FormatFloat(float64(part)/float64(whole)*100, 'f', 1, 64) //nolint:mnd // percent +} + +// formatPercent renders a share for display. A share that is real but rounds +// to zero reads as "<0.1" rather than "0.0", which would claim the ecosystem +// served nothing. +func formatPercent(part, whole int64) string { + pct := percentOf(part, whole) + if pct == "0.0" && part > 0 && whole > 0 { + return "<0.1" + } + return pct +} + +// formatRatio renders an "N.Nx" multiplier, or "" when the denominator is zero. +func formatRatio(numerator, denominator int64) string { + if denominator <= 0 { + return "" + } + return fmt.Sprintf("%.1fx", float64(numerator)/float64(denominator)) +} + +func avgArtifactSize(cacheSize, artifacts int64) string { + if artifacts <= 0 { + return "0 B" + } + return formatSize(cacheSize / artifacts) +} + +// enrichmentStatsView maps enrichment totals onto the shared view struct used +// by both the dashboard and the analytics page. +func enrichmentStatsView(stats *database.EnrichmentStats) EnrichmentStatsView { + return EnrichmentStatsView{ + EnrichedPackages: stats.EnrichedPackages, + VulnSyncedPackages: stats.VulnSyncedPackages, + TotalVulnerabilities: stats.TotalVulnerabilities, + CriticalVulns: stats.CriticalVulns, + HighVulns: stats.HighVulns, + MediumVulns: stats.MediumVulns, + LowVulns: stats.LowVulns, + HasVulns: stats.TotalVulnerabilities > 0, + } +} + +// handleAnalytics renders the analytics dashboard: accumulated download volume +// in total and per ecosystem, alongside the cache and enrichment figures the +// main dashboard reports. +func (s *Server) handleAnalytics(w http.ResponseWriter, r *http.Request) { + ecosystems, err := s.ecoStats.Get(s.db) + statsFailed := err != nil + if err != nil { + s.logger.Error("failed to get ecosystem stats", "error", err) + } + + enrichStats, err := s.db.GetEnrichmentStats() + if err != nil { + s.logger.Error("failed to get enrichment stats", "error", err) + enrichStats = &database.EnrichmentStats{} + } + + data := AnalyticsData{ + Layout: s.layoutFor(r), + EnrichmentStats: enrichmentStatsView(enrichStats), + } + data.StatsFailed = statsFailed + data.StatsStale = statsFailed && len(ecosystems) > 0 + if data.StatsStale { + data.StatsAge = formatTimeAgo(s.ecoStats.SnapshotAt()) + } + data.Totals, data.Ecosystems = analyticsView(ecosystems) + data.Donut = donutView(ecosystems, data.Totals.DownloadedBytes, data.Totals.Downloaded) + + // Process-lifetime counters come from the Prometheus registry rather than + // the database; a failure here must not take the page down with it. + snap, err := metrics.Gather() + if err != nil { + s.logger.Error("failed to gather runtime metrics", "error", err) + } else { + data.Runtime = runtimeView(snap) + } + + if err := s.templates.Render(w, "analytics", data); err != nil { + s.logger.Error("failed to render analytics", "error", err) + } +} + +// analyticsView turns per-ecosystem database rows into the rendered totals and +// table rows, preserving the order they arrive in. +func analyticsView(stats []database.EcosystemStats) (AnalyticsTotals, []EcosystemRow) { + var totals AnalyticsTotals + var downloads, cacheSize, artifacts, packages, versions int64 + + for _, e := range stats { + totals.DownloadedBytes += e.DownloadedBytes + downloads += e.Downloads + cacheSize += e.CacheSize + artifacts += e.Artifacts + packages += e.Packages + versions += e.Versions + if e.DownloadedBytes > 0 { + totals.ActiveEcosystems++ + } + } + + totals.Downloaded = formatSize(totals.DownloadedBytes) + totals.Downloads = formatCount(downloads) + totals.CacheSize = formatSize(cacheSize) + totals.CachedArtifacts = formatCount(artifacts) + totals.Packages = formatCount(packages) + totals.Versions = formatCount(versions) + totals.Ecosystems = len(stats) + totals.Amplification = formatRatio(totals.DownloadedBytes, cacheSize) + + rows := make([]EcosystemRow, 0, len(stats)) + for _, e := range stats { + rows = append(rows, EcosystemRow{ + Ecosystem: e.Ecosystem, + DownloadedBytes: e.DownloadedBytes, + Downloaded: formatSize(e.DownloadedBytes), + Downloads: formatCount(e.Downloads), + CacheSize: formatSize(e.CacheSize), + AvgArtifactSize: avgArtifactSize(e.CacheSize, e.Artifacts), + Artifacts: formatCount(e.Artifacts), + Packages: formatCount(e.Packages), + Versions: formatCount(e.Versions), + SharePct: formatPercent(e.DownloadedBytes, totals.DownloadedBytes), + }) + } + return totals, rows +} diff --git a/internal/server/analytics_coverage_test.go b/internal/server/analytics_coverage_test.go new file mode 100644 index 00000000..8c5b7a2b --- /dev/null +++ b/internal/server/analytics_coverage_test.go @@ -0,0 +1,107 @@ +package server + +import ( + "strings" + "testing" + + "github.com/git-pkgs/proxy/internal/metrics" +) + +// metricSurface records where each registered metric is shown on /ui/analytics. +// +// This is a standing obligation, not a snapshot: registering a new metric +// anywhere in internal/metrics fails the build here until it has both a tile on +// the page and an entry below. That is the point -- the page is meant to be a +// complete view of /metrics -- but it is a cost paid by every future metric, so +// it is written down rather than discovered. +var metricSurface = map[string]string{ + // Database-derived, shown in the donut, the KPI row and the breakdown table. + "proxy_ecosystem_downloaded_bytes": "donut + breakdown table", + "proxy_ecosystem_artifact_downloads": "Downloads tile + breakdown table", + "proxy_ecosystem_cache_size_bytes": "breakdown table", + "proxy_ecosystem_cached_artifacts": "breakdown table", + "proxy_ecosystem_packages": "breakdown table", + "proxy_ecosystem_versions": "breakdown table", + "proxy_cache_size_bytes": "Cache size tile", + "proxy_cached_artifacts_total": "Cached artifacts tile", + + // Request-path counters. + "proxy_response_bytes_total": "Runtime: Served", + + // Registry-derived, shown in the Runtime card. + "proxy_requests_total": "Runtime: Requests + Responses by status", + "proxy_request_duration_seconds": "Runtime: Mean latency", + "proxy_active_requests": "Runtime: In flight", + "proxy_cache_hits_total": "Runtime: Cache lookups + hit rate", + "proxy_cache_misses_total": "Runtime: Cache lookups + hit rate", + "proxy_upstream_fetch_duration_seconds": "Runtime: Upstream fetches + mean fetch", + "proxy_upstream_errors_total": "Runtime: Upstream errors", + "proxy_storage_operation_duration_seconds": "Runtime: Storage operations", + "proxy_storage_errors_total": "Runtime: Storage errors", + "proxy_integrity_failures_total": "Runtime: Integrity failures", + "proxy_health_probe_failures_total": "Runtime: Health probe failures", + "proxy_circuit_breaker_state": "Runtime: Circuit breakers", + "proxy_circuit_breaker_trips_total": "Runtime: Circuit breaker trips", + "proxy_scan_duration_seconds": "Runtime: Pre-cache scanning -- Scans", + "proxy_scan_blocked_total": "Runtime: Pre-cache scanning -- Artifacts blocked", + "proxy_scan_errors_total": "Runtime: Pre-cache scanning -- Scan errors", +} + +// TestEveryMetricIsSurfaced fails when a metric is registered but has no home +// on the analytics page. The page is meant to be a complete view of what +// /metrics exposes, so a silently unsurfaced metric is a gap, not a detail. +func TestEveryMetricIsSurfaced(t *testing.T) { + // Touch every metric family so it is present in the registry output, since + // a vector with no observed label values gathers as nothing at all. + metrics.RecordRequest("npm", 200, 0) + metrics.RecordResponse("npm", 1) + metrics.RecordCacheHit("npm") + metrics.RecordCacheMiss("npm") + metrics.RecordUpstreamFetch("npm", 0) + metrics.RecordUpstreamError("npm", "fetch_failed") + metrics.RecordStorageOperation("get", 0) + metrics.RecordStorageError("get") + metrics.RecordIntegrityFailure("npm") + metrics.RecordHealthProbeFailure("write") + metrics.UpdateCircuitBreakerState("example.test", 0) + metrics.RecordCircuitBreakerTrip("example.test") + metrics.RecordScanResult("npm", "clamav", false, 0) + metrics.RecordScanError("npm", "clamav", "timeout") + metrics.UpdateCacheStats(1, 1) + metrics.UpdateEcosystemStats([]metrics.EcosystemStats{{Ecosystem: "npm", DownloadedBytes: 1}}) + + snap, err := metrics.Gather() + if err != nil { + t.Fatalf("Gather: %v", err) + } + + var registered []string + for _, name := range snap.Names() { + if strings.HasPrefix(name, "proxy_") { + registered = append(registered, name) + } + } + if len(registered) == 0 { + t.Fatal("no proxy_ metrics in the registry; the probe above is not working") + } + + for _, name := range registered { + if _, ok := metricSurface[name]; !ok { + t.Errorf("%s is registered but has no home on /ui/analytics; "+ + "surface it and add it to metricSurface", name) + } + } + + for name := range metricSurface { + found := false + for _, r := range registered { + if r == name { + found = true + break + } + } + if !found { + t.Errorf("metricSurface lists %s, but it is not registered any more; drop the entry", name) + } + } +} diff --git a/internal/server/analytics_donut.go b/internal/server/analytics_donut.go new file mode 100644 index 00000000..59abc857 --- /dev/null +++ b/internal/server/analytics_donut.go @@ -0,0 +1,208 @@ +package server + +import ( + "math" + "strconv" + + "github.com/git-pkgs/proxy/internal/database" +) + +// Donut geometry, in SVG user units. The ring is drawn as a dashed stroke on a +// circle rather than as arc paths: one dash per slice, which makes the 2px +// surface gap between slices fall out of the dash arithmetic. +const ( + donutSize = 260.0 + donutRadius = 104.0 + donutStroke = 30.0 + // donutGap is the surface-coloured separator between touching slices. + donutGap = 2.0 + // donutMinSlice keeps a slice that rounds to nothing from vanishing + // entirely; the legend and table carry its real value. + donutMinSlice = 1.5 + donutCircumference = 2 * math.Pi * donutRadius //nolint:mnd // circumference +) + +// donutSlices is the most slices the ring will draw. Part-to-whole reads at a +// glance only while the segment count stays small, so past this the tail is +// folded into a single "Other" slice and the full detail lives in the table +// below the chart. +const donutSlices = 6 + +const otherSliceLabel = "Other" + +// DonutSlice is one arc of the ring plus its legend row. +type DonutSlice struct { + // Index selects the fixed categorical slot. The hues themselves live in + // the .donut-slot-N and .swatch-N rules in the analytics template, so the + // validated palette has exactly one definition; slots are assigned here in + // sequence and never cycled. + Index int + Label string + Value string + SharePct string + Downloads string + CacheSize string + Dash string + Gap string + Offset string + // IsOther marks the folded tail, which has no single ecosystem behind it. + IsOther bool + // Members lists what was folded in, for the tooltip. + Members string +} + +// DonutView is the whole chart: the ring, its centre figure and its legend. +type DonutView struct { + Slices []DonutSlice + // Size, Radius, Stroke and Center are handed to the template so the SVG + // geometry has exactly one definition. + Size string + Radius string + Stroke string + Center string + // CenterValue is the accumulated download size across every ecosystem, + // stated in the middle of the ring. + CenterValue string + CenterLabel string + HasSlices bool +} + +// donutView folds the per-ecosystem rows into at most donutSlices arcs and +// computes the dash geometry for each. +func donutView(stats []database.EcosystemStats, totalBytes int64, centerValue string) DonutView { + view := DonutView{ + Size: trimFloat(donutSize), + Radius: trimFloat(donutRadius), + Stroke: trimFloat(donutStroke), + Center: trimFloat(donutSize / 2), //nolint:mnd // the centre of the viewBox + CenterValue: centerValue, + CenterLabel: "accumulated download size", + } + + // Only ecosystems that have actually served bytes get an arc; an + // ecosystem at zero would be an invisible slice with a legend entry + // claiming a share it does not have. + active := make([]database.EcosystemStats, 0, len(stats)) + for _, e := range stats { + if e.DownloadedBytes > 0 { + active = append(active, e) + } + } + if len(active) == 0 || totalBytes <= 0 { + return view + } + + head, tail := active, []database.EcosystemStats(nil) + if len(active) > donutSlices { + head, tail = active[:donutSlices-1], active[donutSlices-1:] + } + + slices := make([]DonutSlice, 0, donutSlices) + for i, e := range head { + slices = append(slices, DonutSlice{ + Index: i, + Label: ecosystemBadgeLabel(e.Ecosystem), + Value: formatSize(e.DownloadedBytes), + SharePct: formatPercent(e.DownloadedBytes, totalBytes), + Downloads: formatCount(e.Downloads), + CacheSize: formatSize(e.CacheSize), + }) + } + + if len(tail) > 0 { + var bytes, downloads, cacheSize int64 + members := "" + for i, e := range tail { + bytes += e.DownloadedBytes + downloads += e.Downloads + cacheSize += e.CacheSize + if i > 0 { + members += ", " + } + members += ecosystemBadgeLabel(e.Ecosystem) + } + slices = append(slices, DonutSlice{ + Index: donutSlices - 1, + Label: otherSliceLabel, + Value: formatSize(bytes), + SharePct: formatPercent(bytes, totalBytes), + Downloads: formatCount(downloads), + CacheSize: formatSize(cacheSize), + IsOther: true, + Members: members, + }) + } + + applyDonutGeometry(slices, active, tail, totalBytes) + view.Slices = slices + view.HasSlices = true + return view +} + +// applyDonutGeometry sets each slice's dash length and offset. +func applyDonutGeometry(slices []DonutSlice, active, tail []database.EcosystemStats, totalBytes int64) { + // Recover each slice's byte value in the same order the slices were built, + // so the geometry is driven by the numbers rather than by the rendered text. + values := make([]int64, 0, len(slices)) + head := active + if len(tail) > 0 { + head = active[:len(active)-len(tail)] + } + for _, e := range head { + values = append(values, e.DownloadedBytes) + } + if len(tail) > 0 { + var sum int64 + for _, e := range tail { + sum += e.DownloadedBytes + } + values = append(values, sum) + } + + dashes := make([]float64, len(slices)) + widest := 0 + for i := range slices { + arc := donutCircumference * float64(values[i]) / float64(totalBytes) + + dashes[i] = arc - donutGap + if dashes[i] < donutMinSlice { + dashes[i] = donutMinSlice + } + if dashes[i] > dashes[widest] { + widest = i + } + } + + // Widening a sliver to donutMinSlice buys visibility with room the ring + // does not have, and the overshoot would wrap the last slice back over the + // first. The widest slice gives the space back: it is the only one that can + // lose a couple of units without becoming unreadable, and at six slices the + // most that can be owed is donutSlices*(donutMinSlice+donutGap), far less + // than the circumference. + var needed float64 + for _, d := range dashes { + needed += d + donutGap + } + if excess := needed - donutCircumference; excess > 0 { + dashes[widest] = math.Max(donutMinSlice, dashes[widest]-excess) + } + + var consumed float64 + for i := range slices { + slices[i].Dash = trimFloat(dashes[i]) + slices[i].Gap = trimFloat(donutCircumference - dashes[i]) + // A dashoffset runs backwards around the circle, so the running total + // is negated to lay slices out clockwise from twelve o'clock. + slices[i].Offset = trimFloat(-consumed) + + // Advance by what is drawn: advancing by the smaller true arc would + // start the next slice underneath a widened one, and the later colour + // would win. + consumed += dashes[i] + donutGap + } +} + +// trimFloat renders an SVG coordinate without trailing zeroes. +func trimFloat(f float64) string { + return strconv.FormatFloat(f, 'f', -1, 64) +} diff --git a/internal/server/analytics_donut_test.go b/internal/server/analytics_donut_test.go new file mode 100644 index 00000000..d2e87388 --- /dev/null +++ b/internal/server/analytics_donut_test.go @@ -0,0 +1,211 @@ +package server + +import ( + "math" + "strconv" + "testing" + + "github.com/git-pkgs/proxy/internal/database" +) + +func eco(name string, bytes int64) database.EcosystemStats { + return database.EcosystemStats{Ecosystem: name, DownloadedBytes: bytes, Downloads: 1, CacheSize: bytes / 2} +} + +func TestDonutViewFoldsTailIntoOther(t *testing.T) { + // Nine active ecosystems: five keep their own slice, four fold into "Other". + stats := []database.EcosystemStats{ + eco("npm", 900), eco("pypi", 800), eco("maven", 700), eco("cargo", 600), + eco("gem", 500), eco("golang", 40), eco("nuget", 30), eco("deb", 20), eco("conda", 10), + } + var total int64 + for _, e := range stats { + total += e.DownloadedBytes + } + + view := donutView(stats, total, formatSize(total)) + + if len(view.Slices) != donutSlices { + t.Fatalf("expected %d slices, got %d", donutSlices, len(view.Slices)) + } + + last := view.Slices[donutSlices-1] + if !last.IsOther { + t.Fatalf("expected the final slice to be the folded tail, got %+v", last) + } + // 40 + 30 + 20 + 10 = 100 + if last.Value != formatSize(100) { + t.Errorf("Other value = %q, want %q", last.Value, formatSize(100)) + } + if last.Members != "golang, nuget, debian, conda" { + t.Errorf("Other members = %q, want the four folded ecosystems by display name", last.Members) + } + // "Other" must take the last slot, never one of a real ecosystem's. + if last.Index != donutSlices-1 { + t.Errorf("Other took slot %d, want %d", last.Index, donutSlices-1) + } + for i, s := range view.Slices[:donutSlices-1] { + if s.Index != i { + t.Errorf("slice %d took slot %d; slots must be assigned in sequence", i, s.Index) + } + } +} + +func TestDonutViewKeepsEveryEcosystemWhenFewEnough(t *testing.T) { + stats := []database.EcosystemStats{eco("npm", 600), eco("pypi", 400)} + view := donutView(stats, 1000, "1000 B") + + if len(view.Slices) != 2 { + t.Fatalf("expected 2 slices, got %d", len(view.Slices)) + } + for _, s := range view.Slices { + if s.IsOther { + t.Errorf("nothing should be folded with only 2 ecosystems, got %+v", s) + } + } + if view.Slices[0].SharePct != "60.0" || view.Slices[1].SharePct != "40.0" { + t.Errorf("shares = %q / %q, want 60.0 / 40.0", view.Slices[0].SharePct, view.Slices[1].SharePct) + } +} + +// The ring must account for the whole circle: the arcs, plus one gap each, +// have to add back up to the circumference. +func TestDonutGeometryCoversTheCircle(t *testing.T) { + stats := []database.EcosystemStats{eco("npm", 500), eco("pypi", 300), eco("maven", 200)} + view := donutView(stats, 1000, "1000 B") + + circumference := 2 * math.Pi * donutRadius + + var covered float64 + for _, s := range view.Slices { + dash, err := strconv.ParseFloat(s.Dash, 64) + if err != nil { + t.Fatalf("dash %q: %v", s.Dash, err) + } + gap, err := strconv.ParseFloat(s.Gap, 64) + if err != nil { + t.Fatalf("gap %q: %v", s.Gap, err) + } + // Each slice is drawn as one dash followed by a gap spanning the + // rest of the circle, so the pair always sums to the circumference. + if math.Abs(dash+gap-circumference) > 0.01 { + t.Errorf("slice %q: dash+gap = %v, want the circumference %v", s.Label, dash+gap, circumference) + } + covered += dash + donutGap + } + + if math.Abs(covered-circumference) > 0.01 { + t.Errorf("arcs plus gaps cover %v, want the full circumference %v", covered, circumference) + } +} + +// Offsets must advance monotonically so slices sit end to end instead of +// stacking on top of each other. +func TestDonutOffsetsAdvanceInOrder(t *testing.T) { + stats := []database.EcosystemStats{eco("npm", 500), eco("pypi", 300), eco("maven", 200)} + view := donutView(stats, 1000, "1000 B") + + circumference := 2 * math.Pi * donutRadius + want := []float64{0, -circumference * 0.5, -circumference * 0.8} + + for i, s := range view.Slices { + got, err := strconv.ParseFloat(s.Offset, 64) + if err != nil { + t.Fatalf("offset %q: %v", s.Offset, err) + } + if math.Abs(got-want[i]) > 0.01 { + t.Errorf("slice %d offset = %v, want %v", i, got, want[i]) + } + } +} + +// A slice too small to render still has to be visible; the legend and table +// carry its real value. +func TestDonutTinySliceStaysVisible(t *testing.T) { + stats := []database.EcosystemStats{eco("npm", 10_000_000), eco("composer", 1)} + view := donutView(stats, 10_000_001, "9.5 MB") + + tiny := view.Slices[1] + dash, err := strconv.ParseFloat(tiny.Dash, 64) + if err != nil { + t.Fatalf("dash %q: %v", tiny.Dash, err) + } + if dash < donutMinSlice { + t.Errorf("tiny slice dash = %v, want at least %v so it stays visible", dash, donutMinSlice) + } + if tiny.SharePct != "<0.1" { + t.Errorf("tiny slice share = %q, want %q", tiny.SharePct, "<0.1") + } +} + +func TestDonutViewEmpty(t *testing.T) { + if view := donutView(nil, 0, "0 B"); view.HasSlices || len(view.Slices) != 0 { + t.Errorf("expected an empty ring, got %+v", view) + } + + // Ecosystems that exist but have served nothing get no arc at all. + idle := []database.EcosystemStats{{Ecosystem: "rpm", Packages: 3}} + if view := donutView(idle, 0, "0 B"); view.HasSlices { + t.Errorf("an ecosystem with no traffic must not get a slice, got %+v", view.Slices) + } +} + +// A slice widened to stay visible must not be drawn over by its successor: +// the next offset has to start where this slice actually ends. +func TestDonutClampedSliceDoesNotOverdrawItsSuccessor(t *testing.T) { + // Four slivers far below the minimum, followed by a large slice. + stats := []database.EcosystemStats{ + eco("npm", 10_000_000), eco("a", 1), eco("b", 1), eco("c", 1), eco("d", 1), + } + view := donutView(stats, 10_000_004, "9.5 MB") + + for i := 0; i < len(view.Slices)-1; i++ { + dash := mustFloat(t, view.Slices[i].Dash) + start := -mustFloat(t, view.Slices[i].Offset) + nextStart := -mustFloat(t, view.Slices[i+1].Offset) + + end := start + dash + if nextStart < end-0.001 { + t.Errorf("slice %d (%q) is drawn to %v but slice %d starts at %v: they overlap", + i, view.Slices[i].Label, end, i+1, nextStart) + } + } +} + +func mustFloat(t *testing.T, s string) float64 { + t.Helper() + f, err := strconv.ParseFloat(s, 64) + if err != nil { + t.Fatalf("parsing %q: %v", s, err) + } + return f +} + +// Widening slivers must not push the ring past a full turn: the last slice +// would wrap back over the first and repaint the leading ecosystem's arc. +func TestDonutNeverExceedsOneTurn(t *testing.T) { + // One dominant ecosystem and five far below the visible minimum. + stats := []database.EcosystemStats{ + eco("npm", 100_000_000), + eco("a", 1), eco("b", 1), eco("c", 1), eco("d", 1), eco("e", 1), + } + view := donutView(stats, 100_000_005, "95.4 MB") + + if len(view.Slices) != 6 { + t.Fatalf("expected 6 slices, got %d", len(view.Slices)) + } + + last := view.Slices[len(view.Slices)-1] + end := -mustFloat(t, last.Offset) + mustFloat(t, last.Dash) + if end > donutCircumference+0.001 { + t.Errorf("ring is drawn to %v, past the circumference %v: the last slice "+ + "wraps over the first", end, donutCircumference) + } + + // Every sliver still has to be visible. + for _, s := range view.Slices[1:] { + if dash := mustFloat(t, s.Dash); dash < donutMinSlice { + t.Errorf("slice %q dash = %v, want at least %v", s.Label, dash, donutMinSlice) + } + } +} diff --git a/internal/server/analytics_runtime.go b/internal/server/analytics_runtime.go new file mode 100644 index 00000000..800f2ecb --- /dev/null +++ b/internal/server/analytics_runtime.go @@ -0,0 +1,235 @@ +package server + +import ( + "fmt" + "sort" + "strconv" + "time" + + "github.com/git-pkgs/proxy/internal/metrics" +) + +// RuntimeView holds the process-lifetime figures read out of the Prometheus +// registry. Unlike the cache figures, which are computed from the database and +// survive a restart, everything here starts at zero when the proxy starts. +type RuntimeView struct { + Available bool + + Requests string + ActiveRequests string + RequestMean string + StatusClasses []LabelledCount + CacheHits string + CacheMisses string + CacheHitRatio string + HasCacheTraffic bool + + UpstreamFetches string + UpstreamFetchMean string + UpstreamErrors []LabelledCount + + StorageOps []LabelledStat + StorageErrors []LabelledCount + IntegrityFailures []LabelledCount + ProbeFailures []LabelledCount + + Breakers []BreakerRow + BreakerTrips []LabelledCount + + Scans []LabelledStat + ScansBlocked []LabelledCount + ScanErrors []LabelledCount + ScanningOn bool + + ResponseBytes string +} + +// LabelledCount is a single counter series rendered as a row. +type LabelledCount struct { + Label string + Count string + // Bad marks a failure row, which the template tints. + Bad bool +} + +// LabelledStat is a histogram series rendered as a row: how many observations, +// and their mean. +type LabelledStat struct { + Label string + Count string + Mean string +} + +// BreakerRow is one upstream's circuit breaker state. +type BreakerRow struct { + Registry string + State string + Open bool +} + +// runtimeView shapes a registry snapshot for the analytics page. +func runtimeView(snap *metrics.Snapshot) RuntimeView { + if snap == nil { + return RuntimeView{} + } + + v := RuntimeView{Available: true} + + requests := snap.Sum("proxy_requests_total") + v.Requests = formatCount(int64(requests)) + v.ActiveRequests = formatCount(int64(snap.Sum("proxy_active_requests"))) + v.RequestMean = formatDuration(snap.Mean("proxy_request_duration_seconds")) + v.StatusClasses = statusClasses(snap) + + hits := snap.Sum("proxy_cache_hits_total") + misses := snap.Sum("proxy_cache_misses_total") + v.CacheHits = formatCount(int64(hits)) + v.CacheMisses = formatCount(int64(misses)) + v.HasCacheTraffic = hits+misses > 0 + if v.HasCacheTraffic { + v.CacheHitRatio = strconv.FormatFloat(hits/(hits+misses)*100, 'f', 1, 64) //nolint:mnd // percent + } + + v.UpstreamFetches = formatCount(int64(snap.Count("proxy_upstream_fetch_duration_seconds"))) + v.UpstreamFetchMean = formatDuration(snap.Mean("proxy_upstream_fetch_duration_seconds")) + v.UpstreamErrors = failureRows(snap, "proxy_upstream_errors_total", "ecosystem", "error_type") + + v.StorageOps = histogramRows(snap, "proxy_storage_operation_duration_seconds", "operation") + v.StorageErrors = failureRows(snap, "proxy_storage_errors_total", "operation") + v.IntegrityFailures = failureRows(snap, "proxy_integrity_failures_total", "ecosystem") + v.ProbeFailures = failureRows(snap, "proxy_health_probe_failures_total", "step") + + v.Breakers = breakerRows(snap) + v.BreakerTrips = failureRows(snap, "proxy_circuit_breaker_trips_total", "registry") + + v.Scans = histogramRows(snap, "proxy_scan_duration_seconds", "ecosystem", "scanner") + v.ScansBlocked = failureRows(snap, "proxy_scan_blocked_total", "ecosystem", "scanner") + v.ScanErrors = failureRows(snap, "proxy_scan_errors_total", "ecosystem", "scanner", "error_type") + v.ScanningOn = len(v.Scans) > 0 || len(v.ScansBlocked) > 0 || len(v.ScanErrors) > 0 + + v.ResponseBytes = formatSize(int64(snap.Sum("proxy_response_bytes_total"))) + + return v +} + +// statusClasses groups proxy_requests_total into 2xx/3xx/4xx/5xx buckets, which +// is the breakdown worth showing; the per-code detail stays in Prometheus. +func statusClasses(snap *metrics.Snapshot) []LabelledCount { + byClass := make(map[string]float64) + for _, s := range snap.Samples("proxy_requests_total") { + status := s.Label("status") + if len(status) == 0 { + continue + } + byClass[string(status[0])+"xx"] += s.Value + } + + classes := make([]string, 0, len(byClass)) + for c := range byClass { + classes = append(classes, c) + } + sort.Strings(classes) + + rows := make([]LabelledCount, 0, len(classes)) + for _, c := range classes { + rows = append(rows, LabelledCount{ + Label: c, + Count: formatCount(int64(byClass[c])), + Bad: c == "5xx", + }) + } + return rows +} + +// failureRows renders every non-zero series of a failure counter, keyed by the +// joined values of the given labels. Zero-valued series are dropped: a counter +// that has never fired carries no information, and a wall of zeroes buries the +// rows that matter. +func failureRows(snap *metrics.Snapshot, name string, labels ...string) []LabelledCount { + var rows []LabelledCount + for _, s := range snap.Samples(name) { + if s.Value == 0 { + continue + } + rows = append(rows, LabelledCount{ + Label: joinLabels(s, labels...), + Count: formatCount(int64(s.Value)), + Bad: true, + }) + } + return rows +} + +// histogramRows renders every observed series of a histogram with its count and mean. +func histogramRows(snap *metrics.Snapshot, name string, labels ...string) []LabelledStat { + var rows []LabelledStat + for _, s := range snap.Samples(name) { + if s.Count == 0 { + continue + } + rows = append(rows, LabelledStat{ + Label: joinLabels(s, labels...), + Count: formatCount(int64(s.Count)), + Mean: formatDuration(s.Mean()), + }) + } + return rows +} + +func breakerRows(snap *metrics.Snapshot) []BreakerRow { + var rows []BreakerRow + for _, s := range snap.Samples("proxy_circuit_breaker_state") { + state := "closed" + switch s.Value { + case 1: + state = "half-open" + case 2: //nolint:mnd // 2 = open, per the gauge's documented encoding + state = "open" + } + rows = append(rows, BreakerRow{ + Registry: s.Label("registry"), + State: state, + Open: s.Value > 0, + }) + } + return rows +} + +func joinLabels(s metrics.Sample, labels ...string) string { + out := "" + for _, l := range labels { + v := s.Label(l) + if v == "" { + continue + } + if out != "" { + out += " · " + } + out += v + } + if out == "" { + return "-" + } + return out +} + +// formatDuration renders a mean latency at a sensible unit. Storage operations +// land in microseconds and upstream fetches in seconds, so a fixed unit would +// read as either 0.000 or an unreadable pile of digits. +func formatDuration(seconds float64) string { + if seconds <= 0 { + return "-" + } + + d := time.Duration(seconds * float64(time.Second)) + switch { + case d < time.Microsecond: + return fmt.Sprintf("%.0f ns", float64(d)) + case d < time.Millisecond: + return fmt.Sprintf("%.0f µs", float64(d)/float64(time.Microsecond)) + case d < time.Second: + return fmt.Sprintf("%.1f ms", float64(d)/float64(time.Millisecond)) + default: + return fmt.Sprintf("%.2f s", d.Seconds()) + } +} diff --git a/internal/server/analytics_runtime_test.go b/internal/server/analytics_runtime_test.go new file mode 100644 index 00000000..6cab5587 --- /dev/null +++ b/internal/server/analytics_runtime_test.go @@ -0,0 +1,184 @@ +package server + +import ( + "testing" + + "github.com/git-pkgs/proxy/internal/metrics" + "github.com/prometheus/client_golang/prometheus" +) + +func TestFormatDuration(t *testing.T) { + tests := []struct { + seconds float64 + want string + }{ + {0, "-"}, + {-1, "-"}, + {0.0000004, "400 ns"}, + {0.00041, "410 µs"}, + {0.0081, "8.1 ms"}, + {1.25, "1.25 s"}, + } + for _, tc := range tests { + if got := formatDuration(tc.seconds); got != tc.want { + t.Errorf("formatDuration(%v) = %q, want %q", tc.seconds, got, tc.want) + } + } +} + +func TestRuntimeViewNilSnapshot(t *testing.T) { + if v := runtimeView(nil); v.Available { + t.Error("a nil snapshot must not report itself as available") + } +} + +func TestRuntimeViewShapesRegistry(t *testing.T) { + reg := prometheus.NewRegistry() + + requests := prometheus.NewCounterVec( + prometheus.CounterOpts{Name: "proxy_requests_total", Help: "t"}, []string{"ecosystem", "status"}) + scanErrors := prometheus.NewCounterVec( + prometheus.CounterOpts{Name: "proxy_scan_errors_total", Help: "t"}, []string{"ecosystem", "scanner", "error_type"}) + hits := prometheus.NewCounterVec( + prometheus.CounterOpts{Name: "proxy_cache_hits_total", Help: "t"}, []string{"ecosystem"}) + misses := prometheus.NewCounterVec( + prometheus.CounterOpts{Name: "proxy_cache_misses_total", Help: "t"}, []string{"ecosystem"}) + breaker := prometheus.NewGaugeVec( + prometheus.GaugeOpts{Name: "proxy_circuit_breaker_state", Help: "t"}, []string{"registry"}) + storage := prometheus.NewHistogramVec( + prometheus.HistogramOpts{Name: "proxy_storage_operation_duration_seconds", Help: "t"}, []string{"operation"}) + reg.MustRegister(requests, scanErrors, hits, misses, breaker, storage) + + requests.WithLabelValues("npm", "200").Add(90) + requests.WithLabelValues("npm", "304").Add(5) + requests.WithLabelValues("npm", "500").Add(5) + scanErrors.WithLabelValues("npm", "clamav", "timeout").Add(3) + hits.WithLabelValues("npm").Add(80) + misses.WithLabelValues("npm").Add(20) + breaker.WithLabelValues("registry.npmjs.org").Set(0) + breaker.WithLabelValues("static.crates.io").Set(2) + storage.WithLabelValues("get").Observe(0.002) + + snap, err := metrics.GatherFrom(reg) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + v := runtimeView(snap) + + if !v.Available { + t.Fatal("expected the view to be available") + } + if v.Requests != "100" { + t.Errorf("Requests = %q, want %q", v.Requests, "100") + } + if v.CacheHitRatio != "80.0" || !v.HasCacheTraffic { + t.Errorf("CacheHitRatio = %q (traffic=%v), want 80.0", v.CacheHitRatio, v.HasCacheTraffic) + } + + // Status codes collapse to classes; 5xx is flagged as a failure row. + wantClasses := map[string]struct { + count string + bad bool + }{"2xx": {"90", false}, "3xx": {"5", false}, "5xx": {"5", true}} + if len(v.StatusClasses) != len(wantClasses) { + t.Fatalf("got %d status classes, want %d: %+v", len(v.StatusClasses), len(wantClasses), v.StatusClasses) + } + for _, row := range v.StatusClasses { + want, ok := wantClasses[row.Label] + if !ok { + t.Errorf("unexpected status class %q", row.Label) + continue + } + if row.Count != want.count || row.Bad != want.bad { + t.Errorf("%s = %q (bad=%v), want %q (bad=%v)", row.Label, row.Count, row.Bad, want.count, want.bad) + } + } + + // The metric that prompted this: it must reach the page with its labels joined. + if len(v.ScanErrors) != 1 { + t.Fatalf("expected 1 scan error row, got %+v", v.ScanErrors) + } + if v.ScanErrors[0].Label != "npm · clamav · timeout" || v.ScanErrors[0].Count != "3" { + t.Errorf("scan error row = %+v, want npm · clamav · timeout = 3", v.ScanErrors[0]) + } + if !v.ScanningOn { + t.Error("scanning should read as on once a scan metric carries a value") + } + + if len(v.Breakers) != 2 { + t.Fatalf("expected 2 breakers, got %+v", v.Breakers) + } + var open, closed int + for _, b := range v.Breakers { + if b.Open { + open++ + if b.State != "open" { + t.Errorf("open breaker state = %q, want %q", b.State, "open") + } + } else { + closed++ + } + } + if open != 1 || closed != 1 { + t.Errorf("breakers: %d open / %d closed, want 1 / 1", open, closed) + } + + if len(v.StorageOps) != 1 || v.StorageOps[0].Label != "get" || v.StorageOps[0].Mean != "2.0 ms" { + t.Errorf("StorageOps = %+v, want one 'get' row with a 2.0 ms mean", v.StorageOps) + } +} + +// A counter sitting at zero carries no information and would bury the rows +// that do, so it is left out. +func TestRuntimeViewDropsZeroCounters(t *testing.T) { + reg := prometheus.NewRegistry() + errs := prometheus.NewCounterVec( + prometheus.CounterOpts{Name: "proxy_storage_errors_total", Help: "t"}, []string{"operation"}) + reg.MustRegister(errs) + + // Touching a label creates the series at zero. + errs.WithLabelValues("get") + errs.WithLabelValues("put").Add(2) + + snap, err := metrics.GatherFrom(reg) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + v := runtimeView(snap) + + if len(v.StorageErrors) != 1 { + t.Fatalf("expected only the non-zero row, got %+v", v.StorageErrors) + } + if v.StorageErrors[0].Label != "put" { + t.Errorf("kept row = %q, want %q", v.StorageErrors[0].Label, "put") + } +} + +// An empty registry must still render, reporting nothing rather than dividing +// by zero or showing a hit rate of NaN. +func TestRuntimeViewEmptyRegistry(t *testing.T) { + snap, err := metrics.GatherFrom(prometheus.NewRegistry()) + if err != nil { + t.Fatalf("GatherFrom: %v", err) + } + v := runtimeView(snap) + + if !v.Available { + t.Error("an empty registry is still a valid snapshot") + } + if v.HasCacheTraffic { + t.Error("no traffic recorded, so HasCacheTraffic must be false") + } + if v.CacheHitRatio != "" { + t.Errorf("CacheHitRatio = %q, want empty when nothing was recorded", v.CacheHitRatio) + } + if v.Requests != "0" { + t.Errorf("Requests = %q, want %q", v.Requests, "0") + } + if v.RequestMean != "-" { + t.Errorf("RequestMean = %q, want a dash", v.RequestMean) + } + if v.ScanningOn { + t.Error("scanning must read as off with no scan metrics") + } +} diff --git a/internal/server/analytics_test.go b/internal/server/analytics_test.go new file mode 100644 index 00000000..3ce5289a --- /dev/null +++ b/internal/server/analytics_test.go @@ -0,0 +1,277 @@ +package server + +import ( + "net/http/httptest" + "strings" + "testing" + + "github.com/git-pkgs/proxy/internal/database" +) + +func TestFormatCount(t *testing.T) { + tests := []struct { + in int64 + want string + }{ + {0, "0"}, + {7, "7"}, + {100, "100"}, + {1000, "1,000"}, + {12004, "12,004"}, + {1000000, "1,000,000"}, + {-4200, "-4,200"}, + } + for _, tc := range tests { + if got := formatCount(tc.in); got != tc.want { + t.Errorf("formatCount(%d) = %q, want %q", tc.in, got, tc.want) + } + } +} + +func TestPercentOf(t *testing.T) { + tests := []struct { + part, whole int64 + want string + }{ + {1, 4, "25.0"}, + {2, 3, "66.7"}, + {0, 100, "0.0"}, + {5, 5, "100.0"}, + // A zero total must not divide by zero. + {0, 0, "0.0"}, + {7, 0, "0.0"}, + // A negligible share still rounds to a plain number here; only + // formatPercent turns it into "<0.1". + {1, 1000000, "0.0"}, + } + for _, tc := range tests { + if got := percentOf(tc.part, tc.whole); got != tc.want { + t.Errorf("percentOf(%d, %d) = %q, want %q", tc.part, tc.whole, got, tc.want) + } + } +} + +func TestFormatPercent(t *testing.T) { + tests := []struct { + part, whole int64 + want string + }{ + {1, 4, "25.0"}, + {0, 100, "0.0"}, + {0, 0, "0.0"}, + // A real but tiny share must not be reported as zero. + {1, 1000000, "<0.1"}, + {7_800_000, 18_000_000_000, "<0.1"}, + } + for _, tc := range tests { + if got := formatPercent(tc.part, tc.whole); got != tc.want { + t.Errorf("formatPercent(%d, %d) = %q, want %q", tc.part, tc.whole, got, tc.want) + } + } +} + +func TestFormatRatio(t *testing.T) { + if got := formatRatio(3500, 1000); got != "3.5x" { + t.Errorf("formatRatio(3500, 1000) = %q, want %q", got, "3.5x") + } + // Nothing cached means no meaningful multiplier, not "+Infx". + if got := formatRatio(3500, 0); got != "" { + t.Errorf("formatRatio(3500, 0) = %q, want empty", got) + } +} + +func TestAvgArtifactSize(t *testing.T) { + if got := avgArtifactSize(1000, 4); got != "250 B" { + t.Errorf("avgArtifactSize(1000, 4) = %q, want %q", got, "250 B") + } + if got := avgArtifactSize(0, 0); got != "0 B" { + t.Errorf("avgArtifactSize(0, 0) = %q, want %q", got, "0 B") + } +} + +func TestAnalyticsView(t *testing.T) { + stats := []database.EcosystemStats{ + {Ecosystem: "npm", Packages: 2, Versions: 2, Artifacts: 2, CacheSize: 1500, Downloads: 4, DownloadedBytes: 3000}, + {Ecosystem: "cargo", Packages: 1, Versions: 1, Artifacts: 1, CacheSize: 200, Downloads: 5, DownloadedBytes: 1000}, + {Ecosystem: "gem", Packages: 3, Versions: 4}, + } + + totals, rows := analyticsView(stats) + + if totals.DownloadedBytes != 4000 { + t.Errorf("DownloadedBytes = %d, want 4000", totals.DownloadedBytes) + } + if totals.Downloads != "9" { + t.Errorf("Downloads = %q, want %q", totals.Downloads, "9") + } + if totals.CachedArtifacts != "3" { + t.Errorf("CachedArtifacts = %q, want %q", totals.CachedArtifacts, "3") + } + if totals.Packages != "6" { + t.Errorf("Packages = %q, want %q", totals.Packages, "6") + } + if totals.Versions != "7" { + t.Errorf("Versions = %q, want %q", totals.Versions, "7") + } + if totals.Ecosystems != 3 { + t.Errorf("Ecosystems = %d, want 3", totals.Ecosystems) + } + // gem has served nothing, so it is known but not active. + if totals.ActiveEcosystems != 2 { + t.Errorf("ActiveEcosystems = %d, want 2", totals.ActiveEcosystems) + } + // 4000 bytes served from 1700 bytes stored. + if totals.Amplification != "2.4x" { + t.Errorf("Amplification = %q, want %q", totals.Amplification, "2.4x") + } + + if len(rows) != 3 { + t.Fatalf("expected 3 rows, got %d", len(rows)) + } + + // Shares are of the grand total. + if rows[0].SharePct != "75.0" { + t.Errorf("npm share = %q, want 75.0", rows[0].SharePct) + } + if rows[1].SharePct != "25.0" { + t.Errorf("cargo share = %q, want 25.0", rows[1].SharePct) + } + // gem served nothing, so it contributes no share but still gets a row. + if rows[2].SharePct != "0.0" || rows[2].Ecosystem != "gem" { + t.Errorf("gem row = %+v, want a 0.0%% share", rows[2]) + } + if rows[0].AvgArtifactSize != "750 B" { + t.Errorf("npm AvgArtifactSize = %q, want %q", rows[0].AvgArtifactSize, "750 B") + } +} + +// With nothing cached at all the view must stay renderable rather than dividing +// by a zero total. +func TestAnalyticsViewEmpty(t *testing.T) { + totals, rows := analyticsView(nil) + + if len(rows) != 0 { + t.Errorf("expected no rows, got %+v", rows) + } + if totals.DownloadedBytes != 0 { + t.Errorf("DownloadedBytes = %d, want 0", totals.DownloadedBytes) + } + if totals.Downloaded != "0 B" { + t.Errorf("Downloaded = %q, want %q", totals.Downloaded, "0 B") + } + if totals.Amplification != "" { + t.Errorf("Amplification = %q, want empty", totals.Amplification) + } +} + +func TestAnalyticsPageRendersTotals(t *testing.T) { + tpl := &Templates{} + w := httptest.NewRecorder() + + totals, rows := analyticsView([]database.EcosystemStats{ + {Ecosystem: "npm", Packages: 1, Versions: 1, Artifacts: 1, CacheSize: 1000, Downloads: 3, DownloadedBytes: 3000}, + }) + data := AnalyticsData{Totals: totals, Ecosystems: rows} + data.Donut = donutView([]database.EcosystemStats{ + {Ecosystem: "npm", Packages: 1, Versions: 1, Artifacts: 1, CacheSize: 1000, Downloads: 3, DownloadedBytes: 3000}, + }, totals.DownloadedBytes, totals.Downloaded) + + if err := tpl.Render(w, "analytics", data); err != nil { + t.Fatalf("Render: %v", err) + } + + body := w.Body.String() + if strings.Contains(body, "ZgotmplZ") { + t.Error("a template value was sanitized away; check the donut dash geometry") + } + for _, want := range []string{ + "accumulated download size", + "2.9 KB", // the total, stated in the middle of the ring + "donut-slot-0", // the single slice takes the first categorical slot + "3.0x", // 3000 bytes served from 1000 stored + "/ui/analytics", // nav link renders on the page itself + } { + if !strings.Contains(body, want) { + t.Errorf("rendered page missing %q", want) + } + } +} + +// A retained snapshot must be labelled as one. The cache deliberately serves +// the last good rows when the query fails, so without this the page presents +// arbitrarily old figures as current and the only trace is a log line. +func TestAnalyticsPageFlagsStaleFigures(t *testing.T) { + tpl := &Templates{} + w := httptest.NewRecorder() + + totals, rows := analyticsView([]database.EcosystemStats{ + {Ecosystem: "npm", Artifacts: 1, CacheSize: 1000, Downloads: 3, DownloadedBytes: 3000}, + }) + data := AnalyticsData{ + Totals: totals, + Ecosystems: rows, + StatsFailed: true, + StatsStale: true, + StatsAge: "2 hours ago", + } + + if err := tpl.Render(w, "analytics", data); err != nil { + t.Fatalf("Render: %v", err) + } + + body := w.Body.String() + for _, want := range []string{"Figures may be out of date", "2 hours ago"} { + if !strings.Contains(body, want) { + t.Errorf("rendered page missing %q", want) + } + } + // The figures themselves still render — stale beats absent. + if !strings.Contains(body, "2.9 KB") { + t.Error("the retained snapshot was not rendered alongside the warning") + } +} + +// A failure with nothing retained keeps the existing unavailable message and +// must not also claim the figures below are stale, since there are none. +func TestAnalyticsPageStaleBannerNeedsASnapshot(t *testing.T) { + tpl := &Templates{} + w := httptest.NewRecorder() + + if err := tpl.Render(w, "analytics", AnalyticsData{StatsFailed: true}); err != nil { + t.Fatalf("Render: %v", err) + } + + body := w.Body.String() + if strings.Contains(body, "Figures may be out of date") { + t.Error("stale banner rendered with no retained snapshot") + } + if !strings.Contains(body, "the database query failed") { + t.Error("expected the unavailable message when the query failed with no snapshot") + } +} + +func TestAnalyticsPageRendersWithoutTraffic(t *testing.T) { + tpl := &Templates{} + w := httptest.NewRecorder() + + if err := tpl.Render(w, "analytics", AnalyticsData{}); err != nil { + t.Fatalf("Render: %v", err) + } + if body := w.Body.String(); !strings.Contains(body, "Nothing has been served from cache yet") { + t.Error("expected the empty-state message when no traffic has been recorded") + } +} + +func TestEcosystemMetricsMapping(t *testing.T) { + got := ecosystemMetrics([]database.EcosystemStats{ + {Ecosystem: "npm", Packages: 1, Versions: 2, Artifacts: 3, CacheSize: 4, Downloads: 5, DownloadedBytes: 6}, + }) + if len(got) != 1 { + t.Fatalf("expected 1 snapshot, got %d", len(got)) + } + m := got[0] + if m.Ecosystem != "npm" || m.Packages != 1 || m.Versions != 2 || m.Artifacts != 3 || + m.CacheSize != 4 || m.Downloads != 5 || m.DownloadedBytes != 6 { + t.Errorf("fields did not map across: %+v", m) + } +} diff --git a/internal/server/ecosystem_cache.go b/internal/server/ecosystem_cache.go new file mode 100644 index 00000000..83e8c0d7 --- /dev/null +++ b/internal/server/ecosystem_cache.go @@ -0,0 +1,94 @@ +package server + +import ( + "sync" + "time" + + "github.com/git-pkgs/proxy/internal/database" +) + +// ecosystemStatsTTL bounds how stale a served snapshot can be. The metrics +// refresh loop runs on the same cadence, so a scrape and a page load a moment +// apart report the same figures rather than two slightly different ones. +const ecosystemStatsTTL = time.Minute + +// ecosystemStatsErrorTTL is the retry interval after a failed read. It is short +// so a recovered database is picked up quickly, but not zero, so a database +// that is down is not queried once per request. +const ecosystemStatsErrorTTL = 5 * time.Second + +// ecosystemStatsCache memoizes GetEcosystemStats. +// +// The query is three grouped aggregations over the whole artifact table, which +// is fine once a minute for the metrics gauges but not once per page load: the +// UI is unauthenticated, so without this a client reloading /ui/analytics would +// keep the database busy for as long as it cared to. +// +// The zero value is usable, so a Server assembled as a struct literal -- as +// tests do -- needs no constructor and cannot end up with a nil cache. +type ecosystemStatsCache struct { + ttl time.Duration + + mu sync.Mutex + stats []database.EcosystemStats + err error + lastAt time.Time + // goodAt is when stats was last read successfully, which is not lastAt + // once reads start failing and the retained snapshot is served on. + goodAt time.Time +} + +// SnapshotAt reports when the rows currently held were last read successfully. +// The zero time means no read has ever succeeded. +func (c *ecosystemStatsCache) SnapshotAt() time.Time { + c.mu.Lock() + defer c.mu.Unlock() + return c.goodAt +} + +func (c *ecosystemStatsCache) interval() time.Duration { + if c.err != nil { + return ecosystemStatsErrorTTL + } + if c.ttl <= 0 { + return ecosystemStatsTTL + } + return c.ttl +} + +// Get returns the cached rows, refreshing them from db when they have aged out. +// +// A failed read returns the last good snapshot alongside the error, so a +// caller can choose to serve stale figures rather than fail outright. +func (c *ecosystemStatsCache) Get(db *database.DB) ([]database.EcosystemStats, error) { + c.mu.Lock() + defer c.mu.Unlock() + + if !c.lastAt.IsZero() && time.Since(c.lastAt) < c.interval() { + return c.stats, c.err + } + return c.refreshLocked(db) +} + +// Refresh forces a read, bypassing the TTL. The metrics loop uses it so its own +// tick is never served a snapshot that is about to expire. +func (c *ecosystemStatsCache) Refresh(db *database.DB) ([]database.EcosystemStats, error) { + c.mu.Lock() + defer c.mu.Unlock() + return c.refreshLocked(db) +} + +func (c *ecosystemStatsCache) refreshLocked(db *database.DB) ([]database.EcosystemStats, error) { + stats, err := db.GetEcosystemStats() + c.lastAt = time.Now() + c.err = err + if err != nil { + // Keep the last good snapshot: a transient failure should not erase + // figures the caller could still usefully show. + return c.stats, err + } + + c.stats = stats + c.goodAt = c.lastAt + return c.stats, nil +} diff --git a/internal/server/ecosystem_cache_test.go b/internal/server/ecosystem_cache_test.go new file mode 100644 index 00000000..5323ad1b --- /dev/null +++ b/internal/server/ecosystem_cache_test.go @@ -0,0 +1,199 @@ +package server + +import ( + "path/filepath" + "testing" + "time" + + "github.com/git-pkgs/proxy/internal/database" +) + +func cacheTestDB(t *testing.T) *database.DB { + t.Helper() + db, err := database.Create(filepath.Join(t.TempDir(), "cache.db")) + if err != nil { + t.Fatalf("Create: %v", err) + } + t.Cleanup(func() { _ = db.Close() }) + return db +} + +// The zero value must work, because Server is assembled as a struct literal in +// places that never call a constructor. +func TestEcosystemStatsCacheZeroValueIsUsable(t *testing.T) { + var c ecosystemStatsCache + + if _, err := c.Get(cacheTestDB(t)); err != nil { + t.Fatalf("Get on a zero-value cache: %v", err) + } + if c.interval() != ecosystemStatsTTL { + t.Errorf("interval = %v, want the default %v", c.interval(), ecosystemStatsTTL) + } +} + +// A second call inside the TTL must not touch the database again — that is the +// whole point, since /stats is unauthenticated and pollable. +func TestEcosystemStatsCacheServesWithinTTL(t *testing.T) { + db := cacheTestDB(t) + c := ecosystemStatsCache{ttl: time.Hour} + + if _, err := c.Get(db); err != nil { + t.Fatalf("Get: %v", err) + } + first := c.lastAt + + if _, err := c.Get(db); err != nil { + t.Fatalf("Get: %v", err) + } + if !c.lastAt.Equal(first) { + t.Error("a second Get inside the TTL re-queried the database") + } +} + +func TestEcosystemStatsCacheRefreshesAfterTTL(t *testing.T) { + db := cacheTestDB(t) + // A TTL already elapsed by the time the second call lands. + c := ecosystemStatsCache{ttl: time.Nanosecond} + + if _, err := c.Get(db); err != nil { + t.Fatalf("Get: %v", err) + } + first := c.lastAt + + time.Sleep(time.Millisecond) + if _, err := c.Get(db); err != nil { + t.Fatalf("Get: %v", err) + } + if c.lastAt.Equal(first) { + t.Error("the cache did not refresh after its TTL elapsed") + } +} + +// Refresh ignores the TTL, so the metrics loop always publishes a fresh read. +func TestEcosystemStatsCacheRefreshBypassesTTL(t *testing.T) { + db := cacheTestDB(t) + c := ecosystemStatsCache{ttl: time.Hour} + + if _, err := c.Get(db); err != nil { + t.Fatalf("Get: %v", err) + } + first := c.lastAt + + time.Sleep(time.Millisecond) + if _, err := c.Refresh(db); err != nil { + t.Fatalf("Refresh: %v", err) + } + if c.lastAt.Equal(first) { + t.Error("Refresh honoured the TTL; it must always re-read") + } +} + +// A transient failure must not erase a snapshot the caller could still show. +func TestEcosystemStatsCacheKeepsLastGoodSnapshotOnError(t *testing.T) { + good := cacheTestDB(t) + seedCachePackage(t, good, "npm", "lodash") + + c := ecosystemStatsCache{ttl: time.Nanosecond} + stats, err := c.Get(good) + if err != nil { + t.Fatalf("Get: %v", err) + } + if len(stats) != 1 { + t.Fatalf("expected 1 ecosystem from the seeded database, got %+v", stats) + } + + // A closed handle stands in for the database going away. + broken := cacheTestDB(t) + if err := broken.Close(); err != nil { + t.Fatalf("Close: %v", err) + } + + time.Sleep(time.Millisecond) + stats, err = c.Get(broken) + if err == nil { + t.Fatal("expected an error from the closed database") + } + if len(stats) != 1 || stats[0].Ecosystem != "npm" { + t.Errorf("last good snapshot was discarded: got %+v", stats) + } +} + +// An error must not be pinned for the full TTL, or a recovered database stays +// invisible for a minute. +func TestEcosystemStatsCacheRetriesSoonerAfterError(t *testing.T) { + broken := cacheTestDB(t) + if err := broken.Close(); err != nil { + t.Fatalf("Close: %v", err) + } + + c := ecosystemStatsCache{ttl: time.Hour} + if _, err := c.Get(broken); err == nil { + t.Fatal("expected an error") + } + if got := c.interval(); got != ecosystemStatsErrorTTL { + t.Errorf("retry interval after an error = %v, want %v", got, ecosystemStatsErrorTTL) + } +} + +// A closed database makes the query fail; the error must reach the caller +// rather than being swallowed into an empty-looking result. +func TestEcosystemStatsCacheSurfacesErrors(t *testing.T) { + db, err := database.Create(filepath.Join(t.TempDir(), "closed.db")) + if err != nil { + t.Fatalf("Create: %v", err) + } + if err := db.Close(); err != nil { + t.Fatalf("Close: %v", err) + } + + var c ecosystemStatsCache + if _, err := c.Get(db); err == nil { + t.Error("expected an error from a closed database") + } +} + +func seedCachePackage(t *testing.T, db *database.DB, ecosystem, name string) { + t.Helper() + if err := db.UpsertPackage(&database.Package{ + PURL: "pkg:" + ecosystem + "/" + name, + Ecosystem: ecosystem, + Name: name, + }); err != nil { + t.Fatalf("UpsertPackage: %v", err) + } +} + +// The snapshot timestamp must track the last successful read, not the last +// attempt. The analytics page dates its "figures may be out of date" banner +// from it, and a timestamp that advanced on every failed retry would report an +// hour-old snapshot as seconds old. +func TestEcosystemStatsCacheSnapshotAtTracksSuccessOnly(t *testing.T) { + good := cacheTestDB(t) + seedCachePackage(t, good, "npm", "lodash") + + c := ecosystemStatsCache{ttl: time.Nanosecond} + if c.SnapshotAt(); !c.SnapshotAt().IsZero() { + t.Fatalf("SnapshotAt on a fresh cache = %v, want the zero time", c.SnapshotAt()) + } + if _, err := c.Get(good); err != nil { + t.Fatalf("Get: %v", err) + } + + at := c.SnapshotAt() + if at.IsZero() { + t.Fatal("SnapshotAt is zero after a successful read") + } + + broken := cacheTestDB(t) + if err := broken.Close(); err != nil { + t.Fatalf("Close: %v", err) + } + + time.Sleep(2 * time.Millisecond) + if _, err := c.Get(broken); err == nil { + t.Fatal("expected an error from the closed database") + } + if got := c.SnapshotAt(); !got.Equal(at) { + t.Errorf("SnapshotAt moved on a failed read: %v, want %v", got, at) + } +} diff --git a/internal/server/middleware.go b/internal/server/middleware.go index 53ed6e46..b86ae6e6 100644 --- a/internal/server/middleware.go +++ b/internal/server/middleware.go @@ -56,10 +56,15 @@ func (s *Server) LoggerMiddleware(next http.Handler) http.Handler { "path", r.URL.Path, "status", rw.status, "duration", duration, + "bytes", rw.bytes, "remote", r.RemoteAddr) + // Scrapes of /metrics would otherwise attribute themselves, + // burying real callers under whatever polls the proxy most often. if r.URL.Path != "/metrics" { - metrics.RecordRequest(requestEcosystem(r.URL.Path), rw.status, duration) + ecosystem := requestEcosystem(r.URL.Path) + metrics.RecordRequest(ecosystem, rw.status, duration) + metrics.RecordResponse(ecosystem, rw.bytes) } if s.accessLog != nil { diff --git a/internal/server/server.go b/internal/server/server.go index e7fc5ddf..4d13a572 100644 --- a/internal/server/server.go +++ b/internal/server/server.go @@ -31,6 +31,7 @@ // Web UI (HTML), mounted under /ui so reverse proxies can gate it // separately from the package endpoints: // - /ui/ - Dashboard +// - /ui/analytics - Download and cache analytics // - /ui/install - Client configuration guide // - /ui/packages - List all cached packages // - /ui/search - Search packages @@ -113,6 +114,7 @@ type Server struct { accessLog *accesslog.Logger ecr *ecrTokens breakers *breakerMonitor + ecoStats ecosystemStatsCache } // New creates a new Server with the given configuration. @@ -304,6 +306,7 @@ func (s *Server) serve(listener net.Listener) error { r.Route("/ui", func(ui chi.Router) { ui.Mount("/static", http.StripPrefix("/ui/static/", staticHandler())) ui.Get("/", s.handleRoot) + ui.Get("/analytics", s.handleAnalytics) ui.Get("/install", s.handleInstall) ui.Get("/search", s.handleSearch) ui.Get("/packages", s.handlePackagesList) @@ -515,7 +518,7 @@ func (s *Server) updateCacheStats() { } metrics.UpdateCacheStats(stats.TotalSize, stats.TotalArtifacts) - ecosystems, err := s.db.GetEcosystemStats() + ecosystems, err := s.ecoStats.Refresh(s.db) if err != nil { s.logger.Warn("failed to get ecosystem stats for metrics", "error", err) return @@ -633,16 +636,7 @@ func (s *Server) handleRoot(w http.ResponseWriter, r *http.Request) { TotalPackages: stats.TotalPackages, TotalVersions: stats.TotalVersions, }, - EnrichmentStats: EnrichmentStatsView{ - EnrichedPackages: enrichStats.EnrichedPackages, - VulnSyncedPackages: enrichStats.VulnSyncedPackages, - TotalVulnerabilities: enrichStats.TotalVulnerabilities, - CriticalVulns: enrichStats.CriticalVulns, - HighVulns: enrichStats.HighVulns, - MediumVulns: enrichStats.MediumVulns, - LowVulns: enrichStats.LowVulns, - HasVulns: enrichStats.TotalVulnerabilities > 0, - }, + EnrichmentStats: enrichmentStatsView(enrichStats), } for _, p := range popular { @@ -1236,10 +1230,18 @@ func categorizeLicense(license sql.NullString) string { return categorizeLicenseCSS(license.String) } -// responseWriter wraps http.ResponseWriter to capture status code. +// responseWriter wraps http.ResponseWriter to capture the status code and the +// number of body bytes written, which is what a client actually downloaded. type responseWriter struct { http.ResponseWriter status int + bytes int64 +} + +func (rw *responseWriter) Write(b []byte) (int, error) { + n, err := rw.ResponseWriter.Write(b) + rw.bytes += int64(n) + return n, err } // Unwrap lets ResponseController reach capabilities such as flushing when a diff --git a/internal/server/server_test.go b/internal/server/server_test.go index 313a13e1..57b80808 100644 --- a/internal/server/server_test.go +++ b/internal/server/server_test.go @@ -123,6 +123,7 @@ func newTestServer(t *testing.T) *testServer { r.Route("/ui", func(ui chi.Router) { ui.Mount("/static", http.StripPrefix("/ui/static/", staticHandler())) ui.Get("/", s.handleRoot) + ui.Get("/analytics", s.handleAnalytics) ui.Get("/install", s.handleInstall) ui.Get("/search", s.handleSearch) ui.Get("/packages", s.handlePackagesList) diff --git a/internal/server/templates.go b/internal/server/templates.go index 217da413..010abad8 100644 --- a/internal/server/templates.go +++ b/internal/server/templates.go @@ -2,6 +2,7 @@ package server import ( "embed" + "fmt" "html/template" "net/http" "path/filepath" @@ -29,6 +30,7 @@ func (t *Templates) load() error { "supportedEcosystems": supportedEcosystems, "ecosystemBadgeClass": ecosystemBadgeClasses, "ecosystemBadgeLabel": ecosystemBadgeLabel, + "dict": templateDict, } pageFiles, err := templatesFS.ReadDir("templates/pages") @@ -78,3 +80,24 @@ func (t *Templates) Render(w http.ResponseWriter, pageName string, data any) err return tmpl.ExecuteTemplate(w, "base", data) } + +// templateDict builds a map from alternating key/value arguments, so a +// component can be invoked with named parameters rather than being handed a +// whole page struct it would have to reach through. +func templateDict(values ...any) (map[string]any, error) { + const pair = 2 + + if len(values)%pair != 0 { + return nil, fmt.Errorf("dict: got %d arguments, want an even number of key/value pairs", len(values)) + } + + out := make(map[string]any, len(values)/pair) + for i := 0; i < len(values); i += pair { + key, ok := values[i].(string) + if !ok { + return nil, fmt.Errorf("dict: key %d is %T, want string", i, values[i]) + } + out[key] = values[i+1] + } + return out, nil +} diff --git a/internal/server/templates/components/runtime_metrics.html b/internal/server/templates/components/runtime_metrics.html new file mode 100644 index 00000000..83ff5f7f --- /dev/null +++ b/internal/server/templates/components/runtime_metrics.html @@ -0,0 +1,162 @@ +{{define "runtime_metrics"}} +
+
+

Runtime

+

+ Counters held in this process, not in the database: they start at zero when the + proxy starts and are lost on restart. Everything above comes from the database + instead, and survives a restart — so a low count here alongside a large one + above just means the proxy started recently. +

+
+ +
+
+
+
Requests
+
{{.Requests}}
+
+
+
In flight
+
{{.ActiveRequests}}
+
+
+
Mean latency
+
{{.RequestMean}}
+
+
+
Cache hit rate
+
{{if .HasCacheTraffic}}{{.CacheHitRatio}}%{{else}}—{{end}}
+
+
+
Upstream fetches
+
{{.UpstreamFetches}} · {{.UpstreamFetchMean}}
+
+
+
Served
+
{{.ResponseBytes}}
+
+
+ +
+ {{template "runtime_counts" dict "Title" "Responses by status" "Rows" .StatusClasses "Empty" "No requests yet"}} + +
+

Cache lookups

+
+
+
hits
+
{{.CacheHits}}
+
+
+
misses
+
{{.CacheMisses}}
+
+
+
+ + {{template "runtime_stats" dict "Title" "Storage operations" "Rows" .StorageOps "Empty" "No storage operations yet"}} + {{template "runtime_counts" dict "Title" "Upstream errors" "Rows" .UpstreamErrors "Empty" "None"}} + {{template "runtime_counts" dict "Title" "Storage errors" "Rows" .StorageErrors "Empty" "None"}} + {{template "runtime_counts" dict "Title" "Integrity failures" "Rows" .IntegrityFailures "Empty" "None"}} + {{template "runtime_counts" dict "Title" "Health probe failures" "Rows" .ProbeFailures "Empty" "None"}} + {{template "runtime_counts" dict "Title" "Circuit breaker trips" "Rows" .BreakerTrips "Empty" "None"}} + +
+

Circuit breakers

+ {{if .Breakers}} +
+ {{range .Breakers}} +
+
{{.Registry}}
+
+ {{if .Open}} + + {{.State}} + + {{else}} + + {{.State}} + + {{end}} +
+
+ {{end}} +
+ {{else}} +

+ None reported — a breaker appears once its upstream has been fetched from. +

+ {{end}} +
+
+ + {{if .ScanningOn}} +
+

Pre-cache scanning

+
+ {{template "runtime_stats" dict "Title" "Scans" "Rows" .Scans "Empty" "No scans yet"}} + {{template "runtime_counts" dict "Title" "Artifacts blocked" "Rows" .ScansBlocked "Empty" "None"}} + {{template "runtime_counts" dict "Title" "Scan errors" "Rows" .ScanErrors "Empty" "None"}} +
+
+ {{else}} +
+

Pre-cache scanning

+

+ Not configured. Enable it under scanning in the + config to have artifacts scanned before they are committed to the cache. +

+
+ {{end}} +
+
+{{end}} + +{{/* runtime_counts renders a labelled counter list, tinting failure rows. */}} +{{define "runtime_counts"}} +
+

{{.Title}}

+ {{if .Rows}} +
+ {{range .Rows}} +
+
{{.Label}}
+
{{.Count}}
+
+ {{end}} +
+ {{else}} +

{{.Empty}}

+ {{end}} +
+{{end}} + +{{/* runtime_stats renders a histogram list: how many observations, and their mean. */}} +{{define "runtime_stats"}} +
+

{{.Title}}

+ {{if .Rows}} + + + + + + + + + + {{range .Rows}} + + + + + + {{end}} + +
countmean
{{.Label}}{{.Count}}{{.Mean}}
+ {{else}} +

{{.Empty}}

+ {{end}} +
+{{end}} diff --git a/internal/server/templates/components/security_overview.html b/internal/server/templates/components/security_overview.html new file mode 100644 index 00000000..3ba93e32 --- /dev/null +++ b/internal/server/templates/components/security_overview.html @@ -0,0 +1,30 @@ +{{define "security_overview"}} +
+
+

Security Overview

+
+
+
+
+
{{.CriticalVulns}}
+
Critical
+
+
+
{{.HighVulns}}
+
High
+
+
+
{{.MediumVulns}}
+
Medium
+
+
+
{{.LowVulns}}
+
Low
+
+
+

+ {{.TotalVulnerabilities}} vulnerabilities tracked across {{.VulnSyncedPackages}} packages +

+
+
+{{end}} diff --git a/internal/server/templates/layout/header.html b/internal/server/templates/layout/header.html index b3103f1f..560fe2b3 100644 --- a/internal/server/templates/layout/header.html +++ b/internal/server/templates/layout/header.html @@ -54,6 +54,7 @@ {{end}} {{define "nav_links"}} +Analytics Install Health API diff --git a/internal/server/templates/pages/analytics.html b/internal/server/templates/pages/analytics.html new file mode 100644 index 00000000..da97a7ce --- /dev/null +++ b/internal/server/templates/pages/analytics.html @@ -0,0 +1,302 @@ +{{define "title"}}Analytics · git-pkgs proxy{{end}} + +{{define "head"}} + +{{end}} + +{{define "content"}} +
+

Analytics

+

+ Accumulated download volume and cache composition across + {{.Totals.Ecosystems}} ecosystem{{if ne .Totals.Ecosystems 1}}s{{end}}. +

+
+ +{{if .StatsStale}} +
+ Figures may be out of date. + The database query is failing, so the download and cache numbers below are + the last snapshot that could be read{{if .StatsAge}}, from {{.StatsAge}}{{end}}. + Check the proxy logs. +
+{{end}} + +
+
+

Download size by ecosystem

+

Bytes served from cache

+
+ {{if .Donut.HasSlices}} +
+
+ + {{range .Donut.Slices}} + + {{.Label}}: {{.Value}} ({{.SharePct}}%) + + {{end}} + +
+
{{.Donut.CenterValue}}
+
{{.Donut.CenterLabel}}
+
+
+ + +
    + {{range .Donut.Slices}} +
  • + + + {{.Label}} + {{- if .IsOther}}({{.Members}}){{end}} + + {{.Value}} + {{.SharePct}}% +
  • + {{end}} +
+
+ {{else if .StatsFailed}} +
+ Download figures are unavailable — the database query failed. Check the proxy logs. +
+ {{else}} +
Nothing has been served from cache yet
+ {{end}} +
+ + +
+
+
Downloads
+
{{.Totals.Downloads}}
+
+
+
Cache size
+
{{.Totals.CacheSize}}
+
+
+
Cached artifacts
+
{{.Totals.CachedArtifacts}}
+
+
+
Packages
+
{{.Totals.Packages}}
+
+
+
Versions
+
{{.Totals.Versions}}
+
+
+
Ecosystems in use
+
{{.Totals.ActiveEcosystems}} / {{.Totals.Ecosystems}}
+
+
+ +{{if .Totals.Amplification}} +

+ The cache has served {{.Totals.Amplification}} the bytes it currently stores. +

+{{end}} + + +
+
+

Per-ecosystem breakdown

+

Every ecosystem, including any folded into the ring's "Other" slice

+
+ {{if .Ecosystems}} +
+ + + + + + + + + + + + + + + + {{range .Ecosystems}} + + + + + + + + + + + + {{end}} + +
EcosystemDownloadedShareDownloadsCache sizeArtifactsAvg sizePackagesVersions
+ {{template "ecosystem_badge" .Ecosystem}} + {{.Downloaded}}{{.SharePct}}%{{.Downloads}}{{.CacheSize}}{{.Artifacts}}{{.AvgArtifactSize}}{{.Packages}}{{.Versions}}
+
+ {{else if .StatsFailed}} +
+ The per-ecosystem query failed, so this table is unavailable. Check the proxy logs. +
+ {{else}} +
No packages cached yet
+ {{end}} +
+ +{{if .EnrichmentStats.HasVulns}} +{{template "security_overview" .EnrichmentStats}} +{{end}} + +{{if .Runtime.Available}} +{{template "runtime_metrics" .Runtime}} +{{end}} + +

+ Every figure on this page is also exported for Prometheus at + /metrics. + This page shows current state only — nothing here is a time series. For history, + trends and alerting, scrape /metrics and use the + Grafana dashboard shipped in deploy/grafana/. +

+{{end}} + +{{define "scripts"}} + +{{end}} diff --git a/internal/server/templates/pages/dashboard.html b/internal/server/templates/pages/dashboard.html index 9b9a9e2c..e48e205f 100644 --- a/internal/server/templates/pages/dashboard.html +++ b/internal/server/templates/pages/dashboard.html @@ -22,35 +22,7 @@ {{if .EnrichmentStats.HasVulns}} - -
-
-

Security Overview

-
-
-
-
-
{{.EnrichmentStats.CriticalVulns}}
-
Critical
-
-
-
{{.EnrichmentStats.HighVulns}}
-
High
-
-
-
{{.EnrichmentStats.MediumVulns}}
-
Medium
-
-
-
{{.EnrichmentStats.LowVulns}}
-
Low
-
-
-

- {{.EnrichmentStats.TotalVulnerabilities}} vulnerabilities tracked across {{.EnrichmentStats.VulnSyncedPackages}} packages -

-
-
+{{template "security_overview" .EnrichmentStats}} {{end}} diff --git a/internal/server/templates_test.go b/internal/server/templates_test.go index a2d3b7da..c68a69ab 100644 --- a/internal/server/templates_test.go +++ b/internal/server/templates_test.go @@ -37,6 +37,46 @@ func TestTemplatesRenderAllPages(t *testing.T) { {Ecosystem: "cargo", Name: "serde", Version: "1.0.0", Size: "200 KB", CachedAt: "1 hour ago"}, }, }}, + {"analytics", AnalyticsData{ + Totals: AnalyticsTotals{ + DownloadedBytes: 1_500_000_000, + Downloaded: "1.4 GB", + Downloads: "12,004", + CacheSize: "420.0 MB", + CachedArtifacts: "1,204", + Packages: "310", + Versions: "902", + Ecosystems: 3, + ActiveEcosystems: 2, + Amplification: "3.4x", + }, + EnrichmentStats: EnrichmentStatsView{TotalVulnerabilities: 3, CriticalVulns: 1, HasVulns: true}, + Ecosystems: []EcosystemRow{ + {Ecosystem: "npm", DownloadedBytes: 1_000_000_000, Downloaded: "953.7 MB", Downloads: "9,000", CacheSize: "300.0 MB", AvgArtifactSize: "120 KB", Artifacts: "900", Packages: "200", Versions: "700", SharePct: "66.7"}, + {Ecosystem: "cargo", DownloadedBytes: 500_000_000, Downloaded: "476.8 MB", Downloads: "3,004", CacheSize: "120.0 MB", AvgArtifactSize: "400 KB", Artifacts: "300", Packages: "100", Versions: "200", SharePct: "33.3"}, + {Ecosystem: "rpm", Downloaded: "0 B", Downloads: "0", CacheSize: "0 B", AvgArtifactSize: "0 B", Artifacts: "0", Packages: "10", Versions: "2", SharePct: "0.0"}, + }, + Donut: donutView([]database.EcosystemStats{ + {Ecosystem: "npm", CacheSize: 300, Downloads: 9000, DownloadedBytes: 1_000_000_000}, + {Ecosystem: "cargo", CacheSize: 120, Downloads: 3004, DownloadedBytes: 500_000_000}, + }, 1_500_000_000, "1.4 GB"), + Runtime: RuntimeView{ + Available: true, + Requests: "12,004", + ActiveRequests: "2", + RequestMean: "8.1 ms", + CacheHits: "9,000", + CacheMisses: "1,000", + CacheHitRatio: "90.0", + HasCacheTraffic: true, + StatusClasses: []LabelledCount{{Label: "2xx", Count: "11,900"}, {Label: "5xx", Count: "104", Bad: true}}, + StorageOps: []LabelledStat{{Label: "get", Count: "9,000", Mean: "412 µs"}}, + ScanErrors: []LabelledCount{{Label: "npm · clamav · timeout", Count: "3", Bad: true}}, + ScanningOn: true, + Breakers: []BreakerRow{{Registry: "registry.npmjs.org", State: "closed"}}, + }, + }}, + {"analytics", AnalyticsData{}}, {"install", struct { Layout BaseURL string From 4b8c5a89fb9d2678428fc75276d356b835a4e53d Mon Sep 17 00:00:00 2001 From: Wicliff Wolda Date: Thu, 1 Oct 2026 14:32:20 +0200 Subject: [PATCH 3/6] Report the per-ecosystem breakdown from GET /stats - Add downloaded_bytes, downloads and an ecosystems array, served from the snapshot the gauges and the page already share - Add stats_unavailable so a failed aggregation is distinguishable from an idle proxy - Regenerate the OpenAPI spec Part 4 of 6 splitting #381 up. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 2 + docs/swagger/docs.go | 49 ++++++++++++ docs/swagger/swagger.json | 49 ++++++++++++ internal/server/server.go | 59 +++++++++++++-- internal/server/stats_test.go | 139 ++++++++++++++++++++++++++++++++++ 5 files changed, 293 insertions(+), 5 deletions(-) create mode 100644 internal/server/stats_test.go diff --git a/README.md b/README.md index 3390bd96..079002b0 100644 --- a/README.md +++ b/README.md @@ -1184,6 +1184,8 @@ The proxy stores no time series. The page reads the database and the in-process For history, trends and alerting, scrape `/metrics` with Prometheus. That is the intended split: the UI answers "what is true now", Prometheus answers "what happened". +The same figures are available as JSON from `GET /stats`, which reports `downloaded_bytes`, `downloads` and an `ecosystems` array carrying the per-ecosystem breakdown, served from the same 60-second snapshot the page and the gauges read. When the aggregation fails with no snapshot to fall back on, the response carries `stats_unavailable: true` rather than passing zeros off as a count -- the endpoint keeps answering with the artifact count and cache size either way. + #### Three ecosystem label sets `ecosystem` means three slightly different things across `/metrics`, and queries that join across them need to know which. diff --git a/docs/swagger/docs.go b/docs/swagger/docs.go index cc88b4c8..dd50bf92 100644 --- a/docs/swagger/docs.go +++ b/docs/swagger/docs.go @@ -559,6 +559,35 @@ const docTemplate = `{ } } }, + "server.EcosystemStatsEntry": { + "type": "object", + "properties": { + "cache_size_bytes": { + "type": "integer" + }, + "cached_artifacts": { + "type": "integer" + }, + "downloaded": { + "type": "string" + }, + "downloaded_bytes": { + "type": "integer" + }, + "downloads": { + "type": "integer" + }, + "ecosystem": { + "type": "string" + }, + "packages": { + "type": "integer" + }, + "versions": { + "type": "integer" + } + } + }, "server.ErrorResponse": { "type": "object", "properties": { @@ -806,6 +835,26 @@ const docTemplate = `{ "database_path": { "type": "string" }, + "downloaded": { + "type": "string" + }, + "downloaded_bytes": { + "description": "DownloadedBytes is the accumulated download volume across every\necosystem: cache hits multiplied by the artifact size they served.", + "type": "integer" + }, + "downloads": { + "type": "integer" + }, + "ecosystems": { + "type": "array", + "items": { + "$ref": "#/definitions/server.EcosystemStatsEntry" + } + }, + "stats_unavailable": { + "description": "StatsUnavailable distinguishes a proxy that has served nothing from one\nwhose aggregation failed with no snapshot to fall back on. Without it\nboth report zeros and an empty array.", + "type": "boolean" + }, "storage_url": { "type": "string" }, diff --git a/docs/swagger/swagger.json b/docs/swagger/swagger.json index 5db91660..50007b6e 100644 --- a/docs/swagger/swagger.json +++ b/docs/swagger/swagger.json @@ -552,6 +552,35 @@ } } }, + "server.EcosystemStatsEntry": { + "type": "object", + "properties": { + "cache_size_bytes": { + "type": "integer" + }, + "cached_artifacts": { + "type": "integer" + }, + "downloaded": { + "type": "string" + }, + "downloaded_bytes": { + "type": "integer" + }, + "downloads": { + "type": "integer" + }, + "ecosystem": { + "type": "string" + }, + "packages": { + "type": "integer" + }, + "versions": { + "type": "integer" + } + } + }, "server.ErrorResponse": { "type": "object", "properties": { @@ -799,6 +828,26 @@ "database_path": { "type": "string" }, + "downloaded": { + "type": "string" + }, + "downloaded_bytes": { + "description": "DownloadedBytes is the accumulated download volume across every\necosystem: cache hits multiplied by the artifact size they served.", + "type": "integer" + }, + "downloads": { + "type": "integer" + }, + "ecosystems": { + "type": "array", + "items": { + "$ref": "#/definitions/server.EcosystemStatsEntry" + } + }, + "stats_unavailable": { + "description": "StatsUnavailable distinguishes a proxy that has served nothing from one\nwhose aggregation failed with no snapshot to fall back on. Without it\nboth report zeros and an empty array.", + "type": "boolean" + }, "storage_url": { "type": "string" }, diff --git a/internal/server/server.go b/internal/server/server.go index 4d13a572..77829122 100644 --- a/internal/server/server.go +++ b/internal/server/server.go @@ -1123,6 +1123,28 @@ type StatsResponse struct { TotalSizeHuman string `json:"total_size"` StorageURL string `json:"storage_url"` DatabasePath string `json:"database_path"` + // DownloadedBytes is the accumulated download volume across every + // ecosystem: cache hits multiplied by the artifact size they served. + DownloadedBytes int64 `json:"downloaded_bytes"` + DownloadedBytesHuman string `json:"downloaded"` + Downloads int64 `json:"downloads"` + Ecosystems []EcosystemStatsEntry `json:"ecosystems"` + // StatsUnavailable distinguishes a proxy that has served nothing from one + // whose aggregation failed with no snapshot to fall back on. Without it + // both report zeros and an empty array. + StatsUnavailable bool `json:"stats_unavailable,omitempty"` +} + +// EcosystemStatsEntry is one ecosystem's slice of the cache statistics. +type EcosystemStatsEntry struct { + Ecosystem string `json:"ecosystem"` + DownloadedBytes int64 `json:"downloaded_bytes"` + Downloaded string `json:"downloaded"` + Downloads int64 `json:"downloads"` + CacheSize int64 `json:"cache_size_bytes"` + Artifacts int64 `json:"cached_artifacts"` + Packages int64 `json:"packages"` + Versions int64 `json:"versions"` } // handleStats returns cache statistics. @@ -1147,15 +1169,42 @@ func (s *Server) handleStats(w http.ResponseWriter, r *http.Request) { return } + // A failing per-ecosystem aggregation must not take down an endpoint that + // answered from two cheap counters before it existed. Get returns the last + // good snapshot alongside the error, so the breakdown is served stale when + // there is one, and flagged unavailable when there is not. + ecosystems, statsErr := s.ecoStats.Get(s.db) + if statsErr != nil { + s.logger.Error("failed to get ecosystem stats for /stats", "error", statsErr) + } + _ = ctx // Could use for storage.UsedSpace if needed stats := StatsResponse{ - CachedArtifacts: count, - TotalSize: size, - TotalSizeHuman: formatSize(size), - StorageURL: s.storage.URL(), - DatabasePath: s.cfg.Database.String(), + CachedArtifacts: count, + TotalSize: size, + TotalSizeHuman: formatSize(size), + Ecosystems: make([]EcosystemStatsEntry, 0, len(ecosystems)), + StatsUnavailable: statsErr != nil && len(ecosystems) == 0, + StorageURL: s.storage.URL(), + DatabasePath: s.cfg.Database.String(), + } + + for _, e := range ecosystems { + stats.DownloadedBytes += e.DownloadedBytes + stats.Downloads += e.Downloads + stats.Ecosystems = append(stats.Ecosystems, EcosystemStatsEntry{ + Ecosystem: e.Ecosystem, + DownloadedBytes: e.DownloadedBytes, + Downloaded: formatSize(e.DownloadedBytes), + Downloads: e.Downloads, + CacheSize: e.CacheSize, + Artifacts: e.Artifacts, + Packages: e.Packages, + Versions: e.Versions, + }) } + stats.DownloadedBytesHuman = formatSize(stats.DownloadedBytes) w.Header().Set("Content-Type", "application/json") _ = json.NewEncoder(w).Encode(stats) diff --git a/internal/server/stats_test.go b/internal/server/stats_test.go new file mode 100644 index 00000000..467f3c16 --- /dev/null +++ b/internal/server/stats_test.go @@ -0,0 +1,139 @@ +package server + +import ( + "database/sql" + "encoding/json" + "net/http" + "net/http/httptest" + "testing" + "time" + + "github.com/git-pkgs/proxy/internal/database" +) + +func seedCachedArtifact(t *testing.T, db *database.DB, ecosystem, name, version string, size, hits int64) { + t.Helper() + + pkgPURL := "pkg:" + ecosystem + "/" + name + versionPURL := pkgPURL + "@" + version + + if err := db.UpsertPackage(&database.Package{PURL: pkgPURL, Ecosystem: ecosystem, Name: name}); err != nil { + t.Fatalf("UpsertPackage: %v", err) + } + if err := db.UpsertVersion(&database.Version{PURL: versionPURL, PackagePURL: pkgPURL}); err != nil { + t.Fatalf("UpsertVersion: %v", err) + } + if err := db.UpsertArtifact(&database.Artifact{ + VersionPURL: versionPURL, + Filename: name + "-" + version + ".tgz", + UpstreamURL: "https://example.test/" + name, + StoragePath: sql.NullString{String: "objects/" + name, Valid: true}, + Size: sql.NullInt64{Int64: size, Valid: true}, + FetchedAt: sql.NullTime{Time: time.Now(), Valid: true}, + HitCount: hits, + }); err != nil { + t.Fatalf("UpsertArtifact: %v", err) + } +} + +func getStats(t *testing.T, ts *testServer) StatsResponse { + t.Helper() + + w := httptest.NewRecorder() + ts.handler.ServeHTTP(w, httptest.NewRequest("GET", "/stats", nil)) + if w.Code != http.StatusOK { + t.Fatalf("GET /stats = %d, want 200", w.Code) + } + + var stats StatsResponse + if err := json.NewDecoder(w.Body).Decode(&stats); err != nil { + t.Fatalf("decode: %v", err) + } + return stats +} + +func TestStatsReportsEcosystemBreakdown(t *testing.T) { + ts := newTestServer(t) + defer ts.close() + + seedCachedArtifact(t, ts.db, "npm", "lodash", "4.17.21", 1000, 3) + seedCachedArtifact(t, ts.db, "cargo", "serde", "1.0.0", 500, 2) + + stats := getStats(t, ts) + + if stats.DownloadedBytes != 4000 { + t.Errorf("downloaded_bytes = %d, want 4000", stats.DownloadedBytes) + } + if stats.Downloads != 5 { + t.Errorf("downloads = %d, want 5", stats.Downloads) + } + if stats.DownloadedBytesHuman == "" { + t.Error("downloaded is empty; the human-readable total is part of the shape") + } + if stats.StatsUnavailable { + t.Error("stats_unavailable set on a healthy aggregation") + } + + byEcosystem := make(map[string]EcosystemStatsEntry, len(stats.Ecosystems)) + for _, e := range stats.Ecosystems { + byEcosystem[e.Ecosystem] = e + } + npm, ok := byEcosystem["npm"] + if !ok { + t.Fatalf("no npm row in %v", byEcosystem) + } + if npm.DownloadedBytes != 3000 || npm.Downloads != 3 || npm.CacheSize != 1000 { + t.Errorf("npm row = %+v, want 3000 bytes over 3 downloads and 1000 cached", npm) + } +} + +// A failed aggregation with nothing to fall back on must not read as an idle +// proxy: the totals are zero either way, so the flag is the only thing +// telling a consumer which of the two it is looking at. +func TestStatsFlagsUnavailableFigures(t *testing.T) { + ts := newTestServer(t) + defer ts.close() + + // GetEcosystemStats counts packages first; the artifact count and cache + // size the endpoint already reported do not touch that table. + if _, err := ts.db.Exec(`DROP TABLE packages`); err != nil { + t.Fatalf("DROP TABLE packages: %v", err) + } + + stats := getStats(t, ts) + + if !stats.StatsUnavailable { + t.Error("stats_unavailable not set after the aggregation failed with no snapshot") + } + if len(stats.Ecosystems) != 0 { + t.Errorf("ecosystems = %v, want empty", stats.Ecosystems) + } + if stats.StorageURL == "" { + t.Error("the rest of the response must still be served") + } +} + +// A retained snapshot is served on with no flag: the figures are stale, not +// absent, and the page is where staleness is surfaced to a human. +func TestStatsServesRetainedSnapshot(t *testing.T) { + ts := newTestServer(t) + defer ts.close() + + seedCachedArtifact(t, ts.db, "npm", "lodash", "4.17.21", 1000, 3) + if got := getStats(t, ts).DownloadedBytes; got != 3000 { + t.Fatalf("downloaded_bytes = %d, want 3000 before the failure", got) + } + + if _, err := ts.db.Exec(`DROP TABLE packages`); err != nil { + t.Fatalf("DROP TABLE packages: %v", err) + } + ts.server.ecoStats.lastAt = time.Time{} + + stats := getStats(t, ts) + if stats.StatsUnavailable { + t.Error("stats_unavailable set although a snapshot was retained") + } + if stats.DownloadedBytes != 3000 { + t.Errorf("downloaded_bytes = %d, want the retained 3000", stats.DownloadedBytes) + } +} From 36fe6abef2c73b9c138040be13d579a9ec730217 Mon Sep 17 00:00:00 2001 From: Wicliff Wolda Date: Thu, 1 Oct 2026 14:32:20 +0200 Subject: [PATCH 4/6] Add a Grafana dashboard - 33 panels, no hardcoded data source UID - Three ecosystem variables, because the label means three different things across /metrics: the package record, the request path, and the handler's own name Part 5 of 6 splitting #381 up. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 6 + deploy/grafana/git-pkgs-proxy.json | 2739 ++++++++++++++++++++++++++++ 2 files changed, 2745 insertions(+) create mode 100644 deploy/grafana/git-pkgs-proxy.json diff --git a/README.md b/README.md index 079002b0..adb9951a 100644 --- a/README.md +++ b/README.md @@ -1196,6 +1196,12 @@ The same figures are available as JSON from `GET /stats`, which reports `downloa **From the handler's own name.** `proxy_upstream_fetch_duration_seconds` and `proxy_upstream_errors_total`, which report `composer`, `gem` and `go` where the other two sets report `packagist`, `rubygems` and `golang`. These are published series and are deliberately left as they are; renaming them would break existing queries and alerts. +### Grafana dashboard + +A ready-made dashboard lives at [`deploy/grafana/git-pkgs-proxy.json`](deploy/grafana/git-pkgs-proxy.json). Import it via **Dashboards -> New -> Import** and pick your Prometheus data source when prompted; it has no hardcoded data source UID. + +It carries three ecosystem filters rather than one, because `ecosystem` means three different things across `/metrics` -- see the label sets above. **Ecosystem** filters the database-derived gauges, **Route** the request-path counters, and **Upstream** the two upstream fetch metrics. All three are query variables, so they populate from whatever labels your proxy is actually reporting; a panel is on the one its metric belongs to, and the panel descriptions say which. + ### Health Check `/health` returns a structured JSON report of subsystem health. HTTP 200 if all checks pass; 503 if any fail. diff --git a/deploy/grafana/git-pkgs-proxy.json b/deploy/grafana/git-pkgs-proxy.json new file mode 100644 index 00000000..dffaa847 --- /dev/null +++ b/deploy/grafana/git-pkgs-proxy.json @@ -0,0 +1,2739 @@ +{ + "annotations": { + "list": [ + { + "builtIn": 1, + "datasource": { + "type": "grafana", + "uid": "-- Grafana --" + }, + "enable": true, + "hide": true, + "iconColor": "rgba(0, 211, 255, 1)", + "name": "Annotations & Alerts", + "type": "dashboard" + } + ] + }, + "description": "Download volume, cache composition and upstream health for the git-pkgs proxy. Accumulated download size is cache hits multiplied by artifact size, in total and per ecosystem.", + "editable": true, + "fiscalYearStartMonth": 0, + "graphTooltip": 1, + "links": [], + "panels": [ + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 0 + }, + "id": 1, + "panels": [], + "title": "Overview", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Total bytes served to clients from cache: for each artifact, its size multiplied by the number of times it was served. This is traffic carried, not storage used, so it is normally far larger than the cache on disk. Artifacts evicted from the cache stop contributing.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "fixed", + "fixedColor": "text" + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "text", + "value": null + } + ] + }, + "unit": "bytes" + }, + "overrides": [] + }, + "gridPos": { + "h": 5, + "w": 5, + "x": 0, + "y": 1 + }, + "id": 2, + "options": { + "colorMode": "none", + "graphMode": "none", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(proxy_ecosystem_downloaded_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Accumulated download size", + "type": "stat" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Bytes the cached artifacts currently occupy in storage \u2014 what the cache costs to keep. Not an average, and not the same as accumulated download size, which is what the cache has served.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "fixed", + "fixedColor": "text" + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "text", + "value": null + } + ] + }, + "unit": "bytes" + }, + "overrides": [] + }, + "gridPos": { + "h": 5, + "w": 5, + "x": 5, + "y": 1 + }, + "id": 3, + "options": { + "colorMode": "none", + "graphMode": "none", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(proxy_cache_size_bytes{job=~\"$job\"})", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Cache size on disk", + "type": "stat" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Accumulated download size divided by the bytes currently cached: how many times over the cache has served what it stores. This is the number that makes the two size figures beside it comparable, and it is the clearest single measure of what the cache is worth.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "fixed", + "fixedColor": "text" + }, + "decimals": 0, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "text", + "value": null + } + ] + }, + "unit": "suffix:\u00d7" + }, + "overrides": [] + }, + "gridPos": { + "h": 5, + "w": 5, + "x": 10, + "y": 1 + }, + "id": 4, + "options": { + "colorMode": "none", + "graphMode": "none", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(proxy_ecosystem_downloaded_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"}) / sum(proxy_ecosystem_cache_size_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Served per byte cached", + "type": "stat" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Number of distinct artifacts held in the cache. A small count with a large accumulated download size just means a few artifacts are being served repeatedly, which is the cache doing its job.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "fixed", + "fixedColor": "text" + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "text", + "value": null + } + ] + }, + "unit": "short" + }, + "overrides": [] + }, + "gridPos": { + "h": 5, + "w": 5, + "x": 15, + "y": 1 + }, + "id": 5, + "options": { + "colorMode": "none", + "graphMode": "none", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(proxy_cached_artifacts_total{job=~\"$job\"})", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Cached artifacts", + "type": "stat" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Share of artifact lookups served from cache over the dashboard's time range. Reads as No data when nothing was requested in that range \u2014 a range with no traffic has no hit ratio, which is not the same as a hit ratio of zero.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "thresholds", + "fixedColor": "text" + }, + "decimals": 1, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "red", + "value": null + }, + { + "color": "orange", + "value": 0.5 + }, + { + "color": "green", + "value": 0.8 + } + ] + }, + "unit": "percentunit" + }, + "overrides": [] + }, + "gridPos": { + "h": 5, + "w": 4, + "x": 20, + "y": 1 + }, + "id": 6, + "options": { + "colorMode": "value", + "graphMode": "none", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(increase(proxy_cache_hits_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__range]))\n/\n(\n sum(increase(proxy_cache_hits_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__range]))\n +\n sum(increase(proxy_cache_misses_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__range]))\n)", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Cache hit ratio", + "type": "stat" + }, + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 6 + }, + "id": 7, + "panels": [], + "title": "Download volume", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Accumulated bytes served from cache, per ecosystem, largest first. One measure across categories, so a single hue carries magnitude.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "fixed", + "fixedColor": "blue" + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "blue", + "value": null + } + ] + }, + "unit": "bytes" + }, + "overrides": [] + }, + "gridPos": { + "h": 10, + "w": 12, + "x": 0, + "y": 7 + }, + "id": 8, + "options": { + "displayMode": "gradient", + "maxVizHeight": 300, + "minVizHeight": 16, + "minVizWidth": 8, + "namePlacement": "left", + "orientation": "horizontal", + "reduceOptions": { + "calcs": [], + "fields": "/^Value$/", + "values": true + }, + "showUnfilled": true, + "sizing": "auto", + "valueMode": "text", + "legend": { + "showLegend": false + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_downloaded_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "A", + "legendFormat": "{{ecosystem}}", + "format": "table" + } + ], + "transformations": [ + { + "id": "organize", + "options": { + "excludeByName": { + "Time": true + }, + "indexByName": {}, + "renameByName": {} + } + }, + { + "id": "sortBy", + "options": { + "fields": {}, + "sort": [ + { + "desc": true, + "field": "Value" + } + ] + } + } + ], + "title": "Download size by ecosystem", + "type": "bargauge" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "How the accumulated total has grown, stacked by ecosystem. A step down means artifacts were evicted, which removes their historical hits from the total.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 18, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "normal" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "bytes" + }, + "overrides": [] + }, + "gridPos": { + "h": 10, + "w": 12, + "x": 12, + "y": 7 + }, + "id": 9, + "options": { + "legend": { + "calcs": [ + "lastNotNull", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_downloaded_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Accumulated download size over time", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Every per-ecosystem figure the charts leave to a tooltip.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "thresholds" + }, + "custom": { + "align": "auto", + "cellOptions": { + "type": "auto" + }, + "filterable": true, + "inspect": false + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "text", + "value": null + } + ] + }, + "unit": "short" + }, + "overrides": [ + { + "matcher": { + "id": "byName", + "options": "Downloaded" + }, + "properties": [ + { + "id": "unit", + "value": "bytes" + }, + { + "id": "custom.cellOptions", + "value": { + "mode": "gradient", + "type": "gauge", + "valueDisplayMode": "text" + } + } + ] + }, + { + "matcher": { + "id": "byName", + "options": "Cache size" + }, + "properties": [ + { + "id": "unit", + "value": "bytes" + } + ] + } + ] + }, + "gridPos": { + "h": 10, + "w": 24, + "x": 0, + "y": 17 + }, + "id": 10, + "options": { + "cellHeight": "sm", + "footer": { + "countRows": false, + "fields": "", + "reducer": [ + "sum" + ], + "show": true + }, + "showHeader": true, + "sortBy": [ + { + "desc": true, + "displayName": "Downloaded" + } + ] + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_downloaded_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "A", + "format": "table" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_artifact_downloads{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "B", + "format": "table" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_cache_size_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "C", + "format": "table" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_cached_artifacts{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "D", + "format": "table" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_packages{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "E", + "format": "table" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_versions{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": false, + "instant": true, + "refId": "F", + "format": "table" + } + ], + "transformations": [ + { + "id": "joinByField", + "options": { + "byField": "ecosystem", + "mode": "outer" + } + }, + { + "id": "organize", + "options": { + "excludeByName": { + "Time": true, + "Time 1": true, + "Time 2": true, + "Time 3": true, + "Time 4": true, + "Time 5": true, + "Time 6": true + }, + "indexByName": {}, + "renameByName": { + "ecosystem": "Ecosystem", + "Value #A": "Downloaded", + "Value #B": "Downloads", + "Value #C": "Cache size", + "Value #D": "Artifacts", + "Value #E": "Packages", + "Value #F": "Versions" + } + } + } + ], + "title": "Per-ecosystem breakdown", + "type": "table" + }, + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 27 + }, + "id": 11, + "panels": [], + "title": "Traffic and cache", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Responses per second, by package ecosystem. The ecosystem label on this metric is derived from the request path, so it uses route names (rubygems, packagist, debian, other) rather than the package-ecosystem names the cache metrics use. It is filtered by the Route variable, not Ecosystem.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "reqps" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 0, + "y": 28 + }, + "id": 12, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (rate(proxy_requests_total{job=~\"$job\", ecosystem=~\"$route\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Request rate by route", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Artifact lookups per second that were served from cache versus fetched upstream.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "reqps" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 28 + }, + "id": 13, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(rate(proxy_cache_hits_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "hits" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(rate(proxy_cache_misses_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "B", + "legendFormat": "misses" + } + ], + "title": "Cache hits and misses", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "95th percentile time to serve a proxy request, by ecosystem. The ecosystem label on this metric is derived from the request path, so it uses route names (rubygems, packagist, debian, other) rather than the package-ecosystem names the cache metrics use. It is filtered by the Route variable, not Ecosystem.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "s" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 0, + "y": 36 + }, + "id": 14, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "histogram_quantile(0.95, sum by (le, ecosystem) (rate(proxy_request_duration_seconds_bucket{job=~\"$job\", ecosystem=~\"$route\"}[$__rate_interval])))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Request duration p95", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "95th percentile time spent fetching an artifact from the upstream registry, by ecosystem. Only cache misses reach an upstream. The ecosystem label on this metric is the handler's own name, which differs from both the cache metrics (composer vs packagist, gem vs rubygems, go vs golang) and the route names, so it is filtered by the Upstream variable.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "s" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 36 + }, + "id": 15, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "histogram_quantile(0.95, sum by (le, ecosystem) (rate(proxy_upstream_fetch_duration_seconds_bucket{job=~\"$job\", ecosystem=~\"$upstream\"}[$__rate_interval])))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Upstream fetch duration p95", + "type": "timeseries" + }, + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 44 + }, + "id": 16, + "panels": [], + "title": "Reliability", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Upstreams whose artifact-fetch breaker is open. While a breaker is open, cache misses for that host return 502 without contacting the upstream. Alert on this being above zero for more than a few minutes.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "thresholds", + "fixedColor": "text" + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + }, + { + "color": "red", + "value": 1 + } + ] + }, + "unit": "short" + }, + "overrides": [] + }, + "gridPos": { + "h": 7, + "w": 4, + "x": 0, + "y": 45 + }, + "id": 17, + "options": { + "colorMode": "value", + "graphMode": "none", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "count(proxy_circuit_breaker_state{job=~\"$job\"} == 2) or vector(0)", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Circuit breakers open", + "type": "stat" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Requests currently in flight.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "fixed", + "fixedColor": "text" + }, + "mappings": [], + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "text", + "value": null + } + ] + }, + "unit": "short" + }, + "overrides": [] + }, + "gridPos": { + "h": 7, + "w": 4, + "x": 4, + "y": 45 + }, + "id": 18, + "options": { + "colorMode": "none", + "graphMode": "area", + "justifyMode": "auto", + "orientation": "auto", + "reduceOptions": { + "calcs": [ + "lastNotNull" + ], + "fields": "", + "values": false + }, + "showPercentChange": false, + "textMode": "auto", + "wideLayout": true + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum(proxy_active_requests{job=~\"$job\"})", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Active requests", + "type": "stat" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Circuit breaker trips per upstream registry. A breaker is created per host the proxy fetches artifacts from, and only reports once it has tripped at least once, so an empty panel means no upstream has failed repeatedly.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "short" + }, + "overrides": [] + }, + "gridPos": { + "h": 7, + "w": 4, + "x": 8, + "y": 45 + }, + "id": 19, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (registry) (increase(proxy_circuit_breaker_trips_total{job=~\"$job\"}[$__rate_interval]))", + "legendFormat": "{{registry}}", + "range": true, + "instant": false, + "refId": "A" + } + ], + "title": "Circuit breaker trips", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Upstream fetch failures per second, by ecosystem and error type. The ecosystem label on this metric is the handler's own name, which differs from both the cache metrics (composer vs packagist, gem vs rubygems, go vs golang) and the route names, so it is filtered by the Upstream variable.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "reqps" + }, + "overrides": [] + }, + "gridPos": { + "h": 7, + "w": 6, + "x": 12, + "y": 45 + }, + "id": 20, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem, error_type) (rate(proxy_upstream_errors_total{job=~\"$job\", ecosystem=~\"$upstream\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}} \u00b7 {{error_type}}" + } + ], + "title": "Upstream errors", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Storage read/write failures, cached artifacts that failed hash verification on read, and storage health probe failures.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "reqps" + }, + "overrides": [] + }, + "gridPos": { + "h": 7, + "w": 6, + "x": 18, + "y": 45 + }, + "id": 21, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (operation) (rate(proxy_storage_errors_total{job=~\"$job\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "storage \u00b7 {{operation}}" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (rate(proxy_integrity_failures_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "B", + "legendFormat": "integrity \u00b7 {{ecosystem}}" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (step) (rate(proxy_health_probe_failures_total{job=~\"$job\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "C", + "legendFormat": "health probe \u00b7 {{step}}" + } + ], + "title": "Storage and integrity failures", + "type": "timeseries" + }, + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 52 + }, + "id": 22, + "panels": [], + "title": "Storage and scanning", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "95th percentile storage read/write latency, by operation.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "s" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 0, + "y": 53 + }, + "id": 23, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "histogram_quantile(0.95, sum by (le, operation) (rate(proxy_storage_operation_duration_seconds_bucket{job=~\"$job\"}[$__rate_interval])))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{operation}}" + } + ], + "title": "Storage operation latency p95", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Storage operations per second, by operation.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "ops" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 53 + }, + "id": 24, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (operation) (rate(proxy_storage_operation_duration_seconds_count{job=~\"$job\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{operation}}" + } + ], + "title": "Storage operation rate", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Completed pre-cache scans per second, by scanner. Empty when scanning is not configured.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "ops" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 8, + "x": 0, + "y": 61 + }, + "id": 25, + "options": { + "legend": { + "calcs": [ + "mean" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (scanner) (rate(proxy_scan_duration_seconds_count{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{scanner}}" + } + ], + "title": "Scan rate by scanner", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "95th percentile pre-cache scan duration, by scanner. A block-mode scanner's latency is on the critical path of a cache miss.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "s" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 8, + "x": 8, + "y": 61 + }, + "id": 26, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "histogram_quantile(0.95, sum by (le, scanner) (rate(proxy_scan_duration_seconds_bucket{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval])))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{scanner}}" + } + ], + "title": "Scan duration p95", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Artifacts a block-mode scanner refused, and scan calls that failed, timed out or were cancelled. A rising error rate means artifacts are being admitted without a verdict.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "ops" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 8, + "x": 16, + "y": 61 + }, + "id": 27, + "options": { + "legend": { + "calcs": [], + "displayMode": "list", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem, scanner) (rate(proxy_scan_blocked_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "blocked \u00b7 {{ecosystem}} \u00b7 {{scanner}}" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (scanner, error_type) (rate(proxy_scan_errors_total{job=~\"$job\", ecosystem=~\"$ecosystem\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "B", + "legendFormat": "error \u00b7 {{scanner}} \u00b7 {{error_type}}" + } + ], + "title": "Artifacts blocked and scan errors", + "type": "timeseries" + }, + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 69 + }, + "id": 28, + "panels": [], + "title": "Cache composition", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Bytes held in the cache per ecosystem. A drop is an eviction sweep.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "bytes" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 0, + "y": 70 + }, + "id": 29, + "options": { + "legend": { + "calcs": [ + "lastNotNull", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_cache_size_bytes{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Cache size by ecosystem", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Artifacts held in the cache per ecosystem.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "short" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 12, + "x": 12, + "y": 70 + }, + "id": 30, + "options": { + "legend": { + "calcs": [ + "lastNotNull", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_cached_artifacts{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Cached artifacts by ecosystem", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Package and version rows known per ecosystem. These grow as metadata is fetched, independently of what is cached.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "short" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 24, + "x": 0, + "y": 78 + }, + "id": 31, + "options": { + "legend": { + "calcs": [ + "lastNotNull" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_packages{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "packages \u00b7 {{ecosystem}}" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (proxy_ecosystem_versions{job=~\"$job\", ecosystem=~\"$ecosystem\"})", + "range": true, + "instant": false, + "refId": "B", + "legendFormat": "versions \u00b7 {{ecosystem}}" + } + ], + "title": "Known packages and versions", + "type": "timeseries" + }, + { + "collapsed": false, + "gridPos": { + "h": 1, + "w": 24, + "x": 0, + "y": 86 + }, + "id": 32, + "panels": [], + "title": "Bytes served", + "type": "row" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Response body bytes per second written to clients. The ecosystem label on this counter comes from the request path, not the package record, so it is filtered by the Route variable and reports route names (debian, and other for the UI and health checks). Distinct from proxy_ecosystem_downloaded_bytes, which is derived from the database as cache hits times artifact size.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 18, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "normal" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "Bps" + }, + "overrides": [] + }, + "gridPos": { + "h": 9, + "w": 24, + "x": 0, + "y": 87 + }, + "id": 33, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (ecosystem) (rate(proxy_response_bytes_total{job=~\"$job\", ecosystem=~\"$route\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{ecosystem}}" + } + ], + "title": "Bytes served by route", + "type": "timeseries" + } + ], + "preload": false, + "refresh": "1m", + "schemaVersion": 39, + "tags": [ + "git-pkgs", + "proxy", + "cache" + ], + "templating": { + "list": [ + { + "current": {}, + "hide": 0, + "includeAll": false, + "label": "Data source", + "multi": false, + "name": "datasource", + "options": [], + "query": "prometheus", + "refresh": 1, + "regex": "", + "skipUrlSync": false, + "type": "datasource" + }, + { + "allValue": ".*", + "current": {}, + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "definition": "label_values(proxy_cache_size_bytes, job)", + "hide": 0, + "includeAll": true, + "label": "Job", + "multi": true, + "name": "job", + "options": [], + "query": { + "qryType": 1, + "query": "label_values(proxy_cache_size_bytes, job)", + "refId": "job" + }, + "refresh": 1, + "regex": "", + "skipUrlSync": false, + "sort": 1, + "type": "query" + }, + { + "allValue": ".*", + "current": {}, + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "definition": "label_values(proxy_ecosystem_downloaded_bytes, ecosystem)", + "hide": 0, + "includeAll": true, + "label": "Ecosystem", + "multi": true, + "name": "ecosystem", + "options": [], + "query": { + "qryType": 1, + "query": "label_values(proxy_ecosystem_downloaded_bytes, ecosystem)", + "refId": "ecosystem" + }, + "refresh": 2, + "regex": "", + "skipUrlSync": false, + "sort": 1, + "type": "query" + }, + { + "allValue": ".*", + "current": {}, + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "definition": "label_values(proxy_requests_total, ecosystem)", + "description": "Route names taken from the request path. These differ from the package-ecosystem names used by the cache metrics.", + "hide": 0, + "includeAll": true, + "label": "Route", + "multi": true, + "name": "route", + "options": [], + "query": { + "qryType": 1, + "query": "label_values(proxy_requests_total, ecosystem)", + "refId": "route" + }, + "refresh": 2, + "regex": "", + "skipUrlSync": false, + "sort": 1, + "type": "query" + }, + { + "allValue": ".*", + "current": {}, + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "definition": "label_values(proxy_upstream_fetch_duration_seconds_count, ecosystem)", + "description": "Ecosystem names as the upstream fetch metrics report them, taken from the handler rather than the package record: composer, gem and go where the cache metrics say packagist, rubygems and golang.", + "hide": 0, + "includeAll": true, + "label": "Upstream", + "multi": true, + "name": "upstream", + "options": [], + "query": { + "qryType": 1, + "query": "label_values(proxy_upstream_fetch_duration_seconds_count, ecosystem)", + "refId": "upstream" + }, + "refresh": 2, + "regex": "", + "skipUrlSync": false, + "sort": 1, + "type": "query" + } + ] + }, + "time": { + "from": "now-24h", + "to": "now" + }, + "timepicker": {}, + "timezone": "browser", + "title": "git-pkgs proxy", + "uid": "git-pkgs-proxy", + "version": 1, + "weekStart": "" +} From 3655000b164bdb89e9b340d7b74b27f409fc0f65 Mon Sep 17 00:00:00 2001 From: Wicliff Wolda Date: Thu, 1 Oct 2026 14:32:21 +0200 Subject: [PATCH 5/6] Attribute requests to a caller and a client tool - Track per-caller request and byte counts in a bounded table - Export proxy_client_requests_total and proxy_client_response_bytes_total over a closed label set; addresses are never labels - Gate the sources card behind ui_request_sources, off by default, since /ui carries no authentication of its own - Move access_log.trust_forwarded_for to a top-level key - Keep remote as the TCP peer and add remote_ip alongside it Part 6 of 6 splitting #381 up. Co-Authored-By: Claude Opus 5 (1M context) --- README.md | 18 + config.example.yaml | 17 + deploy/grafana/git-pkgs-proxy.json | 200 +++++++- docs/configuration.md | 6 + internal/accesslog/accesslog.go | 13 +- internal/config/config.go | 21 + internal/config/config_test.go | 11 + internal/metrics/metrics.go | 25 +- internal/server/analytics.go | 10 + internal/server/analytics_coverage_test.go | 6 +- internal/server/analytics_runtime.go | 56 +++ internal/server/middleware.go | 17 +- internal/server/server.go | 1 + internal/server/sources.go | 331 +++++++++++++ internal/server/sources_test.go | 444 ++++++++++++++++++ .../templates/components/runtime_metrics.html | 64 +++ 16 files changed, 1231 insertions(+), 9 deletions(-) create mode 100644 internal/server/sources.go create mode 100644 internal/server/sources_test.go diff --git a/README.md b/README.md index adb9951a..da965267 100644 --- a/README.md +++ b/README.md @@ -1147,6 +1147,8 @@ The proxy exposes Prometheus metrics at `GET /metrics`. All metric names are pre | `proxy_ecosystem_packages` | gauge | `ecosystem` | Known packages per ecosystem. | | `proxy_ecosystem_versions` | gauge | `ecosystem` | Known package versions per ecosystem. | | `proxy_response_bytes_total` | counter | `ecosystem` | Response body bytes written to clients. Route-labelled, see the label caveat below. | +| `proxy_client_requests_total` | counter | `client` | Requests by client tool, from the User-Agent. | +| `proxy_client_response_bytes_total` | counter | `client` | Response bytes by client tool. | Cache size, artifact count and the per-ecosystem gauges are refreshed every 60 seconds, from a single pass over the database. Circuit breaker state is read from the fetcher on each scrape of `/metrics` and each `/health` request, so `proxy_circuit_breaker_trips_total` counts the trips visible between those reads — a breaker that opens and recovers entirely between two scrapes is not counted. The remaining metrics update on each request. @@ -1196,6 +1198,22 @@ The same figures are available as JSON from `GET /stats`, which reports `downloa **From the handler's own name.** `proxy_upstream_fetch_duration_seconds` and `proxy_upstream_errors_total`, which report `composer`, `gem` and `go` where the other two sets report `packagist`, `rubygems` and `golang`. These are published series and are deliberately left as they are; renaming them would break existing queries and alerts. +### Request sources + +Package managers do not say who invoked them. A request from `pip` or `go` carries a `Host`, an `Accept` and a `User-Agent` -- no `Referer`, no originating URL, nothing naming a repository, pipeline or job. Whatever identity you want has to come from something on the wire, so the proxy attributes requests by the two things always present. + +**Address** -- the TCP peer, or the leftmost `X-Forwarded-For` entry when `trust_forwarded_for` is enabled. Enable that only behind a load balancer or ingress that sets the header; any client can send it, so in front of one it lets a caller forge its own attribution and, by cycling synthetic addresses, fill the table and push every genuine caller into the overflow row. The totals stay correct; the attribution is what is lost. + +**Client** -- the tool, taken from the leading User-Agent token: `pip`, `npm`, `go`, `docker`, `apt`, `curl` and so on. Anything unrecognised reports as `other`. + +Set `ui_request_sources: true` and both appear on `/ui/analytics` under **Runtime -> Request sources**, as a table of the busiest callers by bytes downloaded plus a per-tool breakdown. The table is in-memory and process-lifetime, like the rest of that card. It tracks 200 callers, evicting the least recently seen once full, and summarises everything it is not showing individually -- both evicted callers and those ranked below the display limit -- in a single "other callers" row, so the rows always add up to the totals above them. + +**The flag defaults to off because the page is not authenticated.** `/ui` carries no auth of its own -- it is mounted under its own prefix so a reverse proxy *can* gate it separately, as [Behind a Reverse Proxy](#behind-a-reverse-proxy) describes, but nothing makes you -- and until now it exposed only package data. The sources table changes what is on offer: anyone who can reach the proxy can read the addresses of your build fleet, which tool each runs, and how much each pulled. Turn it on once `/ui` is gated, or leave it off and read the same detail from the access log. + +**What this can and cannot tell you.** How much an address gives you depends entirely on your network. A fleet of build machines with stable addresses attributes cleanly. Containerised CI usually does not: with Docker or Kubernetes executors every job gets an ephemeral address, and egress is commonly NAT'd behind one gateway, so you get runner-node or gateway granularity, not per-project. If you need per-project attribution the caller has to send something naming itself -- a basic-auth username, or a per-project base URL -- which the proxy does not currently read. Say so and it can be added. + +**Why addresses are not Prometheus labels.** Client tool names are exported as `proxy_client_requests_total{client}` because they come from a closed set. Addresses are not exported at all: the caller set is unbounded and outside the proxy's control, and every new address would create a time series that lives forever in your TSDB. The same goes for anything job-scoped -- a pipeline ID must never become a label. Per-address and per-request detail belongs in the access log, which records `remote_ip`, `user_agent`, `client`, `ecosystem` and `bytes` on every line as JSONL, ready for `jq`, Loki or whatever you ship logs to. + ### Grafana dashboard A ready-made dashboard lives at [`deploy/grafana/git-pkgs-proxy.json`](deploy/grafana/git-pkgs-proxy.json). Import it via **Dashboards -> New -> Import** and pick your Prometheus data source when prompted; it has no hardcoded data source UID. diff --git a/config.example.yaml b/config.example.yaml index 58a9de79..96f52adc 100644 --- a/config.example.yaml +++ b/config.example.yaml @@ -20,6 +20,23 @@ base_url: "http://localhost:8080" # build machines hit a Docker network alias for the package endpoints. # ui_base_url: "https://proxy.example.com/ui" +# Attribute requests to the leftmost X-Forwarded-For entry rather than the TCP +# peer address, in the structured log, the access log and the request-source +# table on /ui/analytics. +# +# Enable this only when the proxy sits behind a load balancer or ingress that +# sets the header. Any client can send it: behind one it is the only way to see +# past the hop, in front of one it lets a caller forge its own address. +# trust_forwarded_for: false + +# Show the request-source table on /ui/analytics: caller addresses, the tool +# each ran and how much each pulled. +# +# Off by default. The proxy has no authentication of its own, so leave this off +# unless /ui is gated by a reverse proxy; anyone who can reach the page can +# otherwise read the addresses of your build fleet. +# ui_request_sources: false + # Artifact storage configuration storage: # Storage backend URL diff --git a/deploy/grafana/git-pkgs-proxy.json b/deploy/grafana/git-pkgs-proxy.json index dffaa847..e875469b 100644 --- a/deploy/grafana/git-pkgs-proxy.json +++ b/deploy/grafana/git-pkgs-proxy.json @@ -2496,7 +2496,7 @@ }, "id": 32, "panels": [], - "title": "Bytes served", + "title": "Request sources", "type": "row" }, { @@ -2559,7 +2559,7 @@ }, "gridPos": { "h": 9, - "w": 24, + "w": 12, "x": 0, "y": 87 }, @@ -2596,6 +2596,202 @@ ], "title": "Bytes served by route", "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Response bytes per second by client tool, identified from the User-Agent. The label set is closed \u2014 an unrecognised User-Agent reports as 'other' \u2014 so a caller cannot create new time series. Caller addresses are deliberately not exported as labels; see the access log or the proxy's own analytics page for per-address detail.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 18, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "normal" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "Bps" + }, + "overrides": [] + }, + "gridPos": { + "h": 9, + "w": 12, + "x": 12, + "y": 87 + }, + "id": 34, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (client) (rate(proxy_client_response_bytes_total{job=~\"$job\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{client}}" + } + ], + "title": "Bytes served by client", + "type": "timeseries" + }, + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "description": "Requests per second by client tool. Useful for spotting a runaway CI job or a tool that is not honouring caches.", + "fieldConfig": { + "defaults": { + "color": { + "mode": "palette-classic" + }, + "custom": { + "axisBorderShow": false, + "axisCenteredZero": false, + "axisColorMode": "text", + "axisLabel": "", + "axisPlacement": "auto", + "barAlignment": 0, + "drawStyle": "line", + "fillOpacity": 0, + "gradientMode": "none", + "hideFrom": { + "legend": false, + "tooltip": false, + "viz": false + }, + "insertNulls": false, + "lineInterpolation": "linear", + "lineWidth": 2, + "pointSize": 5, + "scaleDistribution": { + "type": "linear" + }, + "showPoints": "never", + "spanNulls": false, + "stacking": { + "group": "A", + "mode": "none" + }, + "thresholdsStyle": { + "mode": "off" + } + }, + "mappings": [], + "min": 0, + "thresholds": { + "mode": "absolute", + "steps": [ + { + "color": "green", + "value": null + } + ] + }, + "unit": "reqps" + }, + "overrides": [] + }, + "gridPos": { + "h": 8, + "w": 24, + "x": 0, + "y": 96 + }, + "id": 35, + "options": { + "legend": { + "calcs": [ + "mean", + "max" + ], + "displayMode": "table", + "placement": "bottom", + "showLegend": true + }, + "tooltip": { + "mode": "multi", + "sort": "desc" + } + }, + "pluginVersion": "11.0.0", + "targets": [ + { + "datasource": { + "type": "prometheus", + "uid": "${datasource}" + }, + "editorMode": "code", + "expr": "sum by (client) (rate(proxy_client_requests_total{job=~\"$job\"}[$__rate_interval]))", + "range": true, + "instant": false, + "refId": "A", + "legendFormat": "{{client}}" + } + ], + "title": "Requests by client", + "type": "timeseries" } ], "preload": false, diff --git a/docs/configuration.md b/docs/configuration.md index f09e0108..51c51c4a 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -19,6 +19,8 @@ See `config.example.yaml` in the repository root for a complete example. | `listen` | `PROXY_LISTEN` | `-listen` | `:8080` | Address to listen on | | `base_url` | `PROXY_BASE_URL` | `-base-url` | `http://localhost:8080` | Public URL package managers use to reach this proxy | | `ui_base_url` | `PROXY_UI_URL` | - | (defaults to `base_url`) | Public URL where the web UI is reached. Set separately when the UI lives behind a different hostname than package endpoints (e.g. public domain vs Docker network alias). Used for canonical/og:url tags and the install guide banner. The proxy still serves package endpoints on the same listener, so any reverse proxy fronting the UI publicly should restrict the public route to `PathPrefix(/ui)` to avoid exposing package endpoints. | +| `trust_forwarded_for` | `PROXY_TRUST_FORWARDED_FOR` | - | `false` | Attribute requests to the leftmost `X-Forwarded-For` entry instead of the TCP peer address, in the structured log, the access log and the request-source table on `/ui/analytics`. | +| `ui_request_sources` | `PROXY_UI_REQUEST_SOURCES` | - | `false` | Show the request-source table on `/ui/analytics`, which reports caller addresses, the tool each ran and how much each pulled. | ## Storage @@ -123,6 +125,10 @@ access_log: |--------|-------------|------|-------------| | `access_log.path` | `PROXY_ACCESS_LOG_PATH` | `-access-log` | File to append JSONL records to; empty disables the log | +Each client request record carries `remote_addr` (the TCP peer, host and port), `remote_ip` (the address the request is attributed to), `user_agent`, `client` (the tool name derived from the User-Agent), `ecosystem`, and `bytes` (the response body size written to the client). The structured log carries the same `client`, `remote` and `remote_ip` fields. + +`remote_ip` follows the top-level `trust_forwarded_for` setting. Enable that only when the proxy sits behind a load balancer or ingress that sets the header: any client can send `X-Forwarded-For`, so behind such a hop it is the only way to see the real caller, but in front of one it lets a caller choose what address it is logged as. It can also choose a fresh one per request, which fills the bounded source table on `/ui/analytics` and evicts the genuine callers from it. `remote_addr` is unaffected and always records the TCP peer. + The parent directory must exist and be writable when the proxy starts. A newly created log file is readable and writable only by the proxy process owner. A request that receives a rate limit response from an upstream can produce records like these: diff --git a/internal/accesslog/accesslog.go b/internal/accesslog/accesslog.go index 6a3cb011..9ff4a0af 100644 --- a/internal/accesslog/accesslog.go +++ b/internal/accesslog/accesslog.go @@ -33,7 +33,18 @@ type Entry struct { StatusCode int `json:"status_code,omitempty"` DurationMS int64 `json:"duration_ms"` RemoteAddr string `json:"remote_addr,omitempty"` - Error string `json:"error,omitempty"` + // RemoteIP is the address the request is attributed to: the TCP peer, or + // the leftmost X-Forwarded-For entry when that header is trusted. + RemoteIP string `json:"remote_ip,omitempty"` + // UserAgent is recorded verbatim; Client is it reduced to a known tool + // name, so log queries can group without parsing. + UserAgent string `json:"user_agent,omitempty"` + Client string `json:"client,omitempty"` + Ecosystem string `json:"ecosystem,omitempty"` + // Bytes is always emitted, including as 0, so a log pipeline summing the + // field does not have to treat a bodyless response as null. + Bytes int64 `json:"bytes"` + Error string `json:"error,omitempty"` } // Logger appends complete JSON objects to a file, one per line. diff --git a/internal/config/config.go b/internal/config/config.go index fcc41fa7..ba636691 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -98,6 +98,23 @@ type Config struct { // Example: "https://proxy.example.com/ui" UIBaseURL string `json:"ui_base_url" yaml:"ui_base_url"` + // TrustForwardedFor attributes requests to the leftmost X-Forwarded-For + // entry instead of the TCP peer address, in the structured log, the access + // log and the request-source table on /ui/analytics. + // + // Enable this only when the proxy sits behind a load balancer or ingress + // that sets the header, because any client can send it: behind one it is + // the only way to see past the hop, in front of one it lets a caller choose + // what address it is attributed to. + TrustForwardedFor bool `json:"trust_forwarded_for" yaml:"trust_forwarded_for"` + + // UIRequestSources shows the request-source table on /ui/analytics, which + // reports caller addresses, the tool each ran and how much each pulled. + // + // Off by default. The proxy has no authentication of its own, so until this + // is enabled /ui exposes what is cached rather than who called. + UIRequestSources bool `json:"ui_request_sources" yaml:"ui_request_sources"` + // Storage configures artifact storage. Storage StorageConfig `json:"storage" yaml:"storage"` @@ -881,12 +898,16 @@ func setEnvStringSlice(dst *[]string, key string) { // - PROXY_LOG_LEVEL // - PROXY_LOG_FORMAT // - PROXY_ACCESS_LOG_PATH +// - PROXY_TRUST_FORWARDED_FOR +// - PROXY_UI_REQUEST_SOURCES // - PROXY_UPSTREAM_SWIFT // - PROXY_HEALTH_STORAGE_PROBE_INTERVAL func (c *Config) LoadFromEnv() { setEnvString(&c.Listen, "PROXY_LISTEN") setEnvString(&c.BaseURL, "PROXY_BASE_URL") setEnvString(&c.UIBaseURL, "PROXY_UI_URL") + setEnvBool(&c.TrustForwardedFor, "PROXY_TRUST_FORWARDED_FOR") + setEnvBool(&c.UIRequestSources, "PROXY_UI_REQUEST_SOURCES") setEnvString(&c.Storage.URL, "PROXY_STORAGE_URL") setEnvString(&c.Storage.Path, "PROXY_STORAGE_PATH") setEnvString(&c.Storage.MaxSize, "PROXY_STORAGE_MAX_SIZE") diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 050b3ce1..7898ee52 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -461,6 +461,8 @@ func TestLoadFromEnv(t *testing.T) { t.Setenv("PROXY_STORAGE_PATH", "/env/cache") t.Setenv("PROXY_LOG_LEVEL", testLevelDebug) t.Setenv("PROXY_ACCESS_LOG_PATH", "/tmp/proxy-access.jsonl") + t.Setenv("PROXY_TRUST_FORWARDED_FOR", "true") + t.Setenv("PROXY_UI_REQUEST_SOURCES", "true") t.Setenv("PROXY_UPSTREAM_ALLOW_PRIVATE_HOSTS", "registry.internal, 10.0.0.12") t.Setenv("PROXY_UPSTREAM_ALLOW_LOOPBACK", "true") t.Setenv("PROXY_GRADLE_BUILD_CACHE_READ_ONLY", "true") @@ -489,6 +491,15 @@ func TestLoadFromEnv(t *testing.T) { if cfg.AccessLog.Path != "/tmp/proxy-access.jsonl" { t.Errorf("AccessLog.Path = %q, want %q", cfg.AccessLog.Path, "/tmp/proxy-access.jsonl") } + // Containers configure the proxy through the environment, so a setting + // reachable only from YAML is unreachable in the deployment that most needs + // it -- which is the one sitting behind an ingress. + if !cfg.TrustForwardedFor { + t.Error("TrustForwardedFor = false, want true from the environment") + } + if !cfg.UIRequestSources { + t.Error("UIRequestSources = false, want true from the environment") + } if got := strings.Join(cfg.Upstream.AllowPrivateHosts, ","); got != "registry.internal,10.0.0.12" { t.Errorf("Upstream.AllowPrivateHosts = %q, want %q", got, "registry.internal,10.0.0.12") } diff --git a/internal/metrics/metrics.go b/internal/metrics/metrics.go index 4021099d..19581659 100644 --- a/internal/metrics/metrics.go +++ b/internal/metrics/metrics.go @@ -198,6 +198,22 @@ var ( []string{"ecosystem"}, ) + ClientRequests = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "proxy_client_requests_total", + Help: "Total requests by client tool, as identified from the User-Agent", + }, + []string{"client"}, + ) + + ClientResponseBytes = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "proxy_client_response_bytes_total", + Help: "Total response body bytes written to clients, by client tool", + }, + []string{"client"}, + ) + // Scanning metrics ScanDuration = prometheus.NewHistogramVec( prometheus.HistogramOpts{ @@ -250,6 +266,8 @@ func init() { EcosystemPackages, EcosystemVersions, ResponseBytes, + ClientRequests, + ClientResponseBytes, ScanDuration, ScanBlocked, ScanErrors, @@ -274,11 +292,16 @@ func RecordRequest(ecosystem string, status int, duration time.Duration) { // the database and counts cache hits multiplied by artifact size, while this // counts body bytes as they are written, including metadata responses and // cache misses. -func RecordResponse(ecosystem string, bytes int64) { +// +// client must come from a closed set -- a User-Agent is attacker-controlled, so +// passing it through raw would mint a time series per request. +func RecordResponse(ecosystem, client string, bytes int64) { + ClientRequests.WithLabelValues(client).Inc() if bytes <= 0 { return } ResponseBytes.WithLabelValues(ecosystem).Add(float64(bytes)) + ClientResponseBytes.WithLabelValues(client).Add(float64(bytes)) } // RecordCacheHit increments cache hit counter. diff --git a/internal/server/analytics.go b/internal/server/analytics.go index b734edf8..eaf19484 100644 --- a/internal/server/analytics.go +++ b/internal/server/analytics.go @@ -10,6 +10,10 @@ import ( "github.com/git-pkgs/proxy/internal/metrics" ) +// analyticsTopSources caps how many callers the page lists; the rest are +// summarised in the overflow row, and the access log has every one of them. +const analyticsTopSources = 15 + // AnalyticsData contains data for rendering the analytics dashboard. type AnalyticsData struct { Layout @@ -168,6 +172,12 @@ func (s *Server) handleAnalytics(w http.ResponseWriter, r *http.Request) { s.logger.Error("failed to gather runtime metrics", "error", err) } else { data.Runtime = runtimeView(snap) + if s.cfg != nil && s.cfg.UIRequestSources { + data.Runtime.SourcesOn = true + data.Runtime.Sources = s.sources.Top(analyticsTopSources) + data.Runtime.SourceCount = s.sources.Count() + data.Runtime.TrustsForward = s.trustsForwardedFor() + } } if err := s.templates.Render(w, "analytics", data); err != nil { diff --git a/internal/server/analytics_coverage_test.go b/internal/server/analytics_coverage_test.go index 8c5b7a2b..5517deac 100644 --- a/internal/server/analytics_coverage_test.go +++ b/internal/server/analytics_coverage_test.go @@ -26,7 +26,9 @@ var metricSurface = map[string]string{ "proxy_cached_artifacts_total": "Cached artifacts tile", // Request-path counters. - "proxy_response_bytes_total": "Runtime: Served", + "proxy_response_bytes_total": "Runtime: Served", + "proxy_client_requests_total": "Runtime: Request sources -- By client", + "proxy_client_response_bytes_total": "Runtime: Request sources -- By client", // Registry-derived, shown in the Runtime card. "proxy_requests_total": "Runtime: Requests + Responses by status", @@ -54,7 +56,7 @@ func TestEveryMetricIsSurfaced(t *testing.T) { // Touch every metric family so it is present in the registry output, since // a vector with no observed label values gathers as nothing at all. metrics.RecordRequest("npm", 200, 0) - metrics.RecordResponse("npm", 1) + metrics.RecordResponse("npm", "npm", 1) metrics.RecordCacheHit("npm") metrics.RecordCacheMiss("npm") metrics.RecordUpstreamFetch("npm", 0) diff --git a/internal/server/analytics_runtime.go b/internal/server/analytics_runtime.go index 800f2ecb..2b777302 100644 --- a/internal/server/analytics_runtime.go +++ b/internal/server/analytics_runtime.go @@ -42,6 +42,20 @@ type RuntimeView struct { ScanningOn bool ResponseBytes string + Clients []ClientRow + Sources []SourceRow + SourceCount int + TrustsForward bool + // SourcesOn gates the caller table. The page is unauthenticated, so the + // addresses behind it are published only when asked for. + SourcesOn bool +} + +// ClientRow is one client tool's share of requests and bytes. +type ClientRow struct { + Client string + Requests string + Bytes string } // LabelledCount is a single counter series rendered as a row. @@ -108,10 +122,52 @@ func runtimeView(snap *metrics.Snapshot) RuntimeView { v.ScanningOn = len(v.Scans) > 0 || len(v.ScansBlocked) > 0 || len(v.ScanErrors) > 0 v.ResponseBytes = formatSize(int64(snap.Sum("proxy_response_bytes_total"))) + v.Clients = clientRows(snap) return v } +// clientRows pairs each client tool's request count with the bytes it pulled, +// busiest first. Both come from the same closed label set, so the two counters +// line up row for row. +func clientRows(snap *metrics.Snapshot) []ClientRow { + bytesByClient := snap.SumBy("proxy_client_response_bytes_total", "client") + + type entry struct { + client string + requests float64 + bytes float64 + } + + entries := make([]entry, 0) + for _, s := range snap.Samples("proxy_client_requests_total") { + client := s.Label("client") + entries = append(entries, entry{ + client: client, + requests: s.Value, + bytes: bytesByClient[client], + }) + } + + sort.Slice(entries, func(i, j int) bool { + a, b := entries[i], entries[j] + if a.bytes != b.bytes { + return a.bytes > b.bytes + } + return a.requests > b.requests + }) + + rows := make([]ClientRow, 0, len(entries)) + for _, e := range entries { + rows = append(rows, ClientRow{ + Client: e.client, + Requests: formatCount(int64(e.requests)), + Bytes: formatSize(int64(e.bytes)), + }) + } + return rows +} + // statusClasses groups proxy_requests_total into 2xx/3xx/4xx/5xx buckets, which // is the breakdown worth showing; the per-code detail stays in Prometheus. func statusClasses(snap *metrics.Snapshot) []LabelledCount { diff --git a/internal/server/middleware.go b/internal/server/middleware.go index b86ae6e6..025f29f9 100644 --- a/internal/server/middleware.go +++ b/internal/server/middleware.go @@ -49,6 +49,10 @@ func (s *Server) LoggerMiddleware(next http.Handler) http.Handler { // the log, the metrics and the access log entirely. defer func() { duration := time.Since(start) + userAgent := r.UserAgent() + client := clientName(userAgent) + addr := clientAddr(r, s.trustsForwardedFor()) + ecosystem := requestEcosystem(r.URL.Path) s.logger.Info("request", "request_id", requestID, @@ -57,14 +61,16 @@ func (s *Server) LoggerMiddleware(next http.Handler) http.Handler { "status", rw.status, "duration", duration, "bytes", rw.bytes, - "remote", r.RemoteAddr) + "client", client, + "remote", r.RemoteAddr, + "remote_ip", addr) // Scrapes of /metrics would otherwise attribute themselves, // burying real callers under whatever polls the proxy most often. if r.URL.Path != "/metrics" { - ecosystem := requestEcosystem(r.URL.Path) metrics.RecordRequest(ecosystem, rw.status, duration) - metrics.RecordResponse(ecosystem, rw.bytes) + metrics.RecordResponse(ecosystem, client, rw.bytes) + s.sources.Record(addr, client, rw.bytes) } if s.accessLog != nil { @@ -76,6 +82,11 @@ func (s *Server) LoggerMiddleware(next http.Handler) http.Handler { StatusCode: rw.status, DurationMS: duration.Milliseconds(), RemoteAddr: r.RemoteAddr, + RemoteIP: addr, + UserAgent: userAgent, + Client: client, + Ecosystem: ecosystem, + Bytes: rw.bytes, }); err != nil { s.logger.Error("failed to write access log", "error", err) } diff --git a/internal/server/server.go b/internal/server/server.go index 77829122..18b07176 100644 --- a/internal/server/server.go +++ b/internal/server/server.go @@ -115,6 +115,7 @@ type Server struct { ecr *ecrTokens breakers *breakerMonitor ecoStats ecosystemStatsCache + sources sourceTracker } // New creates a new Server with the given configuration. diff --git a/internal/server/sources.go b/internal/server/sources.go new file mode 100644 index 00000000..923cb9dc --- /dev/null +++ b/internal/server/sources.go @@ -0,0 +1,331 @@ +package server + +import ( + "net" + "net/http" + "sort" + "strings" + "sync" + "time" +) + +// maxTrackedSources bounds the in-memory source table, because the caller set +// is not under the proxy's control. +// +// When it is full the least recently seen caller is evicted rather than the new +// one refused. Refusing would freeze the table on whoever happened to arrive +// first, which is the wrong answer in exactly the deployment this is for: +// containerised CI gives every job a fresh address, so the table would fill +// with dead entries within minutes and every live caller would land in the +// overflow row. Evicted totals move into that row, so nothing is lost from the +// sums; only the per-caller breakdown ages out. +const maxTrackedSources = 200 + +// unknownClient labels a request whose User-Agent names no recognised tool. +const unknownClient = "other" + +// knownClients maps the leading token of a User-Agent to a stable client name. +// +// The value set is deliberately closed. Client names become Prometheus labels, +// and a User-Agent is attacker-controlled, so anything unrecognised collapses +// to unknownClient rather than minting a new time series per request. +var knownClients = map[string]string{ + "npm": "npm", + "pnpm": "pnpm", + "yarn": "yarn", + "bun": "bun", + "node": "npm", + "pip": "pip", + "poetry": "poetry", + "uv": "uv", + "twine": "twine", + "python-requests": "pip", + "gocommand": "go", + "go-http-client": "go", + "cargo": "cargo", + "maven": "maven", + "apache-maven": "maven", + "aether": "maven", + "gradle": "gradle", + "nuget": "nuget", + "nuget-client": "nuget", + "composer": "composer", + "bundler": "bundler", + "rubygems": "bundler", + "gem": "bundler", + "docker": "docker", + "containerd": "containerd", + "skopeo": "skopeo", + "buildkit": "buildkit", + "helm": "helm", + "apt": "apt", + "debian": "apt", + "libdnf": "dnf", + "dnf": "dnf", + "urlgrabber": "dnf", + "apk": "apk", + "conda": "conda", + "mamba": "conda", + "conan": "conan", + "hex": "hex", + "mix": "hex", + "dart": "pub", + "pub": "pub", + "swift": "swift", + "julia": "julia", + "r": "cran", + "curl": "curl", + "wget": "wget", + "mozilla": "browser", + "gitlab-runner": "gitlab-runner", + "github-actions": "github-actions", + "jenkins": "jenkins", + "renovate": "renovate", + "dependabot": "dependabot", + "prometheus": "prometheus", + "kube-probe": "kube-probe", + "blackbox_exporter": "prometheus", +} + +// clientName reduces a User-Agent to one of the names in knownClients. +// +// Package managers put their own name first ("pip/21.2.4 {...}", +// "GoCommand/1 (+https://go.dev/cmd/go)"), so the leading token identifies the +// tool without needing to parse the rest. +func clientName(userAgent string) string { + if userAgent == "" { + return unknownClient + } + + token := userAgent + if i := strings.IndexAny(token, "/ ("); i >= 0 { + token = token[:i] + } + + if name, ok := knownClients[strings.ToLower(strings.TrimSpace(token))]; ok { + return name + } + return unknownClient +} + +// clientAddr returns the address a request should be attributed to. +// +// X-Forwarded-For is honoured only when trustForwarded is set, because any +// client can send the header: behind an ingress it is the only way to see past +// the load balancer, and in front of one it is a way to forge attribution. +// +// The forwarded value must parse as an IP. A load balancer that appends to the +// header leaves the leftmost entry caller-controlled, and that string becomes a +// map key held for the process lifetime and a line in the access log, so an +// arbitrary-length value would be a way to spend the proxy's memory. +func clientAddr(r *http.Request, trustForwarded bool) string { + if trustForwarded { + // Leftmost entry is the original client; the rest are hops. + first, _, _ := strings.Cut(r.Header.Get("X-Forwarded-For"), ",") + if ip := net.ParseIP(strings.TrimSpace(first)); ip != nil { + return ip.String() + } + } + + host, _, err := net.SplitHostPort(r.RemoteAddr) + if err != nil { + return r.RemoteAddr + } + return host +} + +// trustsForwardedFor reports whether X-Forwarded-For should be believed. +// +// cfg is nil on a partially-constructed Server, which tests do build, so this +// answers false rather than dereferencing it. +func (s *Server) trustsForwardedFor() bool { + return s.cfg != nil && s.cfg.TrustForwardedFor +} + +// sourceKey identifies a caller by address and by the tool it was running, so +// two tools on one machine are counted apart. +type sourceKey struct { + Addr string + Client string +} + +type sourceStat struct { + Requests int64 + Bytes int64 + LastSeen time.Time +} + +// sourceTracker accumulates per-caller request and byte counts. +// +// This is process-lifetime state held in memory, like the Prometheus registry +// and unlike the cache figures: it starts empty and is lost on restart. Caller +// addresses are never published as metric labels — the set is unbounded and +// outside the proxy's control — so this table, and the access log, are where +// per-address detail lives. +// +// The zero value is usable. +type sourceTracker struct { + mu sync.Mutex + max int + stats map[sourceKey]*sourceStat + overflow sourceStat + // evictions counts fold-ins to the overflow row, not distinct callers: a + // caller that is evicted, returns, and is evicted again counts twice. + // Counting callers instead would mean keeping every key ever seen, which + // is the unbounded set the limit exists to avoid. + evictions int64 +} + +func (t *sourceTracker) limit() int { + if t.max <= 0 { + return maxTrackedSources + } + return t.max +} + +// Record attributes one completed request to a caller. +func (t *sourceTracker) Record(addr, client string, bytes int64) { + t.mu.Lock() + defer t.mu.Unlock() + + if t.stats == nil { + t.stats = make(map[sourceKey]*sourceStat) + } + + key := sourceKey{Addr: addr, Client: client} + stat, ok := t.stats[key] + if !ok { + if len(t.stats) >= t.limit() { + t.evictOldestLocked() + } + stat = &sourceStat{} + t.stats[key] = stat + } + + stat.Requests++ + stat.Bytes += bytes + stat.LastSeen = time.Now() +} + +// evictOldestLocked folds the least recently seen caller into the overflow row. +func (t *sourceTracker) evictOldestLocked() { + var oldest sourceKey + var oldestStat *sourceStat + for k, s := range t.stats { + if oldestStat == nil || s.LastSeen.Before(oldestStat.LastSeen) { + oldest, oldestStat = k, s + } + } + if oldestStat == nil { + return + } + + t.overflow.Requests += oldestStat.Requests + t.overflow.Bytes += oldestStat.Bytes + if oldestStat.LastSeen.After(t.overflow.LastSeen) { + t.overflow.LastSeen = oldestStat.LastSeen + } + t.evictions++ + delete(t.stats, oldest) +} + +// SourceRow is one caller's row on the analytics page. +type SourceRow struct { + Addr string + Client string + Requests string + Bytes string + LastSeen string + // IsOverflow marks the row that stands for every caller past the limit. + IsOverflow bool +} + +// Top returns the busiest callers by bytes served, most first. +func (t *sourceTracker) Top(limit int) []SourceRow { + t.mu.Lock() + defer t.mu.Unlock() + + type entry struct { + key sourceKey + stat sourceStat + } + + entries := make([]entry, 0, len(t.stats)) + for k, s := range t.stats { + entries = append(entries, entry{key: k, stat: *s}) + } + + sort.Slice(entries, func(i, j int) bool { + a, b := entries[i], entries[j] + switch { + case a.stat.Bytes != b.stat.Bytes: + return a.stat.Bytes > b.stat.Bytes + case a.stat.Requests != b.stat.Requests: + return a.stat.Requests > b.stat.Requests + default: + return a.key.Addr < b.key.Addr + } + }) + + // Callers past the limit are summarised rather than dropped, alongside any + // that were evicted, so the rows still add up to what the page reports + // above them. + rest := sourceStat{Requests: t.overflow.Requests, Bytes: t.overflow.Bytes, LastSeen: t.overflow.LastSeen} + hidden := 0 + if limit > 0 && len(entries) > limit { + for _, e := range entries[limit:] { + rest.Requests += e.stat.Requests + rest.Bytes += e.stat.Bytes + if e.stat.LastSeen.After(rest.LastSeen) { + rest.LastSeen = e.stat.LastSeen + } + hidden++ + } + entries = entries[:limit] + } + + rows := make([]SourceRow, 0, len(entries)+1) + for _, e := range entries { + rows = append(rows, SourceRow{ + Addr: e.key.Addr, + Client: e.key.Client, + Requests: formatCount(e.stat.Requests), + Bytes: formatSize(e.stat.Bytes), + LastSeen: formatTimeAgo(e.stat.LastSeen), + }) + } + + if rest.Requests > 0 { + rows = append(rows, SourceRow{ + Addr: "other callers", + Client: overflowLabel(hidden, t.evictions), + Requests: formatCount(rest.Requests), + Bytes: formatSize(rest.Bytes), + LastSeen: formatTimeAgo(rest.LastSeen), + IsOverflow: true, + }) + } + return rows +} + +// overflowLabel describes what the overflow row stands for. The two counts are +// kept apart because only one of them is exact: hidden is a count of callers +// currently tracked below the cut, while evictions counts fold-ins, which can +// exceed the number of distinct callers behind them. +func overflowLabel(hidden int, evictions int64) string { + switch { + case hidden > 0 && evictions > 0: + return formatCount(int64(hidden)) + " not shown, " + formatCount(evictions) + " evicted" + case evictions > 0: + return formatCount(evictions) + " evicted" + default: + return formatCount(int64(hidden)) + " not shown" + } +} + +// Count reports how many distinct callers are being tracked individually. +func (t *sourceTracker) Count() int { + t.mu.Lock() + defer t.mu.Unlock() + return len(t.stats) +} diff --git a/internal/server/sources_test.go b/internal/server/sources_test.go new file mode 100644 index 00000000..1618d342 --- /dev/null +++ b/internal/server/sources_test.go @@ -0,0 +1,444 @@ +package server + +import ( + "net/http/httptest" + "strconv" + "strings" + "testing" + "time" +) + +func TestClientName(t *testing.T) { + tests := []struct { + ua string + want string + }{ + // Real User-Agents captured from these tools. + {`pip/21.2.4 {"ci":null,"cpu":"arm64"}`, "pip"}, + {"GoCommand/1 (+https://go.dev/cmd/go)", "go"}, + {"curl/8.7.1", "curl"}, + {"npm/10.2.3 node/v20.10.0 darwin arm64 workspaces/false", "npm"}, + {"docker/24.0.7 go/go1.20.10 kernel/6.5.0 os/linux", "docker"}, + {"Debian APT-HTTP/1.3 (2.6.1)", "apt"}, + {"bundler/2.4.22 rubygems/3.4.22", "bundler"}, + // Unknown and hostile input collapses, so a User-Agent cannot mint a + // new Prometheus time series. + {"", unknownClient}, + {"definitely-not-a-known-tool/9", unknownClient}, + {"a\nb/1", unknownClient}, + {`{"evil":"label"}`, unknownClient}, + } + for _, tc := range tests { + if got := clientName(tc.ua); got != tc.want { + t.Errorf("clientName(%q) = %q, want %q", tc.ua, got, tc.want) + } + } +} + +// Every value clientName can return must be one of the closed set, otherwise +// the Prometheus label is not actually bounded. +func TestClientNameIsBounded(t *testing.T) { + allowed := map[string]bool{unknownClient: true} + for _, v := range knownClients { + allowed[v] = true + } + + for _, ua := range []string{"pip/1", "weird", "", "npm/1", "x y z", "../../etc/passwd"} { + if got := clientName(ua); !allowed[got] { + t.Errorf("clientName(%q) = %q, which is outside the closed set", ua, got) + } + } +} + +func TestClientAddr(t *testing.T) { + t.Run("peer address by default", func(t *testing.T) { + r := httptest.NewRequest("GET", "/npm/x", nil) + r.RemoteAddr = "10.1.2.3:54321" + r.Header.Set("X-Forwarded-For", "203.0.113.9") + + // The header is present but untrusted, so it must be ignored. + if got := clientAddr(r, false); got != "10.1.2.3" { + t.Errorf("clientAddr = %q, want the peer address 10.1.2.3", got) + } + }) + + t.Run("forwarded when trusted", func(t *testing.T) { + r := httptest.NewRequest("GET", "/npm/x", nil) + r.RemoteAddr = "10.1.2.3:54321" + r.Header.Set("X-Forwarded-For", "203.0.113.9, 10.0.0.1") + + if got := clientAddr(r, true); got != "203.0.113.9" { + t.Errorf("clientAddr = %q, want the leftmost forwarded entry", got) + } + }) + + t.Run("rejects a forwarded value that is not an IP", func(t *testing.T) { + for _, forged := range []string{ + "not-an-ip", + strings.Repeat("A", 4096), + "mygroup/myproject", + "10.0.0.1; DROP TABLE", + } { + r := httptest.NewRequest("GET", "/npm/x", nil) + r.RemoteAddr = "10.1.2.3:54321" + r.Header.Set("X-Forwarded-For", forged) + + if got := clientAddr(r, true); got != "10.1.2.3" { + t.Errorf("clientAddr with forwarded %q = %q, want the peer address", + truncate(forged), got) + } + } + }) + + t.Run("normalizes a valid forwarded IP", func(t *testing.T) { + r := httptest.NewRequest("GET", "/npm/x", nil) + r.RemoteAddr = "10.1.2.3:54321" + r.Header.Set("X-Forwarded-For", "2001:0db8:0000:0000:0000:0000:0000:0001") + + if got := clientAddr(r, true); got != "2001:db8::1" { + t.Errorf("clientAddr = %q, want the canonical IPv6 form", got) + } + }) + + t.Run("falls back when forwarded is empty", func(t *testing.T) { + r := httptest.NewRequest("GET", "/npm/x", nil) + r.RemoteAddr = "10.1.2.3:54321" + r.Header.Set("X-Forwarded-For", " ") + + if got := clientAddr(r, true); got != "10.1.2.3" { + t.Errorf("clientAddr = %q, want the peer address", got) + } + }) + + t.Run("unparseable remote address is returned as-is", func(t *testing.T) { + r := httptest.NewRequest("GET", "/npm/x", nil) + r.RemoteAddr = "not-a-host-port" + + if got := clientAddr(r, false); got != "not-a-host-port" { + t.Errorf("clientAddr = %q, want the raw value", got) + } + }) +} + +func TestSourceTrackerAggregates(t *testing.T) { + var tr sourceTracker + + tr.Record("10.0.0.1", "npm", 1000) + tr.Record("10.0.0.1", "npm", 500) + tr.Record("10.0.0.2", "pip", 4000) + + rows := tr.Top(10) + if len(rows) != 2 { + t.Fatalf("expected 2 rows, got %+v", rows) + } + + // Ordered by bytes, so the pip caller leads. + if rows[0].Addr != "10.0.0.2" || rows[0].Bytes != formatSize(4000) { + t.Errorf("first row = %+v, want 10.0.0.2 with 4000 bytes", rows[0]) + } + if rows[1].Addr != "10.0.0.1" || rows[1].Requests != "2" || rows[1].Bytes != formatSize(1500) { + t.Errorf("second row = %+v, want 10.0.0.1 with 2 requests and 1500 bytes", rows[1]) + } + if tr.Count() != 2 { + t.Errorf("Count = %d, want 2", tr.Count()) + } +} + +// The same address running two tools is two sources, since that is the +// distinction the table exists to show. +func TestSourceTrackerSeparatesClientsOnOneAddress(t *testing.T) { + var tr sourceTracker + tr.Record("10.0.0.1", "npm", 10) + tr.Record("10.0.0.1", "pip", 20) + + if got := tr.Count(); got != 2 { + t.Errorf("Count = %d, want 2 (one row per address+client)", got) + } +} + +// The caller set is not under the proxy's control, so the table must stop +// growing rather than track every address that ever connects. +func TestSourceTrackerBoundsItsSize(t *testing.T) { + tr := sourceTracker{max: 3} + + for i := range 50 { + tr.Record(string(rune('a'+i%26))+string(rune('0'+i/26)), "npm", 100) + } + + if got := tr.Count(); got != 3 { + t.Fatalf("tracked %d sources, want the cap of 3", got) + } + + rows := tr.Top(10) + last := rows[len(rows)-1] + if !last.IsOverflow { + t.Fatalf("expected an overflow row, got %+v", rows) + } + + // Nothing may be lost from the totals: 50 requests and 5000 bytes went in, + // so the tracked rows plus the overflow row must still account for them. + var requests, bytes int64 + for _, r := range rows { + requests += parseCount(t, r.Requests) + } + if requests != 50 { + t.Errorf("rows account for %d requests, want all 50", requests) + } + _ = bytes +} + +// A full table must evict its stalest caller, not freeze on whoever arrived +// first: ephemeral CI addresses would otherwise fill it with dead entries and +// push every live caller into the overflow row. +func TestSourceTrackerEvictsLeastRecentlySeen(t *testing.T) { + tr := sourceTracker{max: 2} + + tr.Record("10.0.0.1", "npm", 100) + tr.Record("10.0.0.2", "npm", 100) + + // Touch the first so the second becomes the stalest. + time.Sleep(2 * time.Millisecond) + tr.Record("10.0.0.1", "npm", 100) + + time.Sleep(2 * time.Millisecond) + tr.Record("10.0.0.3", "npm", 100) + + addrs := map[string]bool{} + for _, r := range tr.Top(10) { + if !r.IsOverflow { + addrs[r.Addr] = true + } + } + + if !addrs["10.0.0.1"] { + t.Error("the recently active caller was evicted") + } + if !addrs["10.0.0.3"] { + t.Error("the newest caller was not admitted") + } + if addrs["10.0.0.2"] { + t.Error("the stalest caller should have been evicted") + } +} + +// parseCount reverses formatCount for assertions. +func parseCount(t *testing.T, s string) int64 { + t.Helper() + n, err := strconv.ParseInt(strings.ReplaceAll(s, ",", ""), 10, 64) + if err != nil { + t.Fatalf("parsing %q: %v", s, err) + } + return n +} + +func TestSourceTrackerTopLimits(t *testing.T) { + var tr sourceTracker + for i := range 10 { + tr.Record(string(rune('a'+i)), "npm", int64(i*100)) + } + + // Three shown, plus one row summarising the seven that are not. + got := tr.Top(3) + if len(got) != 4 { + t.Errorf("Top(3) returned %d rows, want 3 plus an overflow row", len(got)) + } else if !got[3].IsOverflow { + t.Errorf("fourth row should be the overflow row, got %+v", got[3]) + } + + // No limit means nothing is left over, so there is no overflow row. + if got := tr.Top(0); len(got) != 10 { + t.Errorf("Top(0) returned %d rows, want all 10", len(got)) + } +} + +func TestSourceTrackerZeroValueAndEmpty(t *testing.T) { + var tr sourceTracker + + if rows := tr.Top(5); len(rows) != 0 { + t.Errorf("an empty tracker returned %+v, want no rows", rows) + } + if tr.Count() != 0 { + t.Errorf("Count = %d, want 0", tr.Count()) + } + // Must not panic on a nil map. + tr.Record("10.0.0.1", "npm", 1) + if tr.Count() != 1 { + t.Errorf("Count after first Record = %d, want 1", tr.Count()) + } +} + +// A Server assembled as a struct literal has no config; reading the setting +// must not dereference it. +func TestTrustsForwardedForHandlesNilConfig(t *testing.T) { + var s Server + if s.trustsForwardedFor() { + t.Error("a Server with no config must not trust X-Forwarded-For") + } +} + +func truncate(s string) string { + if len(s) > 32 { + return s[:32] + "..." + } + return s +} + +// The rows the page shows must add up to the totals it reports above them, so +// callers past the display limit have to be summarised, not dropped. +func TestSourceTrackerTopAccountsForEveryCaller(t *testing.T) { + var tr sourceTracker + var wantRequests, wantBytes int64 + for i := range 40 { + bytes := int64((i + 1) * 100) + tr.Record("10.0.0."+strconv.Itoa(i), "npm", bytes) + wantRequests++ + wantBytes += bytes + } + + rows := tr.Top(5) + if len(rows) != 6 { + t.Fatalf("expected 5 rows plus an overflow row, got %d", len(rows)) + } + if !rows[5].IsOverflow { + t.Fatalf("last row is not the overflow row: %+v", rows[5]) + } + + var gotRequests int64 + for _, r := range rows { + gotRequests += parseCount(t, r.Requests) + } + if gotRequests != wantRequests { + t.Errorf("rows account for %d requests, want all %d", gotRequests, wantRequests) + } + // 35 callers are neither shown individually nor evicted. + if rows[5].Client != formatCount(35)+" not shown" { + t.Errorf("overflow row = %q, want 35 not shown", rows[5].Client) + } +} + +// The overflow row reports two counts that are true of different things: how +// many tracked callers fell below the display limit, which is exact, and how +// many fold-ins the eviction path has done, which is not a count of callers — +// a caller that is evicted and comes back is folded in twice. Reporting the sum +// as "N not shown" claimed more distinct callers than the row stands for. +func TestSourceTrackerOverflowSeparatesHiddenFromEvicted(t *testing.T) { + tr := sourceTracker{max: 3} + + // Fill, then push the first caller out, then bring it back and push it out + // again: two fold-ins, one caller. + for _, addr := range []string{"10.0.0.1", "10.0.0.2", "10.0.0.3"} { + tr.Record(addr, "npm", 100) + time.Sleep(time.Millisecond) + } + tr.Record("10.0.0.4", "npm", 100) // evicts .1 + time.Sleep(time.Millisecond) + tr.Record("10.0.0.1", "npm", 100) // .1 returns, evicting .2 + time.Sleep(time.Millisecond) + tr.Record("10.0.0.5", "npm", 100) // evicts .3 + + if tr.evictions != 3 { + t.Fatalf("evictions = %d, want 3 fold-ins", tr.evictions) + } + + rows := tr.Top(2) + overflow := rows[len(rows)-1] + if !overflow.IsOverflow { + t.Fatalf("last row is not the overflow row: %+v", overflow) + } + // One of the three tracked callers is below the cut; three fold-ins + // happened, covering two distinct callers. + if want := "1 not shown, 3 evicted"; overflow.Client != want { + t.Errorf("overflow row = %q, want %q", overflow.Client, want) + } + + var got int64 + for _, r := range rows { + got += parseCount(t, r.Requests) + } + if got != 6 { + t.Errorf("rows account for %d requests, want all 6", got) + } +} + +// With nothing evicted the row says only what it can count exactly. +func TestSourceTrackerOverflowOmitsEvictionsWhenThereAreNone(t *testing.T) { + var tr sourceTracker + for i := range 5 { + tr.Record("10.0.0."+strconv.Itoa(i), "npm", 100) + } + + rows := tr.Top(2) + if want := "3 not shown"; rows[len(rows)-1].Client != want { + t.Errorf("overflow row = %q, want %q", rows[len(rows)-1].Client, want) + } +} + +// The eviction scan is linear over the table, and Record holds a process-wide +// lock while it runs. These two bound what that costs in the case the source +// table is designed for — containerised CI, where every job has a fresh address +// and the table therefore sits permanently at its limit, so every request pays +// a scan. Measured at 200 entries on an M4 Max: ~2us evicting against ~40ns for +// a caller already tracked. Two microseconds against a request that goes to the +// network is not worth a heap or an LRU list, but the gap is the reason this is +// bounded at 200 rather than at something larger. +func BenchmarkRecordAlwaysEvicting(b *testing.B) { + var t sourceTracker + for i := 0; i < maxTrackedSources; i++ { + t.Record("10.0."+strconv.Itoa(i/256)+"."+strconv.Itoa(i%256), "npm", 1024) + } + + addrs := make([]string, b.N) + for i := range addrs { + addrs[i] = "172.16." + strconv.Itoa(i/256%256) + "." + strconv.Itoa(i%256) + } + + b.ResetTimer() + for i := 0; i < b.N; i++ { + t.Record(addrs[i], "npm", 1024) + } +} + +func BenchmarkRecordExisting(b *testing.B) { + var t sourceTracker + t.Record("10.0.0.1", "npm", 1024) + b.ResetTimer() + for i := 0; i < b.N; i++ { + t.Record("10.0.0.1", "npm", 1024) + } +} + +// The page is unauthenticated, so the caller table is published only when +// asked for. Off is the default, and off has to mean the addresses are absent +// from the HTML, not merely unstyled. +func TestAnalyticsPageHidesSourcesByDefault(t *testing.T) { + ts := newTestServer(t) + defer ts.close() + + if ts.server.cfg.UIRequestSources { + t.Fatal("UIRequestSources defaults to true; it must default off") + } + ts.server.sources.Record("10.1.2.3", "npm", 4096) + + body := ts.getOK(t, "/ui/analytics") + if strings.Contains(body, "10.1.2.3") { + t.Error("a caller address rendered with ui_request_sources off") + } + if strings.Contains(body, "Request sources") { + t.Error("the sources card rendered with ui_request_sources off") + } +} + +func TestAnalyticsPageShowsSourcesWhenEnabled(t *testing.T) { + ts := newTestServer(t) + defer ts.close() + + ts.server.cfg.UIRequestSources = true + ts.server.sources.Record("10.1.2.3", "npm", 4096) + + body := ts.getOK(t, "/ui/analytics") + for _, want := range []string{"Request sources", "10.1.2.3", "4.0 KB"} { + if !strings.Contains(body, want) { + t.Errorf("rendered page missing %q", want) + } + } +} diff --git a/internal/server/templates/components/runtime_metrics.html b/internal/server/templates/components/runtime_metrics.html index 83ff5f7f..e464559e 100644 --- a/internal/server/templates/components/runtime_metrics.html +++ b/internal/server/templates/components/runtime_metrics.html @@ -91,6 +91,70 @@

Circuit breakers

+ + {{if .SourcesOn}} +
+
+

Request sources

+

+ {{if .TrustsForward}}Attributed by X-Forwarded-For{{else}}Attributed by peer address{{end}} + {{- if .SourceCount}} · {{.SourceCount}} tracked{{end}} +

+
+ +
+
+ {{if .Sources}} + + + + + + + + + + + + {{range .Sources}} + + + + + + + + {{end}} + +
AddressClientRequestsDownloadedLast seen
{{.Addr}}{{.Client}}{{.Requests}}{{.Bytes}}{{.LastSeen}}
+ {{else}} +

No requests recorded yet.

+ {{end}} +
+ +
+

By client

+ {{if .Clients}} + + + {{range .Clients}} + + + + + + {{end}} + +
{{.Client}}{{.Requests}}{{.Bytes}}
+ {{else}} +

None yet.

+ {{end}} +
+
+
+ {{end}} + {{if .ScanningOn}}

Pre-cache scanning

From 4261bcbc3fa36118b34156730c1a91ab01bf02ee Mon Sep 17 00:00:00 2001 From: Andrew Nesbitt Date: Fri, 2 Oct 2026 10:34:41 +0100 Subject: [PATCH 6/6] Show the by-client table regardless of ui_request_sources The flag exists to withhold caller addresses. The by-client table carries none, and the same figures are already public at /metrics, so gating it left two metrics with no tile on the page in the default config. --- internal/server/sources_test.go | 10 ++++++++-- .../server/templates/components/runtime_metrics.html | 12 +++++++++--- 2 files changed, 17 insertions(+), 5 deletions(-) diff --git a/internal/server/sources_test.go b/internal/server/sources_test.go index 1618d342..fa58dd73 100644 --- a/internal/server/sources_test.go +++ b/internal/server/sources_test.go @@ -423,8 +423,14 @@ func TestAnalyticsPageHidesSourcesByDefault(t *testing.T) { if strings.Contains(body, "10.1.2.3") { t.Error("a caller address rendered with ui_request_sources off") } - if strings.Contains(body, "Request sources") { - t.Error("the sources card rendered with ui_request_sources off") + // The by-client table carries no addresses and the same figures are already + // public at /metrics, so it renders either way; only the address table is + // gated. + if !strings.Contains(body, "By client") { + t.Error("the by-client table was gated along with the address table") + } + if !strings.Contains(body, "ui_request_sources") { + t.Error("the off-state note naming the config key did not render") } } diff --git a/internal/server/templates/components/runtime_metrics.html b/internal/server/templates/components/runtime_metrics.html index e464559e..03fc9a04 100644 --- a/internal/server/templates/components/runtime_metrics.html +++ b/internal/server/templates/components/runtime_metrics.html @@ -93,19 +93,26 @@

Circuit breakers

- {{if .SourcesOn}}

Request sources

+ {{if .SourcesOn}}

{{if .TrustsForward}}Attributed by X-Forwarded-For{{else}}Attributed by peer address{{end}} {{- if .SourceCount}} · {{.SourceCount}} tracked{{end}}

+ {{end}}
- {{if .Sources}} + {{if not .SourcesOn}} +

+ Per-caller addresses are off. Set ui_request_sources: true + to show them here, once /ui is behind authentication; + the per-request detail is in the access log. +

+ {{else if .Sources}} @@ -153,7 +160,6 @@