Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 36 additions & 3 deletions braintrust/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -217,6 +217,38 @@ Size the request for the pod's full local-storage usage:

When you enable `tmpVolume`, make sure the `ephemeralStorage.request` still covers that extra space.

## GKE API Autoscaling

The API can autoscale on GKE using a Horizontal Pod Autoscaler backed by GKE's native `AutoscalingMetric` resource. When enabled, each API pool scales on three signals - CPU (scoped to the `api` container via `ContainerResource`, so sidecars are excluded), Node.js event-loop utilization, and mean event-loop delay.

This is underpinned by a **Preview (Pre-GA)** GKE feature. It requires:

- Braintrust API / data plane **v2.9.0** or later (Prometheus `/metrics` on the API health server)
- GKE **1.35.1-gke.1396000** or later
- The Performance HPA profile and the Autoscaling API enabled on the cluster
- `roles/autoscaling.metricsWriter` granted to all node service accounts
- The Autoscaling API included in your service perimeter when using VPC Service Controls

See [Expose custom metrics for autoscaling](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/expose-custom-metrics-autoscaling) for more details on `AutoscalingMetric` in GKE.

Enable it in your values:

```yaml
api:
autoscaling:
enabled: true
minReplicas: 4
maxReplicas: 50
```

When enabled for a pool, that pool's `replicas` setting is ignored and the HPA controls the replica count. With `api.workloadIsolation.enabled`, ingest and background pools inherit these settings and can override `minReplicas` / `maxReplicas` under `api.workloadIsolation.<pool>.autoscaling`.

## EKS API Autoscaling

On AWS (`cloud: aws`), the same `api.autoscaling` values deploy an in-chart Prometheus scrape of each API pool's health `/metrics` endpoint and a prometheus-adapter that exposes event-loop gauges to HPA via `custom.metrics.k8s.io`. Targets match GKE / ECS defaults (CPU 50% on the `api` container, event-loop utilization `0.4`, delay mean `0.05s`).

Requires API image **v2.9.0+**. With workload isolation enabled, Prometheus scrapes every pool (`api.name`, ingest, and background).

## API workload isolation

`api.workloadIsolation.enabled` creates fixed-capacity `braintrust-api-ingest`
Expand Down Expand Up @@ -259,9 +291,9 @@ ingest, eval, function, and automation routes match `POST`; proxy routes match
all methods. GKE Ingress cannot route by method, so its equivalent integration
classifies matching paths for all methods.

This feature does not enable autoscaling. Configure fixed replica counts under
`api.replicas`, `api.workloadIsolation.ingest.replicas`, and
`api.workloadIsolation.background.replicas`.
Pools use fixed replica counts by default (`api.replicas` and
`api.workloadIsolation.<pool>.replicas`). On GKE, enable `api.autoscaling` to
let each pool scale independently instead.

## Testing

Expand Down Expand Up @@ -311,4 +343,5 @@ Example values files for different cloud providers and configurations are locate

- `examples/google-autopilot/values.yaml`: GKE Autopilot deployment.
- `examples/google-autopilot-cel/values.yaml`: GKE Autopilot deployment with CEL-friendly security settings.
- `examples/google-api-isolation-autoscaling/values.yaml`: Minimal example for API workload isolation and per-pool autoscaling on GKE (combine with an Autopilot or Standard values file).
- `examples/google-standard/values.yaml`: GKE Standard deployment.
48 changes: 48 additions & 0 deletions braintrust/examples/google-api-isolation-autoscaling/values.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Minimal GKE example with API workload isolation + per-pool autoscaling.
#
# Combine this with a normal google-autopilot (or google-standard) values
# file - only the fields below are specific to this pattern.
#
# Sizing:
# - api.autoscaling is shared by every pool
# - api.workloadIsolation.<pool>.autoscaling deep-merges on top
# - With autoscaling enabled, set capacity via minReplicas/maxReplicas
# (Helm `replicas` is not applied to those Deployments)
#
# Also requires GKE 1.35.1+ with AutoscalingMetric, and ingress routes for the
# isolation contract (see README & files/contracts/api-workload-isolation-routes.yaml).

api:
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 40
# Metric targets (inherited by every pool unless overridden below).
cpu:
targetAverageUtilization: 50
eventLoopUtilization:
targetAverageValue: "0.4" # 0-1 ratio (0.4 is 40%)
eventLoopDelayMean:
targetAverageValue: "0.05" # seconds (0.05 is 50ms)
# HPA scale velocity (inherited by every pool unless overridden below).
behavior:
scaleDown:
stabilizationWindowSeconds: 300
scaleUp:
stabilizationWindowSeconds: 60
workloadIsolation:
enabled: true
ingest:
autoscaling:
minReplicas: 3
maxReplicas: 60
# Example pool-specific overrides: higher ELU target, faster scale-out.
eventLoopUtilization:
targetAverageValue: "0.5" # 0-1 ratio (0.5 is 50%)
behavior:
scaleUp:
stabilizationWindowSeconds: 30
background:
autoscaling:
minReplicas: 3
maxReplicas: 50
8 changes: 8 additions & 0 deletions braintrust/examples/google-autopilot-cel/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,14 @@ api:
service:
networking.gke.io/load-balancer-type: "Internal"
replicas: 4
# Alternatively, autoscale the API on CPU and event-loop metrics.
# This is a Preview (Pre-GA) GKE feature requiring 1.35.1-gke.1396000 or later
# and the prerequisites documented in the main values.yaml.
# See api.autoscaling in the main values.yaml file for more details.
# autoscaling:
# enabled: true
# minReplicas: 4
# maxReplicas: 50
service:
type: LoadBalancer
port: 8000
Expand Down
8 changes: 8 additions & 0 deletions braintrust/examples/google-autopilot/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,14 @@ api:
service:
networking.gke.io/load-balancer-type: "Internal"
replicas: 4
# Alternatively, autoscale the API on CPU and event-loop metrics.
# This is a Preview (Pre-GA) GKE feature requiring 1.35.1-gke.1396000 or later
# and the prerequisites documented in the main values.yaml.
# See api.autoscaling in the main values.yaml file for more details.
# autoscaling:
# enabled: true
# minReplicas: 4
# maxReplicas: 50
# Uncomment the following section to use a different image or tag from the version in the Helm release
#image:
#repository: public.ecr.aws/braintrust/standalone-api
Expand Down
9 changes: 9 additions & 0 deletions braintrust/templates/_api-deployment.tpl
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,9 @@ metadata:
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if not (dig "autoscaling" "enabled" false $api) }}
replicas: {{ $api.replicas }}
{{- end }}
strategy:
type: {{ $api.strategy.type }}
{{- with $api.strategy.rollingUpdate }}
Expand Down Expand Up @@ -101,6 +103,9 @@ spec:
{{- end }}
ports:
- containerPort: {{ $api.service.port }}
{{- if dig "autoscaling" "enabled" false $api }}
- containerPort: {{ $api.healthServer.port }}
{{- end }}
resources:
{{- toYaml $api.resources | nindent 12 }}
{{- with $api.livenessProbe }}
Expand Down Expand Up @@ -176,6 +181,10 @@ spec:
{{- with $api.extraEnvVars }}
{{- toYaml . | nindent 12 }}
{{- end }}
{{- if dig "autoscaling" "enabled" false $api }}
- name: ENABLE_PROMETHEUS_METRICS
value: "true"
{{- end }}
{{- if or $api.tmpVolume.enabled (and (eq $root.Values.cloud "azure") $root.Values.azure.enableAzureKeyVaultDriver) $customCA.enabled }}
volumeMounts:
{{- if $api.tmpVolume.enabled }}
Expand Down
15 changes: 15 additions & 0 deletions braintrust/templates/_helpers.tpl
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,21 @@ Internal cluster URL for the AI Gateway service.
http://{{ .Values.aiGateway.service.name | default .Values.aiGateway.name }}.{{ include "braintrust.namespace" . }}:{{ .Values.aiGateway.service.port }}
{{- end -}}

{{/*
Validate API autoscaling prerequisites.
GKE requires AutoscalingMetric (autoscaling.gke.io/v1beta1).
EKS uses in-chart Prometheus + prometheus-adapter (no GKE CRD).
*/}}
{{- define "braintrust.apiAutoscaling.validate" -}}
{{- if eq .Values.cloud "google" }}
{{- if not (.Capabilities.APIVersions.Has "autoscaling.gke.io/v1beta1") }}
{{- fail "api.autoscaling requires the AutoscalingMetric API (autoscaling.gke.io/v1beta1). Use GKE 1.35.1 or later, or verify with: kubectl api-resources | grep autoscalingmetric. For helm template without a cluster, pass --api-versions=autoscaling.gke.io/v1beta1." }}
{{- end }}
{{- else if ne .Values.cloud "aws" }}
{{- fail "api.autoscaling is currently only supported when cloud is google (GKE) or aws (EKS)" }}
{{- end }}
{{- end -}}

{{/*
Render Brainstore container resources with provider-specific ephemeral storage.

Expand Down
55 changes: 55 additions & 0 deletions braintrust/templates/api-autoscaling-metric.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
{{- $root := . -}}
{{- if eq $root.Values.cloud "google" -}}
{{- $pools := include "braintrust.apiPools" . | fromYamlArray -}}
{{- $validated := false -}}
{{- $rendered := 0 -}}
{{- range $pool := $pools -}}
{{- $api := $pool.config -}}
{{- if dig "autoscaling" "enabled" false $api -}}
{{- if not $validated -}}
{{- include "braintrust.apiAutoscaling.validate" $root -}}
{{- $validated = true -}}
{{- end -}}
{{- if gt $rendered 0 }}
---
{{- end }}
{{- $poolLabels := dict -}}
{{- if or $root.Values.api.workloadIsolation.enabled (ne $pool.role "default") -}}
{{- $_ := set $poolLabels "braintrust.dev/api-pool" $pool.role -}}
{{- end -}}
{{- $resourceLabels := mergeOverwrite (deepCopy $root.Values.global.labels) (deepCopy $api.labels) $poolLabels -}}
apiVersion: autoscaling.gke.io/v1beta1
kind: AutoscalingMetric
metadata:
name: {{ $api.name }}
namespace: {{ include "braintrust.namespace" $root }}
{{- with $resourceLabels }}
labels:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with $api.annotations.autoscalingMetric }}
annotations:
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
metrics:
- pod:
selector:
matchLabels:
app: {{ $api.name }}
containers:
- endpoint:
port: {{ $api.healthServer.port }}
path: {{ $api.autoscaling.metricsPath }}
metrics:
- gauge:
# GKE gauge names must match ^[a-z]([-a-z0-9]*[a-z0-9])?
name: braintrust-api-event-loop-utilization-ratio
prometheusMetricName: braintrust_api_event_loop_utilization_ratio
- gauge:
name: braintrust-api-event-loop-delay-mean-seconds
prometheusMetricName: braintrust_api_event_loop_delay_mean_seconds
{{- $rendered = add1 $rendered -}}
{{- end -}}
{{- end -}}
{{- end }}
Loading
Loading