Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 30 additions & 3 deletions braintrust/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -217,6 +217,32 @@ Size the request for the pod's full local-storage usage:

When you enable `tmpVolume`, make sure the `ephemeralStorage.request` still covers that extra space.

## GKE API Autoscaling

The API can autoscale on GKE using a Horizontal Pod Autoscaler backed by GKE's native `AutoscalingMetric` resource. When enabled, each API pool scales on three signals - CPU (scoped to the `api` container via `ContainerResource`, so sidecars are excluded), Node.js event-loop utilization, and mean event-loop delay.

This is underpinned by a **Preview (Pre-GA)** GKE feature. It requires:
Comment thread
soldatchenko marked this conversation as resolved.

- Braintrust API / data plane **v2.9.0** or later (Prometheus `/metrics` on the API health server)
- GKE **1.35.1-gke.1396000** or later
Comment thread
soldatchenko marked this conversation as resolved.
- The Performance HPA profile and the Autoscaling API enabled on the cluster
- `roles/autoscaling.metricsWriter` granted to all node service accounts
- The Autoscaling API included in your service perimeter when using VPC Service Controls
Comment thread
brianvans marked this conversation as resolved.

See [Expose custom metrics for autoscaling](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/expose-custom-metrics-autoscaling) for more details on `AutoscalingMetric` in GKE.

Enable it in your values:

```yaml
api:
autoscaling:
enabled: true
minReplicas: 4
maxReplicas: 50
```

When enabled for a pool, that pool's `replicas` setting is ignored and the HPA controls the replica count. With `api.workloadIsolation.enabled`, ingest and background pools inherit these settings and can override `minReplicas` / `maxReplicas` under `api.workloadIsolation.<pool>.autoscaling`.

## API workload isolation

`api.workloadIsolation.enabled` creates fixed-capacity `braintrust-api-ingest`
Expand Down Expand Up @@ -259,9 +285,9 @@ ingest, eval, function, and automation routes match `POST`; proxy routes match
all methods. GKE Ingress cannot route by method, so its equivalent integration
classifies matching paths for all methods.

This feature does not enable autoscaling. Configure fixed replica counts under
`api.replicas`, `api.workloadIsolation.ingest.replicas`, and
`api.workloadIsolation.background.replicas`.
Pools use fixed replica counts by default (`api.replicas` and
`api.workloadIsolation.<pool>.replicas`). On GKE, enable `api.autoscaling` to
let each pool scale independently instead.

## Testing

Expand Down Expand Up @@ -311,4 +337,5 @@ Example values files for different cloud providers and configurations are locate

- `examples/google-autopilot/values.yaml`: GKE Autopilot deployment.
- `examples/google-autopilot-cel/values.yaml`: GKE Autopilot deployment with CEL-friendly security settings.
- `examples/google-api-isolation-autoscaling/values.yaml`: Minimal example for API workload isolation and per-pool autoscaling on GKE (combine with an Autopilot or Standard values file).
- `examples/google-standard/values.yaml`: GKE Standard deployment.
48 changes: 48 additions & 0 deletions braintrust/examples/google-api-isolation-autoscaling/values.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Minimal GKE example with API workload isolation + per-pool autoscaling.
#
# Combine this with a normal google-autopilot (or google-standard) values
# file - only the fields below are specific to this pattern.
#
# Sizing:
# - api.autoscaling is shared by every pool
# - api.workloadIsolation.<pool>.autoscaling deep-merges on top
# - With autoscaling enabled, set capacity via minReplicas/maxReplicas
# (Helm `replicas` is not applied to those Deployments)
#
# Also requires GKE 1.35.1+ with AutoscalingMetric, and ingress routes for the
# isolation contract (see README & files/contracts/api-workload-isolation-routes.yaml).

api:
autoscaling:
enabled: true
minReplicas: 3
maxReplicas: 40
# Metric targets (inherited by every pool unless overridden below).
cpu:
targetAverageUtilization: 50
eventLoopUtilization:
targetAverageValue: "0.4" # 0-1 ratio (0.4 is 40%)
eventLoopDelayMean:
targetAverageValue: "0.05" # seconds (0.05 is 50ms)
# HPA scale velocity (inherited by every pool unless overridden below).
behavior:
scaleDown:
stabilizationWindowSeconds: 300
scaleUp:
stabilizationWindowSeconds: 60
workloadIsolation:
enabled: true
ingest:
autoscaling:
minReplicas: 3
maxReplicas: 60
# Example pool-specific overrides: higher ELU target, faster scale-out.
eventLoopUtilization:
targetAverageValue: "0.5" # 0-1 ratio (0.5 is 50%)
behavior:
scaleUp:
stabilizationWindowSeconds: 30
background:
autoscaling:
minReplicas: 3
maxReplicas: 50
8 changes: 8 additions & 0 deletions braintrust/examples/google-autopilot-cel/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,14 @@ api:
service:
networking.gke.io/load-balancer-type: "Internal"
replicas: 4
# Alternatively, autoscale the API on CPU and event-loop metrics.
# This is a Preview (Pre-GA) GKE feature requiring 1.35.1-gke.1396000 or later
# and the prerequisites documented in the main values.yaml.
# See api.autoscaling in the main values.yaml file for more details.
# autoscaling:
# enabled: true
# minReplicas: 4
# maxReplicas: 50
service:
type: LoadBalancer
port: 8000
Expand Down
8 changes: 8 additions & 0 deletions braintrust/examples/google-autopilot/values.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,14 @@ api:
service:
networking.gke.io/load-balancer-type: "Internal"
replicas: 4
# Alternatively, autoscale the API on CPU and event-loop metrics.
# This is a Preview (Pre-GA) GKE feature requiring 1.35.1-gke.1396000 or later
# and the prerequisites documented in the main values.yaml.
# See api.autoscaling in the main values.yaml file for more details.
# autoscaling:
# enabled: true
# minReplicas: 4
# maxReplicas: 50
# Uncomment the following section to use a different image or tag from the version in the Helm release
#image:
#repository: public.ecr.aws/braintrust/standalone-api
Expand Down
9 changes: 9 additions & 0 deletions braintrust/templates/_api-deployment.tpl
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,9 @@ metadata:
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
{{- if not (dig "autoscaling" "enabled" false $api) }}
replicas: {{ $api.replicas }}
{{- end }}
strategy:
type: {{ $api.strategy.type }}
{{- with $api.strategy.rollingUpdate }}
Expand Down Expand Up @@ -101,6 +103,9 @@ spec:
{{- end }}
ports:
- containerPort: {{ $api.service.port }}
{{- if dig "autoscaling" "enabled" false $api }}
- containerPort: {{ $api.healthServer.port }}
{{- end }}
resources:
{{- toYaml $api.resources | nindent 12 }}
{{- with $api.livenessProbe }}
Expand Down Expand Up @@ -176,6 +181,10 @@ spec:
{{- with $api.extraEnvVars }}
{{- toYaml . | nindent 12 }}
{{- end }}
{{- if dig "autoscaling" "enabled" false $api }}
- name: ENABLE_PROMETHEUS_METRICS
value: "true"
{{- end }}
{{- if or $api.tmpVolume.enabled (and (eq $root.Values.cloud "azure") $root.Values.azure.enableAzureKeyVaultDriver) $customCA.enabled }}
volumeMounts:
{{- if $api.tmpVolume.enabled }}
Expand Down
12 changes: 12 additions & 0 deletions braintrust/templates/_helpers.tpl
Original file line number Diff line number Diff line change
Expand Up @@ -115,6 +115,18 @@ Internal cluster URL for the AI Gateway service.
http://{{ .Values.aiGateway.service.name | default .Values.aiGateway.name }}.{{ include "braintrust.namespace" . }}:{{ .Values.aiGateway.service.port }}
{{- end -}}

{{/*
Validate API autoscaling prerequisites (GKE + AutoscalingMetric CRD).
*/}}
{{- define "braintrust.apiAutoscaling.validate" -}}
{{- if ne .Values.cloud "google" }}
{{- fail "api.autoscaling is currently only supported when cloud is google (GKE)" }}
{{- end }}
{{- if not (.Capabilities.APIVersions.Has "autoscaling.gke.io/v1beta1") }}
{{- fail "api.autoscaling requires the AutoscalingMetric API (autoscaling.gke.io/v1beta1). Use GKE 1.35.1 or later, or verify with: kubectl api-resources | grep autoscalingmetric. For helm template without a cluster, pass --api-versions=autoscaling.gke.io/v1beta1." }}
{{- end }}
{{- end -}}

{{/*
Render Brainstore container resources with provider-specific ephemeral storage.

Expand Down
53 changes: 53 additions & 0 deletions braintrust/templates/api-autoscaling-metric.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
{{- $root := . -}}
{{- $pools := include "braintrust.apiPools" . | fromYamlArray -}}
{{- $validated := false -}}
{{- $rendered := 0 -}}
{{- range $pool := $pools -}}
{{- $api := $pool.config -}}
{{- if dig "autoscaling" "enabled" false $api -}}
{{- if not $validated -}}
{{- include "braintrust.apiAutoscaling.validate" $root -}}
{{- $validated = true -}}
{{- end -}}
{{- if gt $rendered 0 }}
---
{{- end }}
{{- $poolLabels := dict -}}
{{- if or $root.Values.api.workloadIsolation.enabled (ne $pool.role "default") -}}
{{- $_ := set $poolLabels "braintrust.dev/api-pool" $pool.role -}}
{{- end -}}
{{- $resourceLabels := mergeOverwrite (deepCopy $root.Values.global.labels) (deepCopy $api.labels) $poolLabels -}}
apiVersion: autoscaling.gke.io/v1beta1
kind: AutoscalingMetric
metadata:
name: {{ $api.name }}
namespace: {{ include "braintrust.namespace" $root }}
{{- with $resourceLabels }}
labels:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with $api.annotations.autoscalingMetric }}
annotations:
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
metrics:
- pod:
selector:
matchLabels:
app: {{ $api.name }}
containers:
- endpoint:
port: {{ $api.healthServer.port }}
path: {{ $api.autoscaling.metricsPath }}
metrics:
- gauge:
# GKE gauge names must match ^[a-z]([-a-z0-9]*[a-z0-9])?
name: braintrust-api-event-loop-utilization-ratio
prometheusMetricName: braintrust_api_event_loop_utilization_ratio
- gauge:
name: braintrust-api-event-loop-delay-mean-seconds
prometheusMetricName: braintrust_api_event_loop_delay_mean_seconds
{{- $rendered = add1 $rendered -}}
{{- end -}}
{{- end }}
70 changes: 70 additions & 0 deletions braintrust/templates/api-hpa.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
{{- $root := . -}}
{{- $pools := include "braintrust.apiPools" . | fromYamlArray -}}
{{- $validated := false -}}
{{- $rendered := 0 -}}
{{- range $pool := $pools -}}
{{- $api := $pool.config -}}
{{- if dig "autoscaling" "enabled" false $api -}}
{{- if not $validated -}}
{{- include "braintrust.apiAutoscaling.validate" $root -}}
{{- $validated = true -}}
{{- end -}}
{{- if gt $rendered 0 }}
---
{{- end }}
{{- $poolLabels := dict -}}
{{- if or $root.Values.api.workloadIsolation.enabled (ne $pool.role "default") -}}
{{- $_ := set $poolLabels "braintrust.dev/api-pool" $pool.role -}}
{{- end -}}
{{- $resourceLabels := mergeOverwrite (deepCopy $root.Values.global.labels) (deepCopy $api.labels) $poolLabels -}}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: {{ $api.name }}
namespace: {{ include "braintrust.namespace" $root }}
{{- with $resourceLabels }}
labels:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with $api.annotations.hpa }}
annotations:
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: {{ $api.name }}
minReplicas: {{ $api.autoscaling.minReplicas }}
maxReplicas: {{ $api.autoscaling.maxReplicas }}
metrics:
# ContainerResource scopes CPU to the api container so sidecars / extraContainers
# do not skew utilization. Custom Pods metrics are already API-scoped via /metrics.
- type: ContainerResource
containerResource:
name: cpu
container: api
target:
type: Utilization
averageUtilization: {{ $api.autoscaling.cpu.targetAverageUtilization }}
- type: Pods
pods:
metric:
name: autoscaling.gke.io|{{ $api.name }}|braintrust-api-event-loop-utilization-ratio
target:
type: AverageValue
averageValue: {{ $api.autoscaling.eventLoopUtilization.targetAverageValue | quote }}
- type: Pods
pods:
metric:
name: autoscaling.gke.io|{{ $api.name }}|braintrust-api-event-loop-delay-mean-seconds
target:
type: AverageValue
averageValue: {{ $api.autoscaling.eventLoopDelayMean.targetAverageValue | quote }}
{{- with $api.autoscaling.behavior }}
behavior:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- $rendered = add1 $rendered -}}
{{- end -}}
{{- end }}
17 changes: 17 additions & 0 deletions braintrust/tests/api-autoscaling-crd_test.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
suite: test API autoscaling CRD requirement
templates:
- api-hpa.yaml
# Do not advertise autoscaling.gke.io here — this suite asserts the fail path.
tests:
- it: should fail when AutoscalingMetric API is unavailable
template: api-hpa.yaml
values:
- __fixtures__/base-values.yaml
set:
cloud: google
api.autoscaling.enabled: true
release:
namespace: "braintrust"
asserts:
- failedTemplate:
errorPattern: "api\\.autoscaling requires the AutoscalingMetric API \\(autoscaling\\.gke\\.io/v1beta1\\)"
Loading
Loading