Skip to content

Repository files navigation

Freestyle Browser Bench

Run a hosted Chrome throughput demo with one Freestyle API key. It creates the VMs, installs pinned browser/driver builds, runs the workload, prints scores, saves readable and JSON reports, and deletes its test VMs afterwards.

No Docker, SSH, Teleport, existing VM, browser endpoint or Freestyle CLI login is required. You need Node.js 22.22.2 or newer, npm, Internet access, and a Freestyle account with enough VM quota. The VMs are billable.

Quick start

git clone https://github.com/freestyle-sh/browser-bench.git
cd browser-bench
npm ci

Set FREESTYLE_API_KEY in your environment, or create a local .env containing:

FREESTYLE_API_KEY=your-key-here

Then:

npm run bench

The default is one hosted Node driver VM → one separate hosted Chrome VM, with WebSocket compression disabled and Chrome's screenshot surface optimization enabled. The program on your computer provisions the resources and retrieves results; the benchmark itself runs in the driver VM. This is intentionally not the same as Playwright running on your laptop.

Default resources: 8-vCPU / 16-GiB driver + 4-vCPU / 8-GiB browser. There are five unscored warmups and twenty measured sessions. Installation takes several minutes on fresh VMs; progress is printed throughout. The exact VM location is platform-managed, not pinned or independently guaranteed by this tool. Measured CDP round-trip time tells you how short the actual path is.

Preview the resource budget without credentials or charges:

npm run bench -- --dry-run

For smaller account quotas, use --driver-size md --browser-size sm, or --mode local --browser-size sm (one two-vCPU hosted browser). This changes the configuration; do not expect the same performance.

Choose a layout

Mode Where Playwright runs Where Chrome runs Default hardware
split Separate hosted driver VM One VM per worker 8 driver + 4 browser vCPUs per worker
local Your computer One hosted VM per worker 4 browser vCPUs per worker
packed Shared large VM Separate Chrome processes in that same VM One 16-vCPU VM
fleet Separate hosted driver VM One small VM per worker 8 driver + 2 browser vCPUs per worker

split and fleet use the same routing/orchestration, with different default browser sizes. Every worker has a separate Chrome process; even in packed mode we do not collapse multiple workers into contexts in a single Chrome process. All modes create a fresh isolated browser context for each measured session.

# Original laptop/client topology
npm run bench -- --mode local

# One large VM runs four Chromes plus the driver
npm run bench -- --mode packed --workers 4

# Four small VMs, one Chrome each, with a separate driver
npm run bench -- --mode fleet --workers 4

# Additional experimental client-runtime optimization, installed automatically
npm run bench -- --runtime bun

# Packed-only loopback experiment: bypass the public proxy, still use guest auth
npm run bench -- --mode packed --workers 4 --transport loopback

# Compression control
npm run bench -- --compression on

# Longer sample (100 measured sessions per worker)
npm run bench -- --iterations 100

--workers is concurrency, not repetitions. Four workers with twenty iterations means eighty measured sessions, up to four running simultaneously. The driver warms all workers before starting their measured runs together. Parallel results are labeled parallel-workload estimates: do not compare them to sequential leaderboard scores or advertise aggregate actions/sec as single-session speed.

Public authenticated HTTPS/WSS is the default in every layout, including packed. --transport loopback explicitly changes the network path and is available only for packed mode. It is not evidence that remote client connections got faster.

What gets optimized?

  • Shorter observed driver-to-browser path: split/fleet put the driver on hosted Freestyle. Placement is not guaranteed; check the measured RTT.
  • Separate driver CPU resources: split/fleet avoid sharing the browser VM's guest CPU allocation with Playwright. Packed mode lets you measure that tradeoff.
  • WebSocket deflate off by default: only in these disposable VMs' nginx endpoints. Ordinary HTTP gzip/Brotli, PNG compression and website traffic are unchanged. This trades more CDP bandwidth for less compression/decompression work.
  • Optional Bun 1.4.2: a driver-runtime experiment, not a Playwright patch. The latest same-VM comparison found a +0.95-point pooled score gain and 5.8% lower mean task time, with a variable gain across rounds and some slower individual actions/tails. Node stays the default. Hosted Bun is installed side-by-side and verified against its official SHA256 manifest. Local Bun mode requires Bun 1.4.2 on your PATH.
  • Screenshot surface optimization in every mode: Chrome launches with --enable-features=CDPScreenshotNewSurface, matching Playwright 1.62.1's launch default. Connecting over CDP does not add launch flags, so we supply it ourselves. This changes frame synchronization, not PNG format, resolution, or quality. A paired hosted Node 24 test scored 75.72 control / 77.28 enabled, with 13.5% lower median screenshot time. Gains vary across runs; this is not an official score or a production stability certification. No Bun or benchmark action changes are required.
  • Fixed versions and explicit warmups: Node 22.22.2 in guests, Playwright 1.62.1, Chrome for Testing 151.0.7922.34. Warmups and installation are outside the measured task time, consistently across modes.

We deliberately do not enable changes that earlier experiments failed to support: larger proxy buffers, bigger V8 GC heaps, custom kernels, or assuming more browser cores automatically help. There is no CPU hot-add. Every size boots directly from its native public base snapshot. The tool checks for backwards monotonic-clock readings across CPUs before and after each run and refuses a complete score if that validation fails.

There is no changed link selector, skipped link enumeration, suppressed CDP preview generation, removed error stacks, lower-resolution image or lossy PNG substitute. The earlier demo's baseline runner is vendored byte-for-byte and hash-checked. Its optional optimized workload exists in the vendored source for integrity, but this CLI never invokes it.

Understand the output

Console output includes score estimate, success count, median/mean/p95 task times, click/screenshot/CDP latency and aggregate actions/sec. Each run writes:

results/<timestamp>-<mode>/
  report.md          readable results, caveats and per-worker scores
  report.html        self-contained, shareable dashboard; open in your browser
  results.json       raw actions, summaries, environment and clock checks
  resources.json     exact VM/route ownership and cleanup status; no credentials
  raw/               driver samples, including failures and warmups

Scores use the pinned legacy ComputeSDK formula: 40% median actions/sec, 25% median task latency, 20% task p95, 15% screenshot median, multiplied by session success rate, with the original 5% trimming. Failed sessions are retained; incomplete runs do not receive an aggregate score. Warmup failures stop scoring. The CLI exits nonzero on failed sessions, incomplete runs or cleanup failures.

Always read the untrimmed mean, p95, maximum and success count, not just the score. One large navigation stall can barely affect a trimmed score. Aggregate actions/sec uses the measured batch wall time, including connection/context work; it is a different metric from median actions/sec inside each task.

These are not official ComputeSDK results or a provider leaderboard rank. The workload/scoring is adapted from ComputeSDK commit 9ee03c6, not the latest upstream benchmark harness. It uses five fixed Wikipedia articles (Linux, Grace Hopper, Capybara, HTTP and San Francisco), a persistent externally managed browser, default PNG at 1920 × 1080 and the same ten-action sequence. It does not reproduce provider-specific stealth configurations, Virginia driver placement, browser lifecycle scores or every current upstream harness detail. See upstream methodology.

For useful comparisons, keep worker count, article count, runtime, compression, sizes and transport fixed except for the thing you intend to change. Run both orders (A–B then B–A); fresh VM placement, CPU sharing and Wikipedia/CDN state can move the numbers. Do not cherry-pick the fastest run.

Cleanup, security and cost

Default cleanup removes only VMs tagged with this run's random ownership ID and their matching TLS routes. Ctrl+C requests cancellation and cleanup; a currently running install/API call may take up to five minutes to return. Every VM also has a two-hour TTL, including retained VMs, as a backstop for laptop crashes or SIGKILL. The driver job has a one-hour ceiling. Very long runs can hit these limits and will be reported incomplete.

# Keep test VMs temporarily for inspection
npm run bench -- --keep

# Later, safely clean up that exact run (also retries failed cleanup)
npm run bench -- cleanup --run results/<your-run-directory>

The resource manifest is private to your output directory. Cleanup verifies live VM metadata and route destinations; it never deletes every VM in an account. Do not edit/delete the manifest until cleanup succeeds. Retained VMs and endpoint credentials can be inspected through your own Freestyle account; treat CDP access as control of the disposable browser VM.

The tool reuses owned style.dev hostnames with no configured route, preferring its own earlier browser-bench-* names. It never changes an active TLS route. When none are available, it claims a new random hostname. Deleting the TLS route does not release domain ownership: the API has no domain-release operation. Those free claims remain owned and can be reused on later runs. If every claim is actively routed and your domain quota is full, provisioning fails with cleanup.

Your Freestyle API key stays on your computer, is not written to reports, and is not sent into a VM. CDP endpoints have independent random bearer tokens; unauthenticated discovery must return 403 before timing starts. Only the authenticated nginx endpoint is published; Chrome listens on guest loopback. Remote worker credentials are mode-0600 files on disposable VMs; local worker configuration is mode-0600 and removed after the run. Results and .env are ignored by git. Don't publish your raw private output directory without reviewing it.

Chrome runs as an unprivileged guest user with --no-sandbox, using each disposable VM as the isolation boundary. Packed mode does not provide a separate security boundary between browser workers. This is a controlled demo against fixed public pages, not a production multi-tenant browser service. Don't put application secrets or untrusted customer workloads in these VMs. Restart-on-crash is disabled so the runner does not silently hide browser restarts.

Development

npm test
npm run check
npm run bench -- --help

CI runs unit/error-path tests without cloud credentials or billable VMs. Live validation is documented in TESTING.md. Public SDK: Freestyle quickstart.

MIT licensed. ComputeSDK-derived code retains its upstream license.

About

One-key hosted Chrome throughput demos: local, split, packed and fleet layouts with reproducible score reports

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages