Run a hosted Chrome throughput demo with one Freestyle API key. It creates the VMs, installs pinned browser/driver builds, runs the workload, prints scores, saves readable and JSON reports, and deletes its test VMs afterwards.
No Docker, SSH, Teleport, existing VM, browser endpoint or Freestyle CLI login is required. You need Node.js 22.22.2 or newer, npm, Internet access, and a Freestyle account with enough VM quota. The VMs are billable.
git clone https://github.com/freestyle-sh/browser-bench.git
cd browser-bench
npm ciSet FREESTYLE_API_KEY in your environment, or create a local .env containing:
FREESTYLE_API_KEY=your-key-hereThen:
npm run benchThe default is one hosted Node driver VM → one separate hosted Chrome VM, with WebSocket compression disabled and Chrome's screenshot surface optimization enabled. The program on your computer provisions the resources and retrieves results; the benchmark itself runs in the driver VM. This is intentionally not the same as Playwright running on your laptop.
Default resources: 8-vCPU / 16-GiB driver + 4-vCPU / 8-GiB browser. There are five unscored warmups and twenty measured sessions. Installation takes several minutes on fresh VMs; progress is printed throughout. The exact VM location is platform-managed, not pinned or independently guaranteed by this tool. Measured CDP round-trip time tells you how short the actual path is.
Preview the resource budget without credentials or charges:
npm run bench -- --dry-runFor smaller account quotas, use --driver-size md --browser-size sm, or
--mode local --browser-size sm (one two-vCPU hosted browser). This changes the
configuration; do not expect the same performance.
| Mode | Where Playwright runs | Where Chrome runs | Default hardware |
|---|---|---|---|
split |
Separate hosted driver VM | One VM per worker | 8 driver + 4 browser vCPUs per worker |
local |
Your computer | One hosted VM per worker | 4 browser vCPUs per worker |
packed |
Shared large VM | Separate Chrome processes in that same VM | One 16-vCPU VM |
fleet |
Separate hosted driver VM | One small VM per worker | 8 driver + 2 browser vCPUs per worker |
split and fleet use the same routing/orchestration, with different default
browser sizes. Every worker has a separate Chrome process; even in packed mode
we do not collapse multiple workers into contexts in a single Chrome process.
All modes create a fresh isolated browser context for each measured session.
# Original laptop/client topology
npm run bench -- --mode local
# One large VM runs four Chromes plus the driver
npm run bench -- --mode packed --workers 4
# Four small VMs, one Chrome each, with a separate driver
npm run bench -- --mode fleet --workers 4
# Additional experimental client-runtime optimization, installed automatically
npm run bench -- --runtime bun
# Packed-only loopback experiment: bypass the public proxy, still use guest auth
npm run bench -- --mode packed --workers 4 --transport loopback
# Compression control
npm run bench -- --compression on
# Longer sample (100 measured sessions per worker)
npm run bench -- --iterations 100--workers is concurrency, not repetitions. Four workers with twenty
iterations means eighty measured sessions, up to four running simultaneously.
The driver warms all workers before starting their measured runs together.
Parallel results are labeled parallel-workload estimates: do not compare
them to sequential leaderboard scores or advertise aggregate actions/sec as
single-session speed.
Public authenticated HTTPS/WSS is the default in every layout, including packed.
--transport loopback explicitly changes the network path and is available only
for packed mode. It is not evidence that remote client connections got faster.
- Shorter observed driver-to-browser path: split/fleet put the driver on hosted Freestyle. Placement is not guaranteed; check the measured RTT.
- Separate driver CPU resources: split/fleet avoid sharing the browser VM's guest CPU allocation with Playwright. Packed mode lets you measure that tradeoff.
- WebSocket deflate off by default: only in these disposable VMs' nginx endpoints. Ordinary HTTP gzip/Brotli, PNG compression and website traffic are unchanged. This trades more CDP bandwidth for less compression/decompression work.
- Optional Bun 1.4.2: a driver-runtime experiment, not a Playwright patch. The latest same-VM comparison found a +0.95-point pooled score gain and 5.8% lower mean task time, with a variable gain across rounds and some slower individual actions/tails. Node stays the default. Hosted Bun is installed side-by-side and verified against its official SHA256 manifest. Local Bun mode requires Bun 1.4.2 on your PATH.
- Screenshot surface optimization in every mode: Chrome launches with
--enable-features=CDPScreenshotNewSurface, matching Playwright 1.62.1's launch default. Connecting over CDP does not add launch flags, so we supply it ourselves. This changes frame synchronization, not PNG format, resolution, or quality. A paired hosted Node 24 test scored 75.72 control / 77.28 enabled, with 13.5% lower median screenshot time. Gains vary across runs; this is not an official score or a production stability certification. No Bun or benchmark action changes are required. - Fixed versions and explicit warmups: Node 22.22.2 in guests, Playwright 1.62.1, Chrome for Testing 151.0.7922.34. Warmups and installation are outside the measured task time, consistently across modes.
We deliberately do not enable changes that earlier experiments failed to support: larger proxy buffers, bigger V8 GC heaps, custom kernels, or assuming more browser cores automatically help. There is no CPU hot-add. Every size boots directly from its native public base snapshot. The tool checks for backwards monotonic-clock readings across CPUs before and after each run and refuses a complete score if that validation fails.
There is no changed link selector, skipped link enumeration, suppressed CDP
preview generation, removed error stacks, lower-resolution image or lossy PNG
substitute. The earlier demo's baseline runner is vendored byte-for-byte and
hash-checked. Its optional optimized workload exists in the vendored source
for integrity, but this CLI never invokes it.
Console output includes score estimate, success count, median/mean/p95 task times, click/screenshot/CDP latency and aggregate actions/sec. Each run writes:
results/<timestamp>-<mode>/
report.md readable results, caveats and per-worker scores
report.html self-contained, shareable dashboard; open in your browser
results.json raw actions, summaries, environment and clock checks
resources.json exact VM/route ownership and cleanup status; no credentials
raw/ driver samples, including failures and warmups
Scores use the pinned legacy ComputeSDK formula: 40% median actions/sec, 25% median task latency, 20% task p95, 15% screenshot median, multiplied by session success rate, with the original 5% trimming. Failed sessions are retained; incomplete runs do not receive an aggregate score. Warmup failures stop scoring. The CLI exits nonzero on failed sessions, incomplete runs or cleanup failures.
Always read the untrimmed mean, p95, maximum and success count, not just the score. One large navigation stall can barely affect a trimmed score. Aggregate actions/sec uses the measured batch wall time, including connection/context work; it is a different metric from median actions/sec inside each task.
These are not official ComputeSDK results or a provider leaderboard rank. The workload/scoring is adapted from ComputeSDK commit 9ee03c6, not the latest upstream benchmark harness. It uses five fixed Wikipedia articles (Linux, Grace Hopper, Capybara, HTTP and San Francisco), a persistent externally managed browser, default PNG at 1920 × 1080 and the same ten-action sequence. It does not reproduce provider-specific stealth configurations, Virginia driver placement, browser lifecycle scores or every current upstream harness detail. See upstream methodology.
For useful comparisons, keep worker count, article count, runtime, compression, sizes and transport fixed except for the thing you intend to change. Run both orders (A–B then B–A); fresh VM placement, CPU sharing and Wikipedia/CDN state can move the numbers. Do not cherry-pick the fastest run.
Default cleanup removes only VMs tagged with this run's random ownership ID and their matching TLS routes. Ctrl+C requests cancellation and cleanup; a currently running install/API call may take up to five minutes to return. Every VM also has a two-hour TTL, including retained VMs, as a backstop for laptop crashes or SIGKILL. The driver job has a one-hour ceiling. Very long runs can hit these limits and will be reported incomplete.
# Keep test VMs temporarily for inspection
npm run bench -- --keep
# Later, safely clean up that exact run (also retries failed cleanup)
npm run bench -- cleanup --run results/<your-run-directory>The resource manifest is private to your output directory. Cleanup verifies live VM metadata and route destinations; it never deletes every VM in an account. Do not edit/delete the manifest until cleanup succeeds. Retained VMs and endpoint credentials can be inspected through your own Freestyle account; treat CDP access as control of the disposable browser VM.
The tool reuses owned style.dev hostnames with no configured route, preferring
its own earlier browser-bench-* names. It never changes an active TLS route.
When none are available, it claims a new random hostname. Deleting the TLS route
does not release domain ownership: the API has no domain-release operation.
Those free claims remain owned and can be reused on later runs. If every claim
is actively routed and your domain quota is full, provisioning fails with cleanup.
Your Freestyle API key stays on your computer, is not written to reports, and
is not sent into a VM. CDP endpoints have independent random bearer tokens;
unauthenticated discovery must return 403 before timing starts. Only the
authenticated nginx endpoint is published; Chrome listens on guest loopback.
Remote worker credentials are mode-0600 files on disposable VMs; local worker
configuration is mode-0600 and removed after the run. Results and .env are ignored
by git. Don't publish your raw private output directory without reviewing it.
Chrome runs as an unprivileged guest user with --no-sandbox, using each disposable
VM as the isolation boundary. Packed mode does not provide a separate security
boundary between browser workers. This is a controlled demo against fixed public
pages, not a production multi-tenant browser service. Don't put application secrets
or untrusted customer workloads in these VMs. Restart-on-crash is disabled so the
runner does not silently hide browser restarts.
npm test
npm run check
npm run bench -- --helpCI runs unit/error-path tests without cloud credentials or billable VMs. Live validation is documented in TESTING.md. Public SDK: Freestyle quickstart.
MIT licensed. ComputeSDK-derived code retains its upstream license.