Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Vertex2OpenAI on Cloudflare Workers

v2.1.1 — OpenAI-compatible Chat Completions and Responses for Vertex AI Gemini, with no runtime dependencies. Use your own adapter key in clients; the Worker manages Vertex credentials and token refresh.

Deploy to Cloudflare

What changed since v1.0?

These are cumulative v2.x changes; v1.0 already supported Chat Completions, streaming, tools, reasoning, search, images and both credential types.

Area v1.0 v2.1.1
APIs Chat Completions and model listing Adds /v1/responses, including typed streaming events, tools and reasoning
Models Gemini 2.x and 3.x Gemini 3+ only; configurable model list and capability-aware variants
Credentials Express keys and service accounts Rotation and failover across both credential pools; 429, 5xx and network failures can try the next credential
Tool calls and schemas Basic conversion Gemini 3 thought-signature round trips; native JSON Schema support for $ref, $defs and nullable unions
Images Image variants Always uses native Vertex, including with service accounts; removes unsupported image/search combinations
Streaming Basic SSE translation Linear SSE parsing, correct usage and finish states; truncated, malformed and failed streams no longer look successful
Request handling Basic validation developer instructions on both routes, invalid body shapes return 400, unique Responses/item IDs
Verification No automated suite in the v1.0 tag 284 unit tests, coverage gates, CPU benchmarks and real Vertex E2E tests

New in v2.1.1: fixes Responses error/truncation handling and output snapshots, malformed request handling, developer messages and ID collisions. Responses streaming skips an intermediate SSE encode/parse step: about 18% less CPU in the local pipeline benchmark. Full changes.

Quick start

Requires Node.js 20+ and npm; use Node 22.18+ for test coverage checks.

git clone https://github.com/workHMZ/vertex2openai-cf.git
cd vertex2openai-cf
npm ci
npm run verify
npx wrangler secret put API_KEY                  # strong, random adapter key
npx wrangler secret put GOOGLE_CREDENTIALS_JSON  # service account JSON on one line
npx wrangler deploy

For Express authentication, set VERTEX_EXPRESS_API_KEY instead of GOOGLE_CREDENTIALS_JSON. The interactive npm run deploy helper can configure the adapter key and Express key, then deploy.

Client setting Value
Base URL https://vertex2openai.<subdomain>.workers.dev/v1
API key Your API_KEY
Model For example, gemini-3.8-flash; query /v1/models for configured variants
curl https://your-worker.workers.dev/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"gemini-3.8-flash","messages":[{"role":"user","content":"Hi"}]}'

To upgrade an existing installation: git pull --ff-only, npm ci, npm run verify, then npx wrangler deploy. No new secrets or migrations are required from v2.1.0. From v1.0, replace Gemini 2.x IDs and review the limits below.

Workers Free plan: read before deploying

Text often fits; images do not reliably fit, even at the default size. CPU time measures the Worker's processing, not time spent waiting for Vertex. A request taking several minutes can still use only a few milliseconds of CPU.

Limit Workers Free What it means here
CPU 10 ms per request Some streamed text and cold service-account authentication exceed it; sustained overruns can terminate requests with Error 1102
Requests 100,000/day per account, resets at 00:00 UTC Shared across Workers on the account
Memory 128 MB per isolate Concurrent requests share it; large base64 images and repeated snapshots increase pressure
Subrequests 50 per request Upstream attempts and authentication requests consume this budget
Variable/secret size 5 KB each Multiple service-account JSON keys must fit the total serialized size; do not assume a fixed key count

Limits checked against Cloudflare's official documentation on 2026-09-27. Workers Paid defaults to 30 seconds of CPU per request (configurable up to 5 minutes); this is separate from Vertex billing and quotas.

Actual Cloudflare CPU, measured with wrangler tail on v2.1.1, 2026-09-27, using the production service-account route:

Successful request Samples CPU range Median
Chat text, non-streaming 17 1–7 ms 1 ms
Chat text, streaming 3 3–10 ms 5 ms
Responses text, non-streaming 5 1–4 ms 2 ms
Responses text, streaming 3 5–17 ms 5 ms
Flash-Lite image, Chat non-streaming 1 4 ms 4 ms
Flash-Lite image, Chat streaming 1 11 ms 11 ms
Flash-Lite image, Responses streaming 1 14 ms 14 ms

All 26 production E2E scenarios passed, including images after retrying upstream 429s. Image response sizes were not recorded; these single Flash-Lite samples cannot be compared directly with the older 2.9 MB image measurements (~36/100 ms non-streaming/streaming). The local 18% optimization is not a measured production speedup.

Use Workers Paid and non-streaming requests for image generation. A previous 4K sample produced a ~54 MB HTTP response; Paid does not raise the isolate's 128 MB memory limit or fix Vertex 429 errors. Successful requests above Free's CPU budget are not a reliability guarantee. Historical measurements and methodology are preserved.

Configuration and API

Set secrets with wrangler secret put; API_KEY plus at least one Vertex credential is required.

Setting Purpose
API_KEY Protects the adapter and access to your Vertex billing account
GOOGLE_CREDENTIALS_JSON One service-account JSON object, or comma-separated objects
VERTEX_EXPRESS_API_KEY / VERTEX_API_KEY Express API key(s), comma-separated; the second name is an alias
GCP_LOCATION Plain variable in wrangler.toml; defaults to global
GCP_PROJECT_ID Optional project override; otherwise read from the service-account key
MODELS_CONFIG Optional JSON with vertex_models and vertex_express_models arrays; defaults to src/models.json

GET / is an unauthenticated health check. /v1/models, /v1/chat/completions and /v1/responses require Authorization: Bearer <API_KEY>.

Model option Behavior
[EXPRESS] / [PAY] prefix Pin Express / service-account credentials; without a prefix, prefer Express and fall back to service accounts
-search Google Search grounding on text models
-nothinking / -max Lowest / highest supported thinking level
-2k / -4k Image resolution; image models use only size variants
-openai / -openaisearch Force the OpenAI-compatible endpoint; service-account text models only
Feature Chat Completions Responses
Input messages input, with optional instructions
Output limit max_tokens max_output_tokens
Thinking reasoning_effort reasoning.effort
JSON Schema response_format text.format
Tools tool_calls / tool messages function_call / function_call_output items
Reasoning output reasoning_content reasoning items

Thinking accepts none, minimal, low, medium, high, xhigh and max; the last two map to high. Model-specific adjustments live in src/model-capabilities.ts: for example, gemini-3.8-flash maps minimal thinking to low and drops unsupported frequency/presence penalties.

Compatibility notes

  • Credentials: ordinary, unbound GCP API keys are not Express keys. Use an Express signup key, a key bound to a service account with Vertex permissions, or service-account JSON. Keep keys out of Git and use a strong adapter key.
  • Tools: preserve thought_signature on returned calls for Gemini 3 reasoning continuity. Incoming thoughtSignature and extra_content.google.thought_signature also work. If a client strips signatures, the adapter supplies Vertex's validation placeholder.
  • Responses is stateless: send the full history in input. previous_response_id and conversation return 400; stored responses, GET/DELETE by ID and OpenAI-hosted tools are unsupported. Use -search for grounding.
  • Stream endings: token limits/content filtering produce response.incomplete; upstream errors, malformed frames and missing finish reasons produce response.failed, retaining partial output. Chat reports upstream failures as an error frame before [DONE]. Clients should check terminal status.
  • Timeouts and billing: under throttling, successful streams historically began after 36, 143 or 184 seconds (normally 2–4 seconds). The adapter waits up to 300 seconds per credential; allow a longer client read timeout if using failover. Thinking tokens count toward billed output and are included in completion-token usage.
  • Safety: the adapter currently sends BLOCK_NONE for the four configured harm categories; review buildSafetySettings for your application.

Tests and development

Validation, 2026-09-27 Result Elapsed time
npm run verify Typecheck + 284/284 unit tests; coverage 99.84% lines / 93.21% branches / 100% functions 0.69 s
Production Cloudflare, complete configured SA route 26/26 passed, no failures/cancellations/skips; four 429 responses recovered through retries 194.98 s (3m 14.98s)
Linux workerd + real Vertex, earlier full E2E 25/26 passed, no timeouts/skips; SA image request returned upstream 429 331.84 s (5m 31.84s)
Earlier final affected E2E rerun Responses 6/6 passed; SA image still 429 183.55 s (3m 3.55s)
Earlier direct Vertex check, bypassing Worker Same SA image request returned 429 RESOURCE_EXHAUSTED 0.19 s

E2E counts are scenarios, often containing several requests across both credential types. The earlier Linux run exercised both credential types; production has only SA credentials and now passes all three image modes. Express image generation passed in the earlier isolated tests. Detailed timings, local CPU results and historical test data include the scope of each measurement.

npm run dev       # local Worker; secrets in .dev.vars, restart after editing them
npm run verify    # typecheck, all unit tests and coverage gates
npm run bench     # local synthetic CPU benchmark, no real Vertex calls

Real E2E calls incur Vertex charges. See test instructions for credentials, route enforcement and the optional image skip.

Inspired by gzzhongqi/vertex2openai. MIT License.

About

OpenAI-compatible Cloudflare Workers proxy for Google Vertex AI Gemini.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages