Repository navigation
docs: refresh Cloud verification guides - #49
Conversation
thisisjoshford
left a comment
There was a problem hiding this comment.
Tysm for the updates here @hanakannzashi ! Below is an agent review and agree we should resolve 1 & 2 ( having NVIDIA verification endpoint/example and including JS/Python snippets rather than just curl commands) - the rest are nice to haves IMO. I'm also working to update the verification example repo as mentioned in 3 but feel free to skip that part as we can do a follow-up PR.
-
The NVIDIA verification endpoint is gone. The old model page had the full
POST https://nras.attestation.nvidia.com/v3/attest/gpuexample with a sample EAT response. The new pages only say "verify with the applicable NVIDIA verification service" — a reader can no longer discover where to sendnvidia_payload. Suggest restoring the NRAS URL + minimal request/response example inreference/quote-nonce-signer.mdx(the "Verify GPU evidence" section is the natural home). -
All JS/Python examples were removed. The old pages had worked examples for attestation requests, request/response hashing (with a real fixture), and
ethers/eth_accountsignature recovery. The new pages are curl-only, and signature verification is prose ("recover the Ethereum EIP-191 signer") with no code. Since #44 asked for stale examples to be updated, dropping them entirely feels like the wrong direction — even one JS + one Python snippet per flow would do. -
Link to the verification example repo was dropped. The old model/chat pages pointed to
near-examples/nearai-cloud-verification-example(just updated for the current API, incl.signature_kindhandling); onlynearai-cloud-verifiersurvives, in one reference page. Please restore it alongside the verifier link — it's the easy on-ramp, the verifier is the complete implementation.
Non-blocking
-
signature_kindinference guidance now contradicts the API reference.cloud-api's OpenAPI text says legacy signatures "can still be distinguished structurally" (3 parts = model TEE, 2 = gateway);cloud-api/response-signatures.mdxsays "do not infer it from the shape oftext". The stricter stance is defensible, but the two sources should agree — worth a follow-up to align the API reference. -
cloud/verification/provenance.mdxis orphaned — not indocs.jsonnav and nothing links to it, and since the path didn't exist before there are no external inbound links either. Wire it in or drop it. -
Consider
docs.jsonredirectsinstead of stub pages for the oldchat/model/gateway/tlsURLs. The stubs work (hidden pages still resolve), but redirects would keep four near-empty pages out of search andllms.txt. -
The "why gateway signatures happen" explanation disappeared. The old chat page explained the gateway signs streamed responses because it rewrites chunks for usage accounting. Someone seeing
signature_kind: "gateway"on a streamed completion now has no "why" anywhere — one sentence incloud-api/response-signatures.mdxwould cover it.
Couldn't verify live (plausible, just noting): the error_code/message 200-unavailable signature shape, and the E2EE "unsupported rather than re-routed" behavior.
|
@hanakannzashi I think we can hold on the merge of this PR until nearai/inference-sdk#1 is merged and the verifiable-ai-sdk is released. We probably need to re-implement https://github.com/nearai/nearai-cloud-verifier after the verifiable-ai-sdk is published, and also update the NEAR AI docs accordingly. |
This reverts commit 965d0f5.
4123e07 to
59412ea
Compare
PierreLeGuen
left a comment
There was a problem hiding this comment.
Thanks @hanakannzashi, this is a big improvement. I ran the documented flows against production today: gateway and NEAR model reports, the direct endpoint, NRAS, same-connection TLS on both, and both signature helpers on real ecdsa and ed25519 signatures, streaming and non-streaming. Most of it works exactly as written. Four things should be fixed before merge. The inline suggestions cover all four and can be applied as one batch.
Must fix
- The E2EE examples return 400.
zai-org/GLM-5.1-FP8is now an alias ofz-ai/glm-5.3-flash, and withx-no-aliasing: truethe gateway rejects aliases. There are suggestions for all nine occurrences ine2ee-chat-completions.mdx. - MRCONFIGID is 48 bytes. Every quote I checked is
01 || SHA-256(app_compose) || 15 zero bytes. "Be01followed by that hash" makes an exact check reject valid reports. - RTMR3 replay. GLM-5.3 Flash reports return every RTMR3 event with an empty
digest, so replaying thedigestfields can't reproduce RTMR3. Recomputing each digest fromevent_type,eventandevent_payloadmatched every report I checked. I suggested a short section inquote-nonce-signer.mdxand links to it from the three pages. - Signer scope. Instances of the same model share one signing key, even when their compose hashes differ: three GLM-5.3 Flash configurations here report one signer. So matching a
provider_teesigner to a verified report doesn't show the completion came from that instance. The direct page already says this; the gateway pages imply the opposite ("Candidates are not interchangeable").
Josh's review
- NRAS (#1): done. The flow works verbatim; the verdict is
trueand the nonce matches. - JS/Python (#2): done for signatures. Both helpers verified all 20 of my test signatures. Attestation requests are still curl-only; linking the SDK guides once #51 lands would cover them.
- Example repo (#3): still missing, and the verifier link is gone too. Please restore both, and name a quote verifier again (the old page used
dcap-qvl). "An Intel DCAP verifier" doesn't tell a reader where to start. - Stub pages and redirects (#5, #6): done.
Follow-ups, fine after merge
- The signature covers the decoded body. With
Accept-Encoding: gzipthe gateway compresses, so hashing the received bytes fails. Keepidentity, but drop "retain the entity bytes it received". - Say when the gateway signs: default streams (it strips per-chunk usage), alias rewrites, and the Responses API (always). Also say how to get
provider_teeon a stream:stream_options.continuous_usage_stats: trueis relayed byte for byte. That fully covers Josh's #7. signing_algodefaults:/v1/signatureusesecdsa;/v1/attestation/reportusesed25519for the gateway report andecdsafor model candidates, so one response can mix both. The explicit-algorithm advice is right; say why.- The Node helper needs ESM (
"type": "module"or a.mjsfile) and ethers v6. The Python helper needs 3.9+. model_attestationsreturns at most one report today, and the field is omitted, not[], when there is none. Say "missing or empty".- Model reports never include
report_data. "Optional" can become "not returned". signing_public_keyis present forecdsatoo, and this PR documents ECDSA E2EE. Makeed25519a recommendation, not a requirement.- Give the
app_composehash command:jq -j(notjq -r, which adds a newline), hashing the UTF-8 bytes. - Report data: strip
0xfrom the ECDSA address. The TLS fingerprint is hashed as raw 32 bytes. - NRAS snippet: stop when the nonce check fails, and send the payload on stdin (
--data-binary @-). An 8-GPU payload is about 98 KB. - Say how to check debug mode (bit 0 of TD_ATTRIBUTES).
gateway-attestation.mdx:22: mentionx-no-aliasingwhen addingmodel=.- The signature reference page lists only the gateway page, but the direct pages send readers there too.
- Nothing links to the direct pages yet. #54 adds a visible Experimental page on top of this PR.
The build and link checks pass locally. Once 1–4 are in, I'll approve. @thisisjoshford, your blocking items are addressed apart from the example repo link. Could you take another look?
| 1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy. | ||
| 2. Read `report_data` and measurements from the **verified quote**, not only from fields echoed in the HTTP response. Reject the report if its echoed `request_nonce` or `report_data` is missing or differs from the nonce and report data you verified. | ||
| 3. Check that the verified quote binds the nonce generated by your client and the reported `signing_address`. Use `signing_algo` to interpret that identity and verify response signatures. | ||
| 4. Replay `event_log` and require the result to equal the RTMR3 in the verified quote. |
There was a problem hiding this comment.
Following the event-log link below: a plain digest replay doesn't reproduce RTMR3 for every report.
| 4. Replay `event_log` and require the result to equal the RTMR3 in the verified quote. | |
| 4. Replay `event_log` as described in [Replay the RTMR3 event log](/cloud/verification/reference/quote-nonce-signer#replay-the-rtmr3-event-log) and require the result to equal the RTMR3 in the verified quote. |
| 2. Read `report_data` and measurements from the **verified quote**, not only from fields echoed in the HTTP response. Reject the report if its echoed `request_nonce` or `report_data` is missing or differs from the nonce and report data you verified. | ||
| 3. Check that the verified quote binds the nonce generated by your client and the reported `signing_address`. Use `signing_algo` to interpret that identity and verify response signatures. | ||
| 4. Replay `event_log` and require the result to equal the RTMR3 in the verified quote. | ||
| 5. Obtain the raw `info.tcb_info.app_compose` string. If `tcb_info` is JSON text, decode it first. Hash `app_compose` without parsing or reserializing it, and require the quote's MRCONFIGID to be `01` followed by that SHA-256 hash. |
There was a problem hiding this comment.
MRCONFIGID is 48 bytes. Every quote I checked (both gateway instances and three GLM-5.3 Flash instances) is 01 || SHA-256(app_compose) || 15 zero bytes, so an exact comparison against 33 bytes rejects valid reports.
| 5. Obtain the raw `info.tcb_info.app_compose` string. If `tcb_info` is JSON text, decode it first. Hash `app_compose` without parsing or reserializing it, and require the quote's MRCONFIGID to be `01` followed by that SHA-256 hash. | |
| 5. Obtain the raw `info.tcb_info.app_compose` string. If `tcb_info` is JSON text, decode it first. Hash `app_compose` without parsing or reserializing it, and require the quote's 48-byte MRCONFIGID to equal `01`, then that SHA-256 hash, then 15 zero bytes. |
|
|
||
| 1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy. | ||
| 2. Read report data from the verified quote. Require it to bind the client-generated nonce and `signing_address`; require `request_nonce` to match the client nonce. When `report_data` is present, require it to match the verified quote. Interpret the address using `signing_algo`. | ||
| 3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote. |
There was a problem hiding this comment.
| 3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote. | |
| 3. Replay `event_log` as described in [Replay the RTMR3 event log](/cloud/verification/reference/quote-nonce-signer#replay-the-rtmr3-event-log) and require the resulting RTMR3 to equal the RTMR3 in the verified quote. |
| 1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy. | ||
| 2. Read report data from the verified quote. Require it to bind the client-generated nonce and `signing_address`; require `request_nonce` to match the client nonce. When `report_data` is present, require it to match the verified quote. Interpret the address using `signing_algo`. | ||
| 3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote. | ||
| 4. If `tcb_info` is JSON text, decode it to obtain `app_compose`. Hash the raw `app_compose` string without parsing or reserializing it. Require the quote's MRCONFIGID to begin with `01` followed by that SHA-256 hash. |
There was a problem hiding this comment.
Same 48-byte rule as the gateway page, so readers can do an exact comparison.
| 4. If `tcb_info` is JSON text, decode it to obtain `app_compose`. Hash the raw `app_compose` string without parsing or reserializing it. Require the quote's MRCONFIGID to begin with `01` followed by that SHA-256 hash. | |
| 4. If `tcb_info` is JSON text, decode it to obtain `app_compose`. Hash the raw `app_compose` string without parsing or reserializing it. Require the quote's 48-byte MRCONFIGID to equal `01`, then that SHA-256 hash, then 15 zero bytes. |
|
|
||
| 1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy. | ||
| 2. Check the client-generated nonce and signer identity against the verified quote. Read report data from the verified quote; require `request_nonce` to match the client nonce and, when `report_data` is present, require it to match the verified quote. Use `signing_algo` to interpret that identity and verify response signatures. | ||
| 3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote. |
There was a problem hiding this comment.
| 3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote. | |
| 3. Replay `event_log` as described in [Replay the RTMR3 event log](/cloud/verification/reference/quote-nonce-signer#replay-the-rtmr3-event-log) and require the resulting RTMR3 to equal the RTMR3 in the verified quote. |
| 'x-no-aliasing': 'true', | ||
| }, | ||
| body: JSON.stringify({ | ||
| model: 'zai-org/GLM-5.1-FP8', |
There was a problem hiding this comment.
| model: 'zai-org/GLM-5.1-FP8', | |
| model: 'z-ai/glm-5.3-flash', |
| NONCE="$(openssl rand -hex 32)" | ||
|
|
||
| curl --fail-with-body -G 'https://cloud-api.near.ai/v1/attestation/report' \ | ||
| --data-urlencode 'model=zai-org/GLM-5.1-FP8' \ |
There was a problem hiding this comment.
| --data-urlencode 'model=zai-org/GLM-5.1-FP8' \ | |
| --data-urlencode 'model=z-ai/glm-5.3-flash' \ |
| -H "X-Model-Pub-Key: MODEL_ECDSA_PUBLIC_KEY_HEX" \ | ||
| -H "x-no-aliasing: true" \ | ||
| -d '{ | ||
| "model": "zai-org/GLM-5.1-FP8", |
There was a problem hiding this comment.
| "model": "zai-org/GLM-5.1-FP8", | |
| "model": "z-ai/glm-5.3-flash", |
| After receiving a `provider_tee` response signature, select the one retained report whose signer identity and algorithm match the signature. If the match is absent or ambiguous, reject the response as unverified. | ||
|
|
||
| <Warning> | ||
| `model_attestations[]` is not an inventory of every model instance behind a URL. Candidates are not interchangeable: a response signature must be matched by signer identity and algorithm. |
There was a problem hiding this comment.
Instances of the same model share one signing key. Three GLM-5.3 Flash configurations with different compose hashes (the gateway's candidate, the direct endpoint and the long-context tier) all report the same signing_address. So "not interchangeable" reads as if each instance had its own signer, and a matching signer doesn't pin the completion to the verified instance. The direct page already says this (direct/response-signatures.mdx, "Match the model signer").
| `model_attestations[]` is not an inventory of every model instance behind a URL. Candidates are not interchangeable: a response signature must be matched by signer identity and algorithm. | |
| `model_attestations[]` is not an inventory of every model instance behind a URL. Instances of the same model share one signing key, including instances with different measured configurations, so a matching signer does not establish that the report and the completion came from the same serving instance. |
| | `provider_tee` | `<MODEL_ID>:<request_hash>:<response_hash>` | Exactly one verified NEAR model report with the same `signing_address` and `signing_algo`. | | ||
| | `gateway` | `<request_hash>:<response_hash>` | A verified `gateway_attestation` with the same `signing_address` and `signing_algo`. | | ||
|
|
||
| Only `provider_tee` binds a response to a model-serving TEE. A `gateway` signature binds the response to the gateway, not to a model report. If a model signature is unavailable, model-level response binding is unavailable. |
There was a problem hiding this comment.
Same signer-scope caveat as on the model-attestations page.
| Only `provider_tee` binds a response to a model-serving TEE. A `gateway` signature binds the response to the gateway, not to a model report. If a model signature is unavailable, model-level response binding is unavailable. | |
| Only `provider_tee` binds a response to a model-serving TEE. A `gateway` signature binds the response to the gateway, not to a model report. If a model signature is unavailable, model-level response binding is unavailable. Instances of the same model share one signing key, so a matching `provider_tee` signer does not establish that the verified report and the completion came from the same serving instance. |
- Use the canonical z-ai/glm-5.3-flash in the E2EE examples. The gateway rejects the old alias when x-no-aliasing is set. - State the full 48-byte MRCONFIGID and document RTMR3 replay, including reports whose event digests are empty. - Note that instances of a model share one signing key, so a matching signer does not pin the serving instance. - Add Python and Node.js helpers that verify the quote with dcap-qvl and check debug mode, nonce, signer, MRCONFIGID and RTMR3. - Restore links to the verification example and verifier, and name dcap-qvl for quote verification. - Stop the NRAS snippet on a nonce mismatch and send the payload on stdin. - Clarify the signed bytes (uncompressed body), when the gateway signs, signing_algo defaults, helper runtimes, and missing model_attestations.
|
@hanakannzashi I pushed 34af011 on top of your branch to get this ready to merge before your break. What changed:
Tested against production today:
Heads-up: with an @thisisjoshford your items are addressed now. Could you take another look? |
PierreLeGuen
left a comment
There was a problem hiding this comment.
Approving with the fixes in 34af011. The link check passes, and the updated flows were tested against production as described in my comment above.
Summary
Validation
git diff --check origin/main...HEADjq empty docs.jsonCloses #44