Skip to content

docs: refresh Cloud verification guides - #49

Merged
PierreLeGuen merged 36 commits into
mainfrom
codex/refresh-verification-docs
Sep 23, 2026
Merged

PierreLeGuen merged 36 commits into
mainfrom
codex/refresh-verification-docs

Conversation

@hanakannzashi

@hanakannzashi hanakannzashi commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Reorganize Cloud verification around model, gateway, TLS, response signatures, and image provenance.
  • Make direct model endpoints experimental, remove them from primary navigation and onboarding, and clarify that their evidence is scoped to individual requests.
  • Remove remaining direct-endpoint discovery from Private Inference, OpenCode, and the integration-guide authoring template.
  • Document client nonces, signer matching for load-balanced endpoints, and same-connection TLS SPKI checks.
  • Add a fail-closed image-provenance guide, including clear handling for missing records and HTTP 404s.
  • Align related private-inference copy with the verification guides.

Validation

  • git diff --check origin/main...HEAD
  • jq empty docs.json
  • Rendered the Gateway, experimental direct, Quickstart, E2EE, and OpenCode guides locally with Mint

Closes #44

@thisisjoshford thisisjoshford left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tysm for the updates here @hanakannzashi ! Below is an agent review and agree we should resolve 1 & 2 ( having NVIDIA verification endpoint/example and including JS/Python snippets rather than just curl commands) - the rest are nice to haves IMO. I'm also working to update the verification example repo as mentioned in 3 but feel free to skip that part as we can do a follow-up PR.


  1. The NVIDIA verification endpoint is gone. The old model page had the full POST https://nras.attestation.nvidia.com/v3/attest/gpu example with a sample EAT response. The new pages only say "verify with the applicable NVIDIA verification service" — a reader can no longer discover where to send nvidia_payload. Suggest restoring the NRAS URL + minimal request/response example in reference/quote-nonce-signer.mdx (the "Verify GPU evidence" section is the natural home).

  2. All JS/Python examples were removed. The old pages had worked examples for attestation requests, request/response hashing (with a real fixture), and ethers/eth_account signature recovery. The new pages are curl-only, and signature verification is prose ("recover the Ethereum EIP-191 signer") with no code. Since #44 asked for stale examples to be updated, dropping them entirely feels like the wrong direction — even one JS + one Python snippet per flow would do.

  3. Link to the verification example repo was dropped. The old model/chat pages pointed to near-examples/nearai-cloud-verification-example (just updated for the current API, incl. signature_kind handling); only nearai-cloud-verifier survives, in one reference page. Please restore it alongside the verifier link — it's the easy on-ramp, the verifier is the complete implementation.

Non-blocking

  1. signature_kind inference guidance now contradicts the API reference. cloud-api's OpenAPI text says legacy signatures "can still be distinguished structurally" (3 parts = model TEE, 2 = gateway); cloud-api/response-signatures.mdx says "do not infer it from the shape of text". The stricter stance is defensible, but the two sources should agree — worth a follow-up to align the API reference.

  2. cloud/verification/provenance.mdx is orphaned — not in docs.json nav and nothing links to it, and since the path didn't exist before there are no external inbound links either. Wire it in or drop it.

  3. Consider docs.json redirects instead of stub pages for the old chat/model/gateway/tls URLs. The stubs work (hidden pages still resolve), but redirects would keep four near-empty pages out of search and llms.txt.

  4. The "why gateway signatures happen" explanation disappeared. The old chat page explained the gateway signs streamed responses because it rewrites chunks for usage accounting. Someone seeing signature_kind: "gateway" on a streamed completion now has no "why" anywhere — one sentence in cloud-api/response-signatures.mdx would cover it.

Couldn't verify live (plausible, just noting): the error_code/message 200-unavailable signature shape, and the E2EE "unsupported rather than re-routed" behavior.

@think-in-universe

Copy link
Copy Markdown
Contributor

@hanakannzashi I think we can hold on the merge of this PR until nearai/inference-sdk#1 is merged and the verifiable-ai-sdk is released.

We probably need to re-implement https://github.com/nearai/nearai-cloud-verifier after the verifiable-ai-sdk is published, and also update the NEAR AI docs accordingly.

@PierreLeGuen PierreLeGuen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @hanakannzashi, this is a big improvement. I ran the documented flows against production today: gateway and NEAR model reports, the direct endpoint, NRAS, same-connection TLS on both, and both signature helpers on real ecdsa and ed25519 signatures, streaming and non-streaming. Most of it works exactly as written. Four things should be fixed before merge. The inline suggestions cover all four and can be applied as one batch.

Must fix

  1. The E2EE examples return 400. zai-org/GLM-5.1-FP8 is now an alias of z-ai/glm-5.3-flash, and with x-no-aliasing: true the gateway rejects aliases. There are suggestions for all nine occurrences in e2ee-chat-completions.mdx.
  2. MRCONFIGID is 48 bytes. Every quote I checked is 01 || SHA-256(app_compose) || 15 zero bytes. "Be 01 followed by that hash" makes an exact check reject valid reports.
  3. RTMR3 replay. GLM-5.3 Flash reports return every RTMR3 event with an empty digest, so replaying the digest fields can't reproduce RTMR3. Recomputing each digest from event_type, event and event_payload matched every report I checked. I suggested a short section in quote-nonce-signer.mdx and links to it from the three pages.
  4. Signer scope. Instances of the same model share one signing key, even when their compose hashes differ: three GLM-5.3 Flash configurations here report one signer. So matching a provider_tee signer to a verified report doesn't show the completion came from that instance. The direct page already says this; the gateway pages imply the opposite ("Candidates are not interchangeable").

Josh's review

  • NRAS (#1): done. The flow works verbatim; the verdict is true and the nonce matches.
  • JS/Python (#2): done for signatures. Both helpers verified all 20 of my test signatures. Attestation requests are still curl-only; linking the SDK guides once #51 lands would cover them.
  • Example repo (#3): still missing, and the verifier link is gone too. Please restore both, and name a quote verifier again (the old page used dcap-qvl). "An Intel DCAP verifier" doesn't tell a reader where to start.
  • Stub pages and redirects (#5, #6): done.

Follow-ups, fine after merge

  • The signature covers the decoded body. With Accept-Encoding: gzip the gateway compresses, so hashing the received bytes fails. Keep identity, but drop "retain the entity bytes it received".
  • Say when the gateway signs: default streams (it strips per-chunk usage), alias rewrites, and the Responses API (always). Also say how to get provider_tee on a stream: stream_options.continuous_usage_stats: true is relayed byte for byte. That fully covers Josh's #7.
  • signing_algo defaults: /v1/signature uses ecdsa; /v1/attestation/report uses ed25519 for the gateway report and ecdsa for model candidates, so one response can mix both. The explicit-algorithm advice is right; say why.
  • The Node helper needs ESM ("type": "module" or a .mjs file) and ethers v6. The Python helper needs 3.9+.
  • model_attestations returns at most one report today, and the field is omitted, not [], when there is none. Say "missing or empty".
  • Model reports never include report_data. "Optional" can become "not returned".
  • signing_public_key is present for ecdsa too, and this PR documents ECDSA E2EE. Make ed25519 a recommendation, not a requirement.
  • Give the app_compose hash command: jq -j (not jq -r, which adds a newline), hashing the UTF-8 bytes.
  • Report data: strip 0x from the ECDSA address. The TLS fingerprint is hashed as raw 32 bytes.
  • NRAS snippet: stop when the nonce check fails, and send the payload on stdin (--data-binary @-). An 8-GPU payload is about 98 KB.
  • Say how to check debug mode (bit 0 of TD_ATTRIBUTES).
  • gateway-attestation.mdx:22: mention x-no-aliasing when adding model=.
  • The signature reference page lists only the gateway page, but the direct pages send readers there too.
  • Nothing links to the direct pages yet. #54 adds a visible Experimental page on top of this PR.

The build and link checks pass locally. Once 1–4 are in, I'll approve. @thisisjoshford, your blocking items are addressed apart from the example repo link. Could you take another look?

1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy.
2. Read `report_data` and measurements from the **verified quote**, not only from fields echoed in the HTTP response. Reject the report if its echoed `request_nonce` or `report_data` is missing or differs from the nonce and report data you verified.
3. Check that the verified quote binds the nonce generated by your client and the reported `signing_address`. Use `signing_algo` to interpret that identity and verify response signatures.
4. Replay `event_log` and require the result to equal the RTMR3 in the verified quote.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Following the event-log link below: a plain digest replay doesn't reproduce RTMR3 for every report.

Suggested change
4. Replay `event_log` and require the result to equal the RTMR3 in the verified quote.
4. Replay `event_log` as described in [Replay the RTMR3 event log](/cloud/verification/reference/quote-nonce-signer#replay-the-rtmr3-event-log) and require the result to equal the RTMR3 in the verified quote.

2. Read `report_data` and measurements from the **verified quote**, not only from fields echoed in the HTTP response. Reject the report if its echoed `request_nonce` or `report_data` is missing or differs from the nonce and report data you verified.
3. Check that the verified quote binds the nonce generated by your client and the reported `signing_address`. Use `signing_algo` to interpret that identity and verify response signatures.
4. Replay `event_log` and require the result to equal the RTMR3 in the verified quote.
5. Obtain the raw `info.tcb_info.app_compose` string. If `tcb_info` is JSON text, decode it first. Hash `app_compose` without parsing or reserializing it, and require the quote's MRCONFIGID to be `01` followed by that SHA-256 hash.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MRCONFIGID is 48 bytes. Every quote I checked (both gateway instances and three GLM-5.3 Flash instances) is 01 || SHA-256(app_compose) || 15 zero bytes, so an exact comparison against 33 bytes rejects valid reports.

Suggested change
5. Obtain the raw `info.tcb_info.app_compose` string. If `tcb_info` is JSON text, decode it first. Hash `app_compose` without parsing or reserializing it, and require the quote's MRCONFIGID to be `01` followed by that SHA-256 hash.
5. Obtain the raw `info.tcb_info.app_compose` string. If `tcb_info` is JSON text, decode it first. Hash `app_compose` without parsing or reserializing it, and require the quote's 48-byte MRCONFIGID to equal `01`, then that SHA-256 hash, then 15 zero bytes.


1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy.
2. Read report data from the verified quote. Require it to bind the client-generated nonce and `signing_address`; require `request_nonce` to match the client nonce. When `report_data` is present, require it to match the verified quote. Interpret the address using `signing_algo`.
3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote.
3. Replay `event_log` as described in [Replay the RTMR3 event log](/cloud/verification/reference/quote-nonce-signer#replay-the-rtmr3-event-log) and require the resulting RTMR3 to equal the RTMR3 in the verified quote.

1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy.
2. Read report data from the verified quote. Require it to bind the client-generated nonce and `signing_address`; require `request_nonce` to match the client nonce. When `report_data` is present, require it to match the verified quote. Interpret the address using `signing_algo`.
3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote.
4. If `tcb_info` is JSON text, decode it to obtain `app_compose`. Hash the raw `app_compose` string without parsing or reserializing it. Require the quote's MRCONFIGID to begin with `01` followed by that SHA-256 hash.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same 48-byte rule as the gateway page, so readers can do an exact comparison.

Suggested change
4. If `tcb_info` is JSON text, decode it to obtain `app_compose`. Hash the raw `app_compose` string without parsing or reserializing it. Require the quote's MRCONFIGID to begin with `01` followed by that SHA-256 hash.
4. If `tcb_info` is JSON text, decode it to obtain `app_compose`. Hash the raw `app_compose` string without parsing or reserializing it. Require the quote's 48-byte MRCONFIGID to equal `01`, then that SHA-256 hash, then 15 zero bytes.


1. Verify `intel_quote` with an Intel DCAP verifier and apply your TCB and advisory policy.
2. Check the client-generated nonce and signer identity against the verified quote. Read report data from the verified quote; require `request_nonce` to match the client nonce and, when `report_data` is present, require it to match the verified quote. Use `signing_algo` to interpret that identity and verify response signatures.
3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
3. Replay `event_log` and require the resulting RTMR3 to equal the RTMR3 in the verified quote.
3. Replay `event_log` as described in [Replay the RTMR3 event log](/cloud/verification/reference/quote-nonce-signer#replay-the-rtmr3-event-log) and require the resulting RTMR3 to equal the RTMR3 in the verified quote.

Comment thread cloud/guides/e2ee-chat-completions.mdx Outdated
'x-no-aliasing': 'true',
},
body: JSON.stringify({
model: 'zai-org/GLM-5.1-FP8',

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
model: 'zai-org/GLM-5.1-FP8',
model: 'z-ai/glm-5.3-flash',

Comment thread cloud/guides/e2ee-chat-completions.mdx Outdated
NONCE="$(openssl rand -hex 32)"

curl --fail-with-body -G 'https://cloud-api.near.ai/v1/attestation/report' \
--data-urlencode 'model=zai-org/GLM-5.1-FP8' \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
--data-urlencode 'model=zai-org/GLM-5.1-FP8' \
--data-urlencode 'model=z-ai/glm-5.3-flash' \

Comment thread cloud/guides/e2ee-chat-completions.mdx Outdated
-H "X-Model-Pub-Key: MODEL_ECDSA_PUBLIC_KEY_HEX" \
-H "x-no-aliasing: true" \
-d '{
"model": "zai-org/GLM-5.1-FP8",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"model": "zai-org/GLM-5.1-FP8",
"model": "z-ai/glm-5.3-flash",

After receiving a `provider_tee` response signature, select the one retained report whose signer identity and algorithm match the signature. If the match is absent or ambiguous, reject the response as unverified.

<Warning>
`model_attestations[]` is not an inventory of every model instance behind a URL. Candidates are not interchangeable: a response signature must be matched by signer identity and algorithm.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Instances of the same model share one signing key. Three GLM-5.3 Flash configurations with different compose hashes (the gateway's candidate, the direct endpoint and the long-context tier) all report the same signing_address. So "not interchangeable" reads as if each instance had its own signer, and a matching signer doesn't pin the completion to the verified instance. The direct page already says this (direct/response-signatures.mdx, "Match the model signer").

Suggested change
`model_attestations[]` is not an inventory of every model instance behind a URL. Candidates are not interchangeable: a response signature must be matched by signer identity and algorithm.
`model_attestations[]` is not an inventory of every model instance behind a URL. Instances of the same model share one signing key, including instances with different measured configurations, so a matching signer does not establish that the report and the completion came from the same serving instance.

| `provider_tee` | `<MODEL_ID>:<request_hash>:<response_hash>` | Exactly one verified NEAR model report with the same `signing_address` and `signing_algo`. |
| `gateway` | `<request_hash>:<response_hash>` | A verified `gateway_attestation` with the same `signing_address` and `signing_algo`. |

Only `provider_tee` binds a response to a model-serving TEE. A `gateway` signature binds the response to the gateway, not to a model report. If a model signature is unavailable, model-level response binding is unavailable.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same signer-scope caveat as on the model-attestations page.

Suggested change
Only `provider_tee` binds a response to a model-serving TEE. A `gateway` signature binds the response to the gateway, not to a model report. If a model signature is unavailable, model-level response binding is unavailable.
Only `provider_tee` binds a response to a model-serving TEE. A `gateway` signature binds the response to the gateway, not to a model report. If a model signature is unavailable, model-level response binding is unavailable. Instances of the same model share one signing key, so a matching `provider_tee` signer does not establish that the verified report and the completion came from the same serving instance.

- Use the canonical z-ai/glm-5.3-flash in the E2EE examples. The gateway
  rejects the old alias when x-no-aliasing is set.
- State the full 48-byte MRCONFIGID and document RTMR3 replay, including
  reports whose event digests are empty.
- Note that instances of a model share one signing key, so a matching
  signer does not pin the serving instance.
- Add Python and Node.js helpers that verify the quote with dcap-qvl and
  check debug mode, nonce, signer, MRCONFIGID and RTMR3.
- Restore links to the verification example and verifier, and name
  dcap-qvl for quote verification.
- Stop the NRAS snippet on a nonce mismatch and send the payload on stdin.
- Clarify the signed bytes (uncompressed body), when the gateway signs,
  signing_algo defaults, helper runtimes, and missing model_attestations.
@PierreLeGuen

Copy link
Copy Markdown
Contributor

@hanakannzashi I pushed 34af011 on top of your branch to get this ready to merge before your break.

What changed:

  • The four must-fix items from my review: the canonical model ID in the E2EE examples, the 48-byte MRCONFIGID, a new RTMR3 replay section that handles empty digests, and the shared-signer caveat on the gateway pages.
  • Josh's feat: enhance private inference #2 and fix: implement feedback #3: Python and Node.js helpers in reference/quote-nonce-signer.mdx that verify the quote with dcap-qvl and check debug mode, nonce, signer, MRCONFIGID and RTMR3. The verification example and verifier links are back on the overview page, and the quote step names dcap-qvl.
  • The quick follow-ups:
    • The NRAS snippet stops on a nonce mismatch and sends the payload on stdin.
    • jq -j for the compose hash.
    • The signed bytes are the uncompressed body.
    • When the gateway signs, and the signing_algo defaults.
    • Helper runtimes (ESM and ethers v6 for Node, Python 3.9+).
    • Missing model_attestations[], and report_data not being returned.
    • signing_public_key for both algorithms, and the debug-mode bit.
    • The direct page link on the signature reference page.

Tested against production today:

  • Both helpers pass on gateway and model reports (ecdsa and ed25519, with and without TLS binding) and on the direct endpoint with same-connection TLS.
  • They reject a wrong nonce, a wrong peer certificate, a changed app_compose and a changed event.
  • The E2EE quick start works end to end with z-ai/glm-5.3-flash, and the NRAS snippet returns a passing verdict.
  • mint validate and mint broken-links pass.

Heads-up: with an UpToDate-only policy, the helpers reject the gateway quote today because its TCB status is OutOfDate. That comes from the gateway hosts, not from the docs.

@thisisjoshford your items are addressed now. Could you take another look?

@PierreLeGuen PierreLeGuen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving with the fixes in 34af011. The link check passes, and the updated flows were tested against production as described in my comment above.

@PierreLeGuen
PierreLeGuen merged commit af705e1 into main Sep 23, 2026
2 checks passed
@PierreLeGuen
PierreLeGuen deleted the codex/refresh-verification-docs branch September 23, 2026 17:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Refresh outdated Cloud verification docs

4 participants