Skip to content
Open
44 changes: 44 additions & 0 deletions .github/workflows/devlog-probes.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
name: devlog probes

# Runs the offline unit probes that live beside a devlog plan. The suite exists so
# an exact-head gate actually executes the research fixtures a docs PR cites
# instead of leaving them as locally-verified-only claims.

permissions:
contents: read

on:
pull_request:
paths:
- "devlog/**/probes/**"
- ".github/workflows/devlog-probes.yml"

concurrency:
group: devlog-probes-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true

jobs:
unittest:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
Comment thread
coderabbitai[bot] marked this conversation as resolved.
with:
persist-credentials: false
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.x"
- name: aiohttp for socket fixtures (required; the socket suite imports it unconditionally)
run: python -m pip install aiohttp
- name: Run every discovered probe suite
shell: bash
run: |
set -euo pipefail
found=0
while IFS= read -r -d '' dir; do
found=1
echo "== probes: $dir"
(cd "$dir" && python -m unittest discover -s . -p 'test_*.py' -v)
done < <(find devlog -type d -name probes -print0)
if [[ "$found" -eq 0 ]]; then
echo "no devlog probe directories present"
fi
32 changes: 32 additions & 0 deletions devlog/_plan/260928_remote_thread_provider_policy/000_plan.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# 5848: native remote-list provider policy before a backend relay

Status: **proposal and executable specifications only; no production fix**.
Date: 2026-09-28.

Related: [#5848](https://github.com/lidge-jun/opencodex/issues/5848),
[duplicate #5906](https://github.com/lidge-jun/opencodex/issues/5906),
[upstream #48358](https://github.com/openai/codex/issues/48358).

## Proposed decision

Prefer a narrowly scoped, operator-opt-in policy in **native Codex app-server**
for remote `thread/list` requests that do not supply a provider array. Keep the
normal default-provider behavior, all existing authorization, and history intact.
The mobile client's explicit all-provider request remains the simplest upstream
client correction. Compare both with the previously tested local backend relay;
do not turn that experimental relay into an installed feature in this PR.

The current OpenCodex runtime and [ADR-5848](../../../structure/decisions/ADR-5848-provider-table-remote-history-visibility.md)
remain unchanged. This proposal does **not** close #5848.

## Work remaining

- [x] Identify the native remote-connection boundary and existing provider predicate.
- [x] Specify precedence and test the native-policy proposal offline.
- [x] Retain and rerun the isolated loopback relay alternative.
- [ ] Obtain upstream agreement on the native configuration/API contract.
- [ ] Implement config/schema, trusted-origin plumbing, and Rust regressions upstream.
- [ ] Verify native pagination, managed policy, account changes, and actual mobile resume.
- [ ] Establish released-version/capability detection before adding any OpenCodex toggle.

See [design](010_design.md) and [verification](020_verification.md).
134 changes: 134 additions & 0 deletions devlog/_plan/260928_remote_thread_provider_policy/010_design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,134 @@
# Design comparison and proposed native contract

This is an RFC, not an accepted architecture change or a released setting.
`probes/native_policy.py` is an executable specification, **not native Rust code**.
No probe is imported by OpenCodex or registered in its runtime/CLI/test suite.

## Evidence and scope

The source baseline is OpenAI Codex
`1cc7e2361237ce7244430ee1d581c77f95c57ac8` and OpenCodex `dev`
`eb7f0f0970c2298f8b2d66d170c4d4be869f301b`.

- [Native list predicate](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/app-server/src/request_processors/thread_processor.rs#L5456-L5530): explicit nonempty arrays filter to those ids; `[]` removes the provider predicate; omission defaults to the configured provider, except for related-thread queries.
- [Remote connection origin](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/app-server-transport/src/transport/remote_control/client_tracker.rs): native remote sessions open with `ConnectionOrigin::RemoteControl`; messages become native transport events.
- [Original rationale, #5658](https://github.com/openai/codex/pull/5658): provider annotations/filtering address cross-provider resume/decryption failures. Visibility is not proof that a conversation is safe to resume under a different provider.
- [OpenCodex mitigation, #6007](https://github.com/lidge-jun/opencodex/pull/6007): warnings and documentation without rewriting native history or intercepting native RPC.

## Alternatives

| Approach | Can leave official mobile app unchanged? | Additional ownership | Recommendation |
| --- | --- | --- | --- |
| Mobile explicitly sends `modelProviders: []` (or a chosen list) | No | Client's list behavior and resume UX | Simplest client correction; continue upstream tracking. |
| Native, opt-in, remote-only default list policy | Yes, once a compatible native build ships | Native config and request policy, no extra service | **Preferred implementation proposal** when the host needs explicit control. |
| Local relay selected through `chatgpt_base_url` | Potentially; live service not verified | Shared backend traffic, token forwarding, connection lifecycle | Research fallback, not a default or currently supported remedy. |
| Globally change omitted filter to all providers | Yes | Changes behavior for every caller | Do not use: loses the existing default isolation behavior. |
| Rewrite `openai` history tags to `opencodex` | Yes in some observations | Native SQLite/rollout state | Do not use: violates the retained history-writer boundary. |

The native proposal minimizes new connection and credential handling; this is an
engineering recommendation, not evidence of upstream acceptance or deployment.
Neither upstream route can be delivered by modifying OpenCodex's `/v1` inference
proxy alone. This PR publishes the comparison rather than disguising a probe as a fix.

## Proposed native setting (name subject to upstream review)

**Illustrative only; do not add this to current Codex/OpenCodex configuration.**

```toml
[remote_control]
thread_list_model_providers = ["openai", "opencodex"]
```

Absent setting means opt-out. An explicitly empty array means all providers;
a nonempty array means exactly those ids. Do not automatically infer equivalence
from provider names, model names, or a shared URL. The setting changes listing,
not authentication, permission checks, routing, compaction, or resume semantics.
An operator who only needs two provider ids need not opt into every provider.

### Precedence

1. Existing authentication and managed remote-control policy still run first.
2. A request's explicit provider **array**, including `[]`, always wins.
3. With no array, parent/ancestor queries retain their current no-default-filter behavior.
4. Only a server-identified remote connection may use the operator's configured policy.
5. With no policy, and for every non-remote connection, preserve the existing default.

Typed Rust `Option<Vec<String>>` treats omission and JSON `null` alike. This
proposal deliberately applies its default to both. The raw relay probe preserves
all present JSON keys, including null; that difference is documented and must not
be mistaken for identical behavior between the two alternatives.

Use the native `ConnectionOrigin` attached to the connection. Never infer remote
origin from clientInfo.name, a user-agent, JSON fields, or a caller-supplied header.
The Python enum tests only model this trusted input; they do not prove authentication.

### Native implementation boundary

The upstream change must add and validate the config contract, snapshot it with
the native app-server's effective configuration, carry the trusted connection
origin to the `thread/list` handler, and resolve the provider predicate before
calling the existing store pagination. A new RPC layer is unnecessary.

Conceptual selection, **not a drop-in patch**:

```text
if request supplies provider array: use that array
else if parent/ancestor query: use existing related-thread behavior
else if trusted origin is RemoteControl and operator policy is configured:
use the configured array
else: use existing configured-default behavior
```

Do not add a global exception to `list_threads_common` for all callers. Do not
change `thread/start`, `thread/resume`, `turn/start`, stored provider metadata,
or request error handling. Keep source/cwd/archive/project filters and ordering.
The new configuration's trust/precedence should prevent project content from
silently opting an operator into a broader default; upstream must choose and test
the appropriate configuration layers and any managed restrictions.

Hold one effective policy throughout a listing/pagination sequence. A changed
policy requires a fresh listing cursor; do not combine pages obtained under
incompatible filters. Native tests must pin behavior with actual cursor semantics.
The fixture tests here use integer pagination, not native opaque cursors.

## Local relay alternative: what was and was not shown

[Native startup](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/app-server/src/lib.rs)
uses `config.chatgpt_base_url` for remote control. The
[URL protocol](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/app-server-transport/src/transport/remote_control/protocol.rs)
accepts loopback URLs. A host-side relay is therefore a concrete design alternative:

```text
unchanged mobile <-> existing ChatGPT backend <-> local relay <-> native app-server
```

Connection establishment starts on the host. Enroll/pair/refresh can be forwarded
rather than reimplemented. However, `chatgpt_base_url` is shared with
[authentication configuration](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/core/src/config/auth_keyring.rs)
and other backend consumers. A WebSocket-only implementation is insufficient.
[Enrollment persistence](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/app-server-transport/src/transport/remote_control/enroll.rs)
also keys state by URL/account/client, so changing the URL can affect enrollment
selection; whether live re-pairing is required remains unverified.

The retained relay fixture accepts only literal `127.0.0.1` HTTP upstreams and
synthetic credentials. It is not a production service and cannot be configured
for ChatGPT. It handles regular and single-chunk frames; multi-chunk messages are
passed unchanged. Actual protocol-v3 segmentation, native reconnect/ACK/cursors,
account changes, managed network policy, and full backend compatibility remain
release gates, not claims supported by these tests. See native
[WebSocket handling](https://github.com/openai/codex/blob/1cc7e2361237ce7244430ee1d581c77f95c57ac8/codex-rs/app-server-transport/src/transport/remote_control/websocket.rs).

No global base-URL mutation, token-pool integration, production listener,
configuration migration, history write, or authentication bypass is proposed here.

## Upstream and downstream follow-through

An upstream implementation needs Rust config/schema, origin-scoping, explicit/null/
related-query, pagination, and managed-policy tests. Keep existing defaults intact.
Document the omitted/default behavior and regenerate affected schema/TS fixtures.

After an accepted implementation is released, OpenCodex may consider an opt-in
integration with positive capability/version evidence and precise configuration
ownership/restoration. Merely writing an unknown TOML key is not a feature test.
Until then, retain #6007's warnings and keep #5848 open. Do not suppress the warning
because an experimental setting was written or a fixture passed.
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
# Verification

Date: 2026-09-28. Environment: Python 3.13.5, aiohttp 3.13.3, Linux.
All inputs, accounts, tokens, thread ids, and database rows are synthetic.
No native Codex process, user home, installed config, or real service was accessed.

From the repository root:

```sh
cd devlog/_plan/260928_remote_thread_provider_policy/probes
python -m unittest -v test_native_policy test_probe test_loopback_bridge
```

**Inventory: 61 tests across three suites.**

- 23 tests specify the proposed native policy: opt-out, trusted-origin scoping,
explicit arrays, null semantics, parent/ancestor exceptions, immutable policy,
exact ids, UUID thread-id validation, and synthetic paginated fixture reads.
- 29 retained raw-frame tests cover the relay alternative's ordinary/single-chunk
rewrites, byte-identical pass-through, limits, and unsupported inputs.
- 9 retained localhost HTTP/WebSocket tests use a mock host and mock backend,
checking two-way traffic, fixture auth/headers, explicit filters, enrollment,
unrelated HTTP, and fixture endpoint restrictions.

The 23 + 29 offline tests use the Python standard library only. The 9 socket tests
need aiohttp; no production dependency manifest is changed.

Author-local execution on 2026-09-28 covered the 52 standard-library tests
(test_native_policy + test_probe); the socket suite was not part of that recorded
run. Hosted exact-head execution is authoritative for the remaining 9: the
devlog-probes workflow added by this PR installs aiohttp deterministically (the
step fails the job if install fails) and discovers every devlog/**/probes
directory. On head 787ef30698 the hosted unittest job ran all 61 tests with 0
failures.

```sh
python -m unittest -v test_native_policy test_probe
```

The synthetic database has 5,200 `openai` rows, one `opencodex` row, and one `other`
row. Default filtering yields one; the two-id opt-in yields 5,201; all providers
yields 5,202. Page sizes 1, 37, and 1,000 are exercised. The fixture dump hash stays
unchanged after reads. This demonstrates the specification, not native SQLite
schema compatibility or actual mobile results.

## Not executed / not established

- Native Rust implementation, compilation, config/schema generation, or native tests.
- Actual ChatGPT mobile pairing, full pagination, resume, token renewal, or reconnect.
- OpenCodex Bun typecheck, full suite, structure gate, or repository privacy gate:
Bun and a full checkout are unavailable in the execution environment. A git
checkout attempt failed at DNS resolution. Focused Python tests are not a
substitute for required repository gates; the PR must remain draft.
- Production multi-segment relay support or management/security review.

The PR adds this research unit plus one workflow that runs it: devlog-probes.yml
is a new CI lane, so the executable contract now executes on exact head rather
than only locally. No executable configuration, release artifact, or native
storage is changed. Independent review remains outstanding even with green CI.
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
"""Executable specification for a PROPOSED native Codex policy, not an integration.

No socket, authentication, config, history, or rollout access. The real Rust
implementation must obtain origin from its authenticated server-side connection,
not from JSON, clientInfo.name, or an HTTP header. This enum only models that input.
None as the result means no provider predicate, not no results.
"""
from __future__ import annotations

from dataclasses import dataclass
from enum import Enum
from uuid import UUID
import re

# Rust's Uuid::parse_str accepts exactly four textual shapes - hyphenated, simple,
# braced-hyphenated and the urn:uuid prefix. Python's UUID() is looser: it tolerates
# arbitrary hyphen placement inside a 32-hex payload, so wrapper/hyphen variants
# must be rejected by shape before parsing instead of leaning on the constructor.
_UUID_ACCEPTED_SHAPES = re.compile(
r"(?:"
r"[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}"
r"|[0-9a-fA-F]{32}"
r"|\{[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}\}"
r"|urn:uuid:[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}"
r")\Z"
)


class ConnectionOrigin(Enum):
STDIO = "stdio"
WEBSOCKET = "websocket"
IN_PROCESS = "in_process"
REMOTE_CONTROL = "remote_control"


def _validate_ids(value: tuple[str, ...], label: str) -> None:
if not isinstance(value, tuple):
raise ValueError(f"{label} must be a tuple")
if any(not isinstance(item, str) or not item.strip() for item in value):
raise ValueError(f"{label} must contain non-empty strings")
if len(set(value)) != len(value):
raise ValueError(f"{label} must not contain duplicate identifiers")


def _validate_thread_id(value: str, label: str) -> None:
"""Upstream parses ThreadId as a UUID; a non-UUID string must not widen the list."""
if not isinstance(value, str) or not _UUID_ACCEPTED_SHAPES.match(value):
raise ValueError(f"{label} must be a UUID thread id")
try:
UUID(value)
except (ValueError, AttributeError, TypeError):
raise ValueError(f"{label} must be a UUID thread id")


@dataclass(frozen=True)
class RemoteListPolicy:
# None: opt-out; (): explicitly all; non-empty: exactly these provider ids.
providers: tuple[str, ...] | None = None

def __post_init__(self) -> None:
if self.providers is not None:
_validate_ids(self.providers, "providers")


def resolve_provider_filter(
*,
origin: ConnectionOrigin,
default_provider: str,
requested: tuple[str, ...] | None = None,
parent_thread_id: str | None = None,
ancestor_thread_id: str | None = None,
policy: RemoteListPolicy = RemoteListPolicy(),
) -> tuple[str, ...] | None:
"""Model precedence after typed request decoding and existing authorization.

Rust Option<Vec<String>> decodes both omission and JSON null as None. They
intentionally have the same semantics in this native-policy proposal. The
separate raw-frame relay probe instead preserves every present JSON key.
Empty requested tuple is an explicit client request for all providers.
The policy applies only to thread/list; callers must not reuse it for resume.
"""
if not isinstance(origin, ConnectionOrigin):
raise ValueError("origin must be supplied by the trusted connection context")
_validate_ids((default_provider,), "default_provider")
for label, tid in (
("parent_thread_id", parent_thread_id),
("ancestor_thread_id", ancestor_thread_id),
):
if tid is not None:
_validate_thread_id(tid, label)
if parent_thread_id is not None and ancestor_thread_id is not None:
raise ValueError("parent_thread_id and ancestor_thread_id are mutually exclusive")
if requested is not None:
# Preserve every typed client array, including duplicate/empty ids.
# Native Vec<String> accepts them; this proposal must not add rejection.
if not isinstance(requested, tuple) or any(not isinstance(item, str) for item in requested):
raise ValueError("requested must model a typed string array")
return requested or None
if parent_thread_id is not None or ancestor_thread_id is not None:
return None
if origin is ConnectionOrigin.REMOTE_CONTROL and policy.providers is not None:
return policy.providers or None
return (default_provider,)
Loading
Loading