From d3658de1a7c6f5b32fd213a9761d355e3cf20d2f Mon Sep 17 00:00:00 2001 From: AtmegaBuzz Date: Sat, 26 Sep 2026 18:27:30 +0530 Subject: [PATCH 1/9] Remove docs/ folder; fix now-dangling references to it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit docs/persona/{architecture,A2A,GROUPS,theme,AUDIT_REPORT, IMPLEMENTATION_MEMORY_BRIEF,linkedin-portability-pre-apply,seo-plan}.md are gone from the working tree — still recoverable from history (git log --all --oneline -- '**/architecture.md') since they came in with the persona import. services/persona-api/CLAUDE.md still said 'full design docs live at the repo root' and named architecture.md/A2A.md/GROUPS.md/theme.md as if they were there — reworded so it doesn't point at files that no longer exist. Same for the docs/persona row and further-reading links in AGENTS.md and README.md. Co-Authored-By: Claude Sonnet 5 --- AGENTS.md | 1 - README.md | 5 - docs/persona/A2A.md | 740 ------------------ docs/persona/AUDIT_REPORT.md | 208 ----- docs/persona/GROUPS.md | 252 ------ docs/persona/IMPLEMENTATION_MEMORY_BRIEF.md | 685 ---------------- docs/persona/architecture.md | 375 --------- .../persona/linkedin-portability-pre-apply.md | 270 ------- docs/persona/seo-plan.md | 102 --- docs/persona/theme.md | 700 ----------------- services/persona-api/CLAUDE.md | 18 +- 11 files changed, 9 insertions(+), 3347 deletions(-) delete mode 100644 docs/persona/A2A.md delete mode 100644 docs/persona/AUDIT_REPORT.md delete mode 100644 docs/persona/GROUPS.md delete mode 100644 docs/persona/IMPLEMENTATION_MEMORY_BRIEF.md delete mode 100644 docs/persona/architecture.md delete mode 100644 docs/persona/linkedin-portability-pre-apply.md delete mode 100644 docs/persona/seo-plan.md delete mode 100644 docs/persona/theme.md diff --git a/AGENTS.md b/AGENTS.md index fe7f97d..0315978 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,6 @@ If a task mentions either of those by name, it's the wrong repo. | `services/memory` | Shared context/memory layer: ingest, matching, MCP server, OAuth for AI clients | FastAPI / Python | api.zynd.ai | | `infra/persona-box` | pm2 configs for the box running persona-api + persona-web | — | — | | `infra/api-box` | Caddy + docker-compose for the box running cards-api + memory | — | — | -| `docs/persona` | Persona's architecture docs (identity derivation, A2A protocol, groups rollout) | — | — | | `packages/` | Shared code (DB migrations, contracts). **Planned, not built yet.** | — | — | **History is preserved.** Each service was merged in with `git filter-repo`, diff --git a/README.md b/README.md index 04502b7..403e47c 100644 --- a/README.md +++ b/README.md @@ -25,7 +25,6 @@ against multiple server versions). | `services/memory` | Shared context layer — ingest, matching, MCP server, OAuth for ChatGPT/Claude/Cursor | FastAPI / Python | https://api.zynd.ai | | `infra/persona-box` | pm2 process configs for the persona server | — | — | | `infra/api-box` | Caddy + Docker Compose for the cards/memory server | — | — | -| `docs/persona` | Persona's architecture docs (identity, A2A protocol, groups) | — | — | | `packages/` | Shared DB migrations / API contracts. **Planned, not built yet.** | — | — | Each service came from its own repo (`agent-persona`, `zynd-cards`, @@ -148,9 +147,5 @@ side effect of an unrelated change: ## Further reading - [`AGENTS.md`](./AGENTS.md) — full rules for any AI agent working here. -- [`docs/persona/architecture.md`](./docs/persona/architecture.md) — persona - identity model, heartbeat design, security model. -- [`docs/persona/A2A.md`](./docs/persona/A2A.md) — the agent-to-agent - protocol. - Per-service `CLAUDE.md`/`AGENTS.md`/`README.md` inside `apps/*` and `services/*` — stack-specific conventions. diff --git a/docs/persona/A2A.md b/docs/persona/A2A.md deleted file mode 100644 index 8d9e250..0000000 --- a/docs/persona/A2A.md +++ /dev/null @@ -1,740 +0,0 @@ -# Agent-to-Agent Communication Architecture (Zynd Personas, v3) - -**Status:** Design — source of truth for implementation -**Scope:** The persona-to-persona (agent-to-agent) channel inside `agent-persona`. Human-to-human and human-to-agent flows are out of scope and are not changed. -**Standard:** A2A protocol v0.3 over JSON-RPC 2.0, with the `x-zynd-auth` per-message Ed25519 envelope defined in `zyndai-ts-sdk/src/a2a/` and `zyndai-agent/zyndai_agent/a2a/`. - ---- - -## 1. System Overview - -### 1.1 What we are designing - -A formal, deterministic, fail-safe protocol that lets one persona's AI agent talk to another persona's AI agent on behalf of two human principals — for example, when Alice tells her agent "set up a 30-minute intro call with Bob next week" and her agent has to coordinate with Bob's agent until either a meeting ticket lands or one side declines. - -The replacement covers everything inside the `agent-persona` repo that participates in cross-agent traffic — the webhook ingress, the orchestrator's external mode, the reply path, the connection lifecycle, and the per-thread permission model. It does **not** replace human-typed conversation, the persona registration flow, identity derivation, the heartbeat manager, MCP tool implementations, or the OAuth subsystem. - -### 1.2 Goals - -1. **Continuous, conversational A2A** — agents must be able to hold multi-turn negotiations across minutes or hours without losing state, without polling, and without race windows where a reply is received but never surfaced. -2. **Deterministic state** — every situation an agent can be in maps to exactly one named state. Every input maps to exactly one transition. There are no implicit timeouts, no implicit retries, no "if it didn't reply within X seconds, hope for the best" code paths. -3. **Cryptographic accountability** — every cross-agent message is Ed25519-signed via `x-zynd-auth`. The receiver always knows which entity is talking and can prove it later. No network peer can forge a sender, replay a message, or send something past its expiry window. -4. **Permission enforcement at three layers** — agent card (advertised), connection (`dm_threads.permissions`), and orchestrator tool allowlist. A foreign agent cannot trigger an action the principal hasn't explicitly granted, even if the LLM hallucinates the call. -5. **Resilience** — transient transport failures must be invisible above the protocol layer. Permanent failures must terminate cleanly with a recorded cause that the principal can see. -6. **Compatibility with the SDKs** — agent-persona becomes a first-class A2A peer that any other agent built on `zyndai-ts-sdk` or `zyndai-agent` can talk to without special-casing. - -### 1.3 Key differences from the current implementation - -| Concern | Today (legacy) | New design (this doc) | -|---|---|---| -| Wire format | Custom `AgentMessage` JSON POSTed to `/api/persona/webhooks/{user_id}` and `/api/persona/webhooks/{user_id}/sync` | JSON-RPC 2.0 over `POST /a2a/v1` (the spec endpoint) | -| Discovery | Webhook URL stored in DB / on registry card | A2A `/.well-known/agent-card.json` published per persona, signed with the persona's Ed25519 key | -| Authentication | None at message level (relies on registry pre-knowledge) | Per-message Ed25519 `x-zynd-auth` envelope with nonce, expiry, replay cache | -| Conversation identifier | `dm_threads.id` only | `dm_threads.id` IS the A2A `contextId`. Each request inside the connection gets its own A2A `taskId` | -| Reply discovery | Sender polls `dm_messages` for ~60s, gives up if nothing lands | Receiver pushes terminal/interrupted state back via the caller-supplied `pushNotificationConfig`. UI updates via Supabase realtime on the persistence tables | -| Multi-turn flow | Each message is independent; no formal "still working", no formal "I need more info" | Task FSM with `submitted → working → input-required / auth-required → completed / canceled / failed / rejected`. The "I need more info" state is first-class | -| Connection state | Three values: `pending`, `accepted`, `blocked` | Explicit FSM (`none → requested → accepted | declined | blocked → revoked`) with documented transitions | -| Permissions | Four boolean flags consumed only by the orchestrator's allowlist | Same four flags, but enforced at three layers (card advertisement, transport-level rejection, orchestrator allowlist) and revalidated per task | -| State on failure | Silent: caller sees "no reply yet", receiver may have crashed mid-flight | Every task ends in a named terminal state; failures persist a reason string the principal can read | -| Concurrency | Two webhooks (async + sync) coexist; sender can't tell which the receiver will run | One endpoint, two transports negotiated via the agent card (`JSONRPC` for sync, `message/stream` SSE for streaming). Choice belongs to the caller | -| Replay protection | None | LRU nonce cache per `entity_id` with skew-window TTL; replays are rejected with `ZYND_REPLAY_DETECTED` | - ---- - -## 2. Core Concepts - -### 2.1 Agent - -A persona deployed by a human principal. Identified by `agent_id` (currently `zns:<32 hex>`), backed by an Ed25519 keypair derived from the developer key at `derivation_index`. One agent per principal. The agent is the only thing that participates in A2A traffic; the human never does directly. - -### 2.2 Identity - -Already in place and not redesigned here. Three properties matter for A2A: - -* **`entity_id`** — what the receiver sees on every message. Carried in `x-zynd-auth.entity_id`. -* **`public_key`** — Ed25519 pubkey, base64 (`ed25519:...`). Carried in `x-zynd-auth.public_key`. Hashed to verify against `entity_id`. -* **`developer_proof`** (optional but always sent on first contact) — proves the agent key was HD-derived from a known developer key. Lets the receiver check provenance without an Agent DNS lookup. - -### 2.3 Connection - -The long-lived relationship between two personas. Stored on `dm_threads`. A connection is the unit of permission: trust, allowed tools, takeover state, and history all hang off it. A connection has its own state machine — see §3.1. - -### 2.4 Session (A2A `contextId`) - -A connection's stable identifier as far as the protocol is concerned. **`dm_threads.id` IS the `contextId`.** Every cross-agent message on this connection carries the same `contextId`. The `contextId` is what makes a multi-task negotiation feel like one conversation — both sides can correlate "this task is part of the meeting we've been planning" without inspecting message content. - -### 2.5 Task (A2A `taskId`) - -A single discrete request inside a session. Examples: "ask Bob's agent for availability next week", "propose meeting at T", "ack a counter". A task has its own FSM (§3.2). A task is the unit of: - -* state (where in the FSM it is now) -* history (the messages exchanged for this task) -* artifacts (the structured outputs produced by the receiver — e.g. a meeting ticket id) -* timeouts (idle TTL, terminal retention) -* push delivery (one `pushNotificationConfig` per task) - -A session typically holds many tasks over its lifetime. Each task is independent — failure of one task never propagates to siblings. - -### 2.6 Permissions - -Four boolean flags persisted in `dm_threads.permissions` JSONB, owned by the receiver: - -* `can_request_meetings` (default ON) -* `can_query_availability` (default OFF) -* `can_view_full_profile` (default OFF) -* `can_post_on_my_behalf` (default OFF) - -Permissions live on the connection, not on the task — but each task is gated by them at dispatch time. Changing a permission affects every future task on that connection immediately. In-flight tasks are not retroactively cancelled (they ran with the permissions captured when the task was opened). - -### 2.7 Capabilities - -Two layers: - -* **Agent capabilities** — the persona's `capabilities` array on `persona_agents`. What the agent advertises it can do. Public, on the agent card. Used for discovery only. -* **Tool allowlist** — the actual MCP tools available to the orchestrator on a given task. Computed from `EXTERNAL_DEFAULT_ALLOWED ∪ permission_gates(thread_permissions)`. Computed deterministically per task; never mutated mid-task. - -### 2.8 Message types and primitives - -The wire types are exactly the A2A spec's. We expose four primitive operations as JSON-RPC methods on `/a2a/v1`: - -| Method | Direction | Purpose | -|---|---|---| -| `message/send` | caller → callee | Submit one message. Returns the final Task once it settles (terminal or interrupted). Synchronous from the caller's POV but the callee may take seconds. | -| `message/stream` | caller → callee | Same as send, but the response is an SSE stream of `status-update` and `artifact-update` events until the task is final. | -| `tasks/get` | caller → callee | Read-only Task lookup by id. | -| `tasks/cancel` | caller → callee | Force a task to `canceled`. Only legal on non-terminal tasks. | -| `tasks/resubscribe` | caller → callee | Re-attach an SSE stream to a task already in flight (e.g. after a reconnect). | -| `tasks/pushNotificationConfig/set` | caller → callee | Register or update the URL/token to which the callee will POST terminal-state updates. | -| `tasks/pushNotificationConfig/get` | caller → callee | Read back the current push config. | - -A `Message` itself contains an array of typed `Part`s: - -* `TextPart` — natural-language content (what the LLM produced). -* `DataPart` — structured payload (e.g. `{ proposed_time: "...", title: "..." }`). -* `FilePart` — inline bytes or remote URI. - -This means a single message can carry the LLM's prose and a structured proposal at the same time — no parsing of free text on the receiver side. The orchestrator's tool calls produce `DataPart`s; the orchestrator's natural-language wrap-up produces a `TextPart`. - -### 2.9 Push notifications - -A receiver-driven async delivery primitive. The caller registers a webhook URL (along with an optional bearer token) with the callee; when the task hits a terminal or interrupted state, the callee POSTs a signed wrapper Message containing a `TaskStatusUpdateEvent` to that URL. agent-persona will expose `POST /api/persona/push/{user_id}` for this. Push delivery never replaces the SSE stream or `tasks/get` — it's a wakeup signal so the caller can fetch and surface the result immediately. - ---- - -## 3. Formal State Machine - -There are **two layered FSMs**: the **ConnectionFSM** governs whether two personas can talk at all and on what terms; the **TaskFSM** governs each individual cross-agent request once a connection permits it. The TaskFSM is the A2A spec's lifecycle (we keep it byte-for-byte to stay interoperable). The ConnectionFSM is ours. - -### 3.1 ConnectionFSM - -#### 3.1.1 States - -| State | Meaning | Stored as | -|---|---|---| -| `none` | No `dm_threads` row exists between these two agents. | (absence of row) | -| `requested` | Initiator created the row; receiver hasn't responded yet. No tasks may be opened against this connection. | `dm_threads.status = 'pending'` | -| `accepted` | Both sides agree to communicate. Tasks may be opened. Permissions are honored as configured. | `dm_threads.status = 'accepted'` | -| `declined` | Receiver rejected the request. Connection is dormant; no tasks may be opened. Initiator may not auto-retry. | `dm_threads.status = 'declined'` (new value — see §10.1 schema delta) | -| `blocked` | Receiver actively blocked the initiator. Inbound A2A traffic from this peer is rejected at the transport layer with `ZYND_AUTH_FAILED` reason `untrusted_sender`. | `dm_threads.status = 'blocked'` | -| `revoked` | Either side ended the connection after acceptance. No new tasks may be opened. Existing terminal tasks remain readable for audit. | `dm_threads.status = 'revoked'` (new value) | - -Every connection at every moment is in exactly one of the six states. There is no implicit, derived, or "in between" state — `pending` and `accepted` and the rest are mutually exclusive and exhaustive. - -#### 3.1.2 Events that drive transitions - -| Event | Origin | Notes | -|---|---|---| -| `EV_REQUEST_OPEN` | Initiator | Initiator creates a connection request. | -| `EV_REQUEST_ACCEPT` | Receiver | Receiver accepts via UI. | -| `EV_REQUEST_DECLINE` | Receiver | Receiver declines via UI. | -| `EV_BLOCK` | Receiver | Receiver blocks via UI. Allowed from any non-terminal state. | -| `EV_UNBLOCK` | Receiver | Receiver unblocks via UI. Connection returns to whatever it was before; if it was never accepted, it returns to `requested` only if the original request is still ≤ TTL_REQUEST (30 days). Otherwise it returns to `none`. | -| `EV_REVOKE` | Either side | Either party walks away from an accepted connection. | -| `EV_INBOUND_MESSAGE` | Network | Used to validate transport-level admission, not to change connection state. | - -#### 3.1.3 Transition table - -Every cell is filled in. `—` means the event is illegal in that state and the system MUST reject it (HTTP 4xx for UI-driven events; transport rejection for network events). - -| | EV_REQUEST_OPEN | EV_REQUEST_ACCEPT | EV_REQUEST_DECLINE | EV_BLOCK | EV_UNBLOCK | EV_REVOKE | -|---|---|---|---|---|---|---| -| **none** | → `requested` | — | — | — | — | — | -| **requested** | — (idempotent: stays `requested`, no row duplication) | → `accepted` | → `declined` | → `blocked` (preserves prior status) | — | — | -| **accepted** | — (idempotent) | — (idempotent) | — | → `blocked` (preserves prior status) | — | → `revoked` | -| **declined** | → `requested` (only if ≥ TTL_DECLINE_COOLDOWN since `declined_at`; else stays `declined`) | — | — | → `blocked` | — | — | -| **blocked** | — (rejected at transport) | — | — | — (idempotent) | → restored prior status (or `none` if expired) | — | -| **revoked** | → `requested` (always allowed; treats the new request as fresh) | — | — | → `blocked` | — | — (idempotent) | - -`TTL_REQUEST = 30 days`, `TTL_DECLINE_COOLDOWN = 7 days`. Both are documented invariants, not magic numbers buried in code. - -#### 3.1.4 Diagram (Mermaid) - -```mermaid -stateDiagram-v2 - [*] --> none - - none --> requested: EV_REQUEST_OPEN - - requested --> accepted: EV_REQUEST_ACCEPT - requested --> declined: EV_REQUEST_DECLINE - requested --> blocked: EV_BLOCK - - accepted --> revoked: EV_REVOKE - accepted --> blocked: EV_BLOCK - - declined --> requested: EV_REQUEST_OPEN (after cooldown) - declined --> blocked: EV_BLOCK - - blocked --> requested: EV_UNBLOCK (if prior was requested & not expired) - blocked --> accepted: EV_UNBLOCK (if prior was accepted) - blocked --> none: EV_UNBLOCK (if request expired) - - revoked --> requested: EV_REQUEST_OPEN - revoked --> blocked: EV_BLOCK -``` - -#### 3.1.5 Determinism and absence of deadlock - -* For every (state, event) pair the table specifies exactly one outcome (a transition, an idempotent stay, or a rejection). No event is ambiguous. -* The set of events that change state is finite and human-driven (UI buttons or API calls). There is no cycle that lacks a human exit. From any state, the user can always reach `blocked` or `revoked`, both of which terminate the conversation cleanly. -* `requested → accepted` is the only path that opens task-level traffic. There is no shortcut. - -### 3.2 TaskFSM - -This is the A2A v0.3 spec FSM exactly. We do not redefine it — we adopt it. Documenting it here explicitly so the rest of the system has a single reference. - -#### 3.2.1 States - -| State | Meaning | Terminal? | -|---|---|---| -| `submitted` | Task entry created, message queued, handler not yet picked up. | No | -| `working` | Handler is running. | No | -| `input-required` | Handler suspended; needs the caller to send another message on this `taskId`. | No (interrupted) | -| `auth-required` | Handler suspended; needs the caller to provide an auth credential before it can continue. | No (interrupted) | -| `completed` | Handler returned a result successfully. Final artifact attached. | Yes | -| `canceled` | Either side called `tasks/cancel`. | Yes | -| `failed` | Handler raised, or the receiver hit a hard error (validation, internal, push delivery exhausted). | Yes | -| `rejected` | Receiver refused at dispatch (no handler registered, payload validation failed, permissions denied). | Yes | - -`TERMINAL_STATES = { completed, canceled, failed, rejected }`. `INTERRUPTED_STATES = { input-required, auth-required }`. - -#### 3.2.2 Events - -| Event | Origin | -|---|---| -| `EV_T_NEW_MESSAGE` | First `message/send` for a new `taskId` | -| `EV_T_RESUME_MESSAGE` | `message/send` for an existing `taskId` currently in an interrupted state | -| `EV_T_HANDLER_STARTS` | Internal, when the handler picks up the message | -| `EV_T_HANDLER_ASKS` | Handler called `task.ask(...)` | -| `EV_T_HANDLER_REQUIRES_AUTH` | Handler called `task.requireAuth(...)` | -| `EV_T_HANDLER_COMPLETES` | Handler returned / called `task.complete(...)` | -| `EV_T_HANDLER_FAILS` | Handler threw or called `task.fail(...)` | -| `EV_T_HANDLER_REJECTS` | Validation failed before dispatch, or no handler | -| `EV_T_CANCEL` | `tasks/cancel` from either side | -| `EV_T_IDLE_TIMEOUT` | Task in interrupted state past `IDLE_TTL_INTERRUPTED` | -| `EV_T_PERMISSION_REVOKED` | Connection moved to `blocked` / `revoked` while task was in flight | - -#### 3.2.3 Transition table (per A2A spec, with our additions) - -| | EV_T_NEW_MESSAGE | EV_T_RESUME_MESSAGE | EV_T_HANDLER_STARTS | EV_T_HANDLER_ASKS | EV_T_HANDLER_REQUIRES_AUTH | EV_T_HANDLER_COMPLETES | EV_T_HANDLER_FAILS | EV_T_HANDLER_REJECTS | EV_T_CANCEL | EV_T_IDLE_TIMEOUT | EV_T_PERMISSION_REVOKED | -|---|---|---|---|---|---|---|---|---|---|---|---| -| **(no row)** | → `submitted` | — (rejected: A2A_TASK_NOT_FOUND) | — | — | — | — | — | — | — | — | — | -| **submitted** | — (idempotent) | — (rejected: not interrupted) | → `working` | — | — | — | — | → `rejected` | → `canceled` | — | → `canceled` | -| **working** | — (rejected: task busy) | — (rejected: not interrupted) | — | → `input-required` | → `auth-required` | → `completed` | → `failed` | → `rejected` | → `canceled` | — | → `canceled` | -| **input-required** | — (must reuse `taskId` ⇒ EV_T_RESUME_MESSAGE) | → `working` | — | — | — | — | — | — | → `canceled` | → `failed` (reason: idle) | → `canceled` | -| **auth-required** | — | → `working` | — | — | — | — | — | — | → `canceled` | → `failed` (reason: idle) | → `canceled` | -| **completed** / **canceled** / **failed** / **rejected** | rejects with A2A_TASK_NOT_CANCELABLE on cancel; new messages with the same taskId after retention window land in (no row) and start fresh | — | — | — | — | — | — | — | — | — | — | - -#### 3.2.4 Diagram (Mermaid) - -```mermaid -stateDiagram-v2 - [*] --> submitted: EV_T_NEW_MESSAGE - - submitted --> working: handler starts - submitted --> rejected: validation / no handler - submitted --> canceled: tasks/cancel - - working --> input_required: handler asks - working --> auth_required: handler requires auth - working --> completed: handler returns - working --> failed: handler throws - working --> rejected: invalid output - working --> canceled: tasks/cancel - working --> canceled: connection blocked or revoked - - input_required --> working: caller resumes - input_required --> failed: idle TTL - input_required --> canceled: tasks/cancel - - auth_required --> working: caller resumes with auth - auth_required --> failed: idle TTL - auth_required --> canceled: tasks/cancel - - completed --> [*] - canceled --> [*] - failed --> [*] - rejected --> [*] -``` - -#### 3.2.5 Determinism, completeness, freedom from deadlock - -* The transition table covers every (state, event) pair. Cells marked `—` are explicit rejection responses, not undefined behavior. -* From any non-terminal state, at least three exit paths to terminal states exist (`EV_T_CANCEL`, `EV_T_IDLE_TIMEOUT` for interrupted states, `EV_T_HANDLER_FAILS` while working). No state can sit forever. -* `submitted` cannot stall on the dispatcher: the server thread that creates a `submitted` row also schedules the handler before returning. If the dispatcher itself crashes, the GC sweeper (§3.2.6) eventually transitions the row to `failed`. -* Cancellation is idempotent on terminal states: receiving `tasks/cancel` for a completed task returns `A2A_TASK_NOT_CANCELABLE`, never silently drops. - -#### 3.2.6 Idle GC - -The Task Store runs a periodic sweeper: - -* Tasks in `INTERRUPTED_STATES` past `IDLE_TTL_INTERRUPTED` (default **1 hour**, configurable per agent) move to `failed` with reason `"Task timed out after Ns of inactivity"`. Any suspended handler is unblocked with an internal abort sentinel. -* Tasks in `TERMINAL_STATES` are retained for `TERMINAL_RETENTION` (default **5 minutes**) so callers can `tasks/get` the result, then deleted from the in-memory store. Persistent task records (see §4.4) keep them indefinitely. - -GC events drive transitions through `EV_T_IDLE_TIMEOUT`. They never produce any state not already in §3.2.1. - ---- - -## 4. Message Protocol - -### 4.1 Transport - -* **Endpoint:** `POST {persona_base_url}/a2a/v1`. Mounted on the existing FastAPI app per persona, dispatched by the same `user_id` path scheme already in use, e.g. `https://your-server.com/api/persona/a2a/{user_id}/v1`. The exact mount path is recorded on each agent card under `url`. -* **Discovery:** `GET {persona_base_url}/.well-known/agent-card.json` (same per-persona scheme). The card carries the canonical `url`, `preferredTransport`, `capabilities`, advertised input/output schemas, and the `x-zynd` extension block (entity_id, public_key, fqan, registry, status, developerProof). The card is signed with the persona's Ed25519 key per A2A's detached-JWS scheme. -* **Transports advertised on the card:** `JSONRPC` as primary; `HTTP+JSON` not advertised (keeps the wire single-shape). Streaming and push notifications are advertised as capabilities. - -### 4.2 Envelope - -Every cross-agent call is JSON-RPC 2.0: - -``` -{ - "jsonrpc": "2.0", - "id": "", - "method": "message/send" | "message/stream" | "tasks/get" | "tasks/cancel" | "tasks/resubscribe" | "tasks/pushNotificationConfig/set" | "tasks/pushNotificationConfig/get", - "params": { ... } -} -``` - -The response is either: - -``` -{ "jsonrpc": "2.0", "id": "", "result": } -``` - -or a JSON-RPC error envelope using A2A's documented codes (§4.6). - -### 4.3 Message structure - -A2A `Message`: - -``` -{ - "kind": "message", - "messageId": "", - "role": "user" | "agent", - "parts": [TextPart | DataPart | FilePart, ...], - "taskId": "" // optional — required to continue/resume - "contextId": "" // mandatory for cross-agent traffic on this platform - "metadata": { - "x-zynd-auth": { v, entity_id, public_key, nonce, issued_at, expires_at, fqan?, developer_proof?, signature } - } -} -``` - -Rules the implementation must hold: - -1. **`contextId` MUST equal the `dm_threads.id`** on which the connection lives. Without it the receiver can't correlate the task to a permission set or persist the message. -2. **First message of a task omits `taskId`.** The receiver assigns one. -3. **Continuation messages carry the same `taskId`.** Receiving a `taskId` the receiver doesn't have ⇒ JSON-RPC error `A2A_TASK_NOT_FOUND` (-32001). -4. **`messageId` is unique per message** and used for deduplication on the receiver side at the message layer (not the same as nonce, which is per-signature). - -### 4.4 Authentication: `x-zynd-auth` - -Per-message Ed25519 envelope, defined in `zyndai-ts-sdk/src/a2a/auth.ts` and `zyndai-agent/zyndai_agent/a2a/auth.py`. agent-persona reuses the SDK implementation directly — we do not roll our own crypto. - -Verification chain (receiver): - -1. Pull `auth = metadata["x-zynd-auth"]`. If absent and `auth_mode = strict`, reject with `ZYND_AUTH_FAILED` reason `missing_auth`. agent-persona uses `strict` for inbound. -2. Check version (`v == 1`). -3. Check expiry window (`now < expires_at`) and skew (`now ≥ issued_at - 60s`). -4. Check nonce uniqueness in the per-sender LRU replay cache. -5. Verify the public key hashes to the prefix in `entity_id` (`zns:` or `zns:svc:`). -6. JCS-canonicalize the message with `signature` blanked, prepend `ZYND-A2A-MSG-v1\n`, Ed25519-verify. -7. (If present) verify `developer_proof` against the agent's public key. - -agent-persona signing always includes `developer_proof` on the **first** message of a context, then drops it on subsequent messages to save bytes. The receiver uses this to assert the persona is HD-derived from the same developer key it expects. - -### 4.5 Request / response lifecycle (the canonical path) - -``` -Caller agent (Alice's persona) Callee agent (Bob's persona) -────────────────────────────── ─────────────────────────── -sign Message m1 (role=user, contextId=T) -POST /a2a/v1 verify x-zynd-auth(m1) - method = message/send check connection FSM (§5.4) - params.message = m1 allocate taskId, contextId - params.configuration.pushNotificationConfig = task: submitted - { url: PUSH_URL, token: TKN } register pushConfig - - dispatch handler thread - task: working - run orchestrator (external mode) - ... (may take seconds) ... - handler returns Result - task: completed, append artifact - -←──── JSON-RPC result: Task (state=completed) - (taken from in-memory store; identical to what tasks/get would return) - - deliver_push_if_configured(taskId) - POST PUSH_URL { wrapper: signed Msg - + DataPart{ status-update event }} -←──── push delivery (out-of-band) - -(caller updates UI / persists artifact) -``` - -Two delivery channels for the same final state — the synchronous JSON-RPC reply AND the async push notification. **Both paths converge on the same Task; reading either is sufficient.** The push exists for the case where the caller's process has died or never blocked (`blocking=false`) on the original call. - -### 4.6 Async handling (continuous conversation) - -The protocol supports continuous, multi-turn conversation **without** dropping out of the task lifecycle. Two patterns: - -**Pattern A — interrupted-state loopback (within one task).** Bob's handler calls `task.ask("which afternoon works for you?")`. The task transitions `working → input-required`, the JSON-RPC `result` returns to Alice. Alice's orchestrator inspects the Task, sees `state=input-required`, picks the question off `status.message.parts`, formulates a new message reusing the **same taskId**, and POSTs again. Bob's handler resumes from where it paused (the SDK's `task_store.suspend_until_next_message` already implements this). Loops until Bob's handler completes. - -**Pattern B — task chain (across tasks).** Some negotiations span multiple tasks under the same `contextId`. Example: task #1 = "ask availability", completes. Alice's orchestrator decides on a slot. Task #2 = "propose meeting at T". Each task gets its own taskId; both share the contextId. The receiver's `agent_tasks` table (the meeting ticket) references the contextId, not any single taskId, so it survives task boundaries. - -Pattern A is the default for back-and-forth. Pattern B is used when the work changes shape (a question vs. a proposal vs. an ack). - -### 4.7 JSON-RPC error codes used - -| Code | Name | When | -|---|---|---| -| -32700 | RPC_PARSE_ERROR | Body wasn't valid JSON | -| -32600 | RPC_INVALID_REQUEST | Not JSON-RPC 2.0 | -| -32601 | RPC_METHOD_NOT_FOUND | Unknown method | -| -32602 | RPC_INVALID_PARAMS | Validation failed on params | -| -32603 | RPC_INTERNAL_ERROR | Caught exception in dispatch | -| -32001 | A2A_TASK_NOT_FOUND | Unknown taskId | -| -32002 | A2A_TASK_NOT_CANCELABLE | Cancel attempted on terminal task | -| -32003 | A2A_PUSH_NOTIFICATION_NOT_SUPPORTED | (Reserved; not used by agent-persona) | -| -32004 | A2A_UNSUPPORTED_OPERATION | Connection not `accepted` so this method is not available | -| -32005 | A2A_CONTENT_TYPE_NOT_SUPPORTED | Future use | -| -32006 | A2A_INVALID_AGENT_RESPONSE | Output failed schema validation | -| -32100 | ZYND_AUTH_FAILED | Generic signature/identity failure | -| -32101 | ZYND_REPLAY_DETECTED | Nonce already seen in window | -| -32102 | ZYND_AUTH_EXPIRED | Outside the issued/expires window | - -Only these codes are ever produced by the new transport. Anything else is a bug. - ---- - -## 5. Connection Lifecycle - -### 5.1 Creating a connection - -Initiator goes from `none → requested`. The implementation already does this via `request_connection` (the MCP tool) and the UI's "start chat" path. New protocol invariants on top: - -1. The initiator's first cross-agent message after creating the row MUST be the connection-handshake message (a `DataPart` with `kind: "zynd.connection.request"` and the initiator's introduction text). This is sent via `message/send` and creates the receiver's first task on this connection. -2. The receiver's A2A server, on every inbound message, checks the ConnectionFSM (§5.4) before dispatching. A handshake message on a `requested` connection is **always** dispatched even though the connection isn't `accepted` yet — this is the ONE message exception that lets the receiver see the request content. -3. Until the receiver moves the connection to `accepted` via UI, the receiver's agent does NOT auto-reply. The initiator gets a Task with `state=submitted` and a status message like "awaiting human approval" — they can poll `tasks/get` or wait for the push notification when the human acts. - -### 5.2 Maintaining a session - -A session (contextId) is alive as long as the connection is `accepted`. Maintenance is implicit — there is no keepalive. The persistence layer is the source of truth, and the protocol carries no session state of its own outside of `taskId` continuations. - -### 5.3 Pause / resume — per-side mode (human takeover) - -Each participant independently flips their side of the connection between `agent` (AI handles) and `human` (human will reply manually). This is **NOT** a TaskFSM state — it is a connection-level toggle that affects how the receiver dispatches inbound messages. - -Behavior: - -* Side `mode = agent` (default): inbound message dispatches to the orchestrator handler as normal. -* Side `mode = human`: the receiver's A2A server acks the message at the transport level (the task transitions `submitted → working → completed` with an artifact `DataPart { kind: "zynd.takeover.queued" }` and a TextPart explaining "the principal will respond personally"). No orchestrator runs. The human types into the human channel of the UI; that lives outside this protocol. - -The caller can detect human takeover by inspecting the artifact's data kind. They typically downgrade their UI to "waiting for human" and stop sending agent-style follow-ups. - -Switching modes mid-session has no effect on in-flight tasks (those have already dispatched). Future tasks see the new mode. - -### 5.4 Pre-dispatch admission gate - -Every inbound `message/send` and `message/stream` runs the **admission gate** before allocating a task: - -``` -1. Verify x-zynd-auth (§4.4) → fail ⇒ JSON-RPC error -2. Look up dm_threads by contextId → not found ⇒ check sender_entity_id; - if unknown sender, see §5.1 handshake rule -3. ConnectionFSM check: - state == blocked → reject (-32004 + reason "blocked_by_receiver") - state == declined → reject (-32004 + reason "request_declined") - state == revoked → reject (-32004 + reason "connection_revoked") - state == requested → ONLY allow if message is a handshake DataPart - state == accepted → allow -4. Enforce expires_at on x-zynd-auth (already done in step 1, repeat-defensive here) -5. Compute permission snapshot from dm_threads.permissions; freeze on the task -6. Allocate taskId, transition submitted → working, dispatch handler -``` - -This gate is the single chokepoint where every inbound message is reasoned about. There is no other path into the orchestrator from the network. - -### 5.5 Termination - -Two termination paths: - -* **Soft (`revoked`):** either side decides to end the relationship without animosity. Permissions are zeroed, the connection FSM sits at `revoked`, all in-flight tasks transition `→ canceled` via `EV_T_PERMISSION_REVOKED`. New `message/send` from the other side is rejected at the admission gate. -* **Hard (`blocked`):** receiver actively rejects the peer. Same task cleanup as `revoked`. Additionally, every subsequent inbound message from this `entity_id` is rejected at step 3 of the admission gate without dispatching, even if the connection is later un-blocked. - -Termination is not directly exposed as an A2A method — the protocol is for tasks, not connection management. Termination is a UI action that mutates `dm_threads` and triggers the cleanup as a side effect. - ---- - -## 6. Error Handling and Reliability - -The exhaustiveness rule: **every code path that can fail has a named outcome state**. The matrix below enumerates them. There are no "best effort" branches. - -### 6.1 Failure matrix - -| # | Failure | Where it surfaces | Outcome state | User-visible behavior | -|---|---|---|---|---| -| 1 | Network unreachable on outbound `message/send` | Caller | n/a (caller-side) | Caller surfaces a `delivery_failed` Task (synthetic `failed`) with reason "network unreachable: ". Retry policy applies. | -| 2 | TLS / handshake error | Caller | n/a | Same as #1; reason "tls_error". No retry (could be a permanent misconfiguration on receiver side). | -| 3 | Non-200 HTTP response with no JSON-RPC envelope | Caller | n/a | Caller raises `A2AError` with `code=RPC_INTERNAL_ERROR`, body excerpt for debugging. Retry policy applies. | -| 4 | JSON-RPC error envelope returned (any `-32xxx`) | Caller | n/a | Caller surfaces the error code+message verbatim. No retry on auth/perm errors; retry on `RPC_INTERNAL_ERROR`. | -| 5 | `x-zynd-auth` missing / bad signature / replay / expired | Receiver | (no task created) | JSON-RPC error -32100/-32101/-32102 to caller. No persistence side effect. | -| 6 | Connection in `blocked` / `declined` / `revoked` | Receiver | (no task created) | JSON-RPC error -32004 with reason. Caller's UI marks the connection unusable. | -| 7 | Connection in `requested` and message is not a handshake | Receiver | (no task created) | JSON-RPC error -32004 reason `"awaiting_acceptance"`. Caller's UI surfaces "still waiting for accept". | -| 8 | Unknown `taskId` | Receiver | (no task created) | JSON-RPC error A2A_TASK_NOT_FOUND. Caller's orchestrator drops the in-memory taskId and starts a fresh task. | -| 9 | Payload validation failure | Receiver | task `rejected` | Status message carries reason. | -| 10 | No handler registered (impossible in production but defensive) | Receiver | task `rejected` | Status message: "no handler registered". Operational alert. | -| 11 | Handler raised an unhandled exception | Receiver | task `failed` | Status message has the exception string. Push delivered. | -| 12 | Handler intentionally `task.fail(reason)` | Receiver | task `failed` | Reason as written. | -| 13 | Handler `task.requireAuth(scheme)` and caller doesn't supply | Receiver | task `failed` after IDLE_TTL_INTERRUPTED | Reason: "auth-required idle timeout". | -| 14 | Caller blocks waiting for resume but caller's process dies | Receiver | task `failed` after IDLE_TTL_INTERRUPTED | Reason: "input-required idle timeout". Push notification fires if push config registered. | -| 15 | `tasks/cancel` while `working` | Receiver | task `canceled` | If the handler is in a tight CPU loop the cancel is observed at the next yield; if the handler is awaiting external IO, the cancel takes effect when IO resolves. We do not preempt running threads. | -| 16 | `tasks/cancel` on terminal task | Receiver | (no change) | -32002 A2A_TASK_NOT_CANCELABLE. | -| 17 | Connection moved to blocked/revoked while task in flight | Receiver | task `canceled` (EV_T_PERMISSION_REVOKED) | Status reason mentions which connection event triggered. | -| 18 | Push delivery POST fails (network, 5xx) | Receiver | task remains in its terminal state | Receiver retries push 3× with exponential backoff (1s, 4s, 16s). After exhaustion: log warning, do NOT change task state. The next `tasks/get` from the caller still returns the truthful state. | -| 19 | Push delivery POST returns non-2xx 4xx (caller rejects) | Receiver | task remains terminal, push abandoned | Same: caller can still `tasks/get`. | -| 20 | SSE stream client disconnect mid-task | Receiver | task continues to its natural end | When the handler completes, the unsubscribe was already triggered on disconnect; results are still in the task store. Caller can `tasks/resubscribe`. | -| 21 | Inbound body too large | Receiver | (no task created) | HTTP 413, no JSON-RPC envelope (this is a transport-layer failure before parse). | -| 22 | Replay cache memory pressure | Receiver | n/a | Per-sender bucket capped at 4096 entries; oldest evicted. May cause an "old replay accepted" if a sender exceeds 4096 messages within the skew window — acceptable since each message is also expiry-bounded. | - -### 6.2 Retry logic - -Retries are **caller-side only**. The receiver never retries handler dispatch. Handler-internal retries (e.g. an MCP tool retrying an OAuth refresh) are the tool's responsibility, not the protocol's. - -Caller retry rules — applied by the orchestrator when wrapping `message/send`: - -| Outcome | Retry? | Backoff | -|---|---|---| -| Network unreachable / TLS / 5xx without JSON-RPC | yes, max **3** attempts | 1s, 4s, 16s | -| `RPC_INTERNAL_ERROR` (-32603) | yes, max **2** attempts | 4s, 16s | -| `RPC_INVALID_PARAMS` (-32602) | no | — | -| `RPC_METHOD_NOT_FOUND` (-32601) | no | — | -| `A2A_TASK_NOT_FOUND` (-32001) | no (start fresh task) | — | -| `A2A_UNSUPPORTED_OPERATION` (-32004) | no (connection state, not transient) | — | -| `ZYND_AUTH_FAILED` / `_EXPIRED` / `_REPLAY_DETECTED` | no (re-sign with fresh nonce/timestamp ⇒ not a retry, a new send; max 1 fresh send) | — | -| 4xx other than the above | no | — | - -Retry-on-success bookkeeping: each retry uses a **fresh `messageId`** and **fresh `nonce`/`signature`**. The receiver dedupes on `messageId` if it's seen recently for this contextId — see §6.3 invariant 3. - -### 6.3 Timeouts - -* **Caller-side outbound HTTP timeout:** 5 minutes for `message/send`; 30 minutes for `message/stream`; 30 seconds for `tasks/get|cancel|pushNotificationConfig/*`. These are per-call, not per-task. -* **Idle TTL on interrupted tasks:** 1 hour (configurable per agent). Hits §6.1 #13/#14. -* **Terminal retention:** 5 minutes in-memory, indefinite in `a2a_tasks` persistence. -* **Handshake (request) TTL:** 30 days. After this a `requested` connection auto-transitions to `revoked`. -* **`x-zynd-auth` validity window:** 60 seconds default (configurable down to 10s for high-security flows, up to 5 min for low-network-quality flows). -* **Replay cache window:** matches the validity window; when a nonce expires it's swept. - -### 6.4 Recovery paths - -For every failure outcome, there is a recorded route back to a healthy state: - -* **Crashed receiver process while task `working`:** on restart, the in-memory task store is empty. The persistent `a2a_tasks` row is in `working`. A reconciliation pass on startup transitions all `a2a_tasks.state == 'working'` rows to `failed` reason `"server_restart"`. Push notifications fire. -* **Crashed caller process while waiting:** caller had registered a push config. On the next cold start the orchestrator queries `tasks/get` for any tasks it had marked "in flight" in its own DB and surfaces results that landed during downtime. -* **Persistent push delivery failure:** task is terminal but caller never knew. `tasks/get` still works. Next time the caller's orchestrator opens a task on this connection, it queries pending tasks first and clears the backlog. -* **Connection blocked with in-flight tasks on the caller side:** caller observes `A2A_UNSUPPORTED_OPERATION` reason `blocked` on the next send; existing tasks transition `canceled` via the receiver-side EV_T_PERMISSION_REVOKED. - ---- - -## 7. Permissions and Capability Model - -### 7.1 Three layers of enforcement - -| Layer | What it does | Where it runs | Failure mode | -|---|---|---|---| -| Card advertisement | Lists what the persona's agent CAN do (tags, skills, x-zynd capabilities) | `/.well-known/agent-card.json` | A caller asking for something not advertised gets a polite refusal but is not protocol-blocked — the card is informational, not authoritative for permission. | -| Connection permissions | Per-thread booleans (§2.6) | Admission gate (§5.4) freezes them on the task; orchestrator reads them inside `external_permissions` | A foreign agent sending a message that requires a permission the connection lacks: the orchestrator builds an allowlist that excludes the gated tool; if the LLM still calls it, it's hard-blocked. | -| Tool allowlist | Set of MCP tool names the orchestrator will dispatch in external mode | Orchestrator pre-flight check before MCP `_call` | Tool call rejected; result fed back to the LLM as `{ error: "permission_denied", message: "..." }` so the LLM can produce a graceful refusal. | - -The three layers are independent. A capability appearing on the card does not grant connection permission. A connection permission does not bypass the orchestrator's hard-block (defense in depth: even if a misconfiguration leaks the permission flag, the allowlist is the final gate). - -### 7.2 Permission semantics - -* **Connection-scoped, not task-scoped.** Permissions live on `dm_threads`, not on tasks. -* **Captured on dispatch.** The admission gate reads the current permission snapshot when allocating the task and pins it on the task record. A permission change made after dispatch does NOT retroactively affect the in-flight task. -* **Default conservative.** New connections start with only `can_request_meetings = true`. The receiver opts the foreign side into anything else explicitly. -* **Receiver-owned.** Only the receiver can set `dm_threads.permissions` via `PATCH /api/persona/threads/{thread_id}/permissions`. The initiator never gets to set their own permissions on this connection. -* **Symmetrical for negotiation.** Each side has its own permission snapshot — Alice's permissions on the connection govern what Bob's agent can ask Alice's agent to do, and vice versa. - -### 7.3 Where permissions enter the architecture - -* **§5.4 admission gate** — looks them up; included on the task record so it can't drift. -* **§3.1 ConnectionFSM** — permissions are inherited across `requested → accepted` (defaults). They survive `accepted ↔ blocked ↔ accepted` round trips. They are zeroed on `revoked`. -* **§8.2 MCP integration** — the allowlist is computed from these and the LLM only sees allowed tools. -* **Agent card extension** — the card MAY include an `x-zynd.permissions.required_for_action` map listing which permission a foreign agent would need to invoke each advertised skill. This is informational, not enforced. - -### 7.4 Capability advertisement (separate from permissions) - -`persona_agents.capabilities` is converted into the agent card's `skills[]` array on every card rebuild. This is what shows up to other agents during discovery. The capability list never grants anything — it just makes the persona findable. The receiver still enforces permissions per connection. - ---- - -## 8. MCP Integration - -### 8.1 When MCP is invoked - -MCP tools (the existing ContextAware registry) are invoked **only inside the orchestrator's handler**, which runs after the admission gate has accepted the inbound message. There is no path from the network to MCP that bypasses the gate. - -### 8.2 Permission checks before invocation - -For each tool call the LLM emits in external mode: - -``` -1. Look up the per-task pinned external_permissions -2. Compute allowlist = EXTERNAL_DEFAULT_ALLOWED ∪ ⋃{ permission_gates[k] | external_permissions[k] } -3. If tool_name ∉ allowlist: - inject { error: "permission_denied", message: "..." } into the conversation - skip MCP._call entirely -4. Else: - execute MCP._call(tool_name, args) - wrap result back into the conversation -``` - -This is already implemented in the current orchestrator (§4.2 of `agent-persona/architecture.md`) and is preserved. The change in this design is that the permission set is **captured on the task** at dispatch time, not pulled fresh on every tool call. Pulling fresh would race with permission edits and produce inconsistent task histories. - -Note also: the `propose_meeting` direction-fix in the existing orchestrator (foreign agent triggers a proposal on behalf of its principal) is preserved — it's a property of the `propose_meeting` tool, not of the protocol. - -### 8.3 MCP failure handling within the state system - -| MCP outcome | TaskFSM impact | -|---|---| -| Tool returns `{ error: "..." }` | No FSM impact. Result fed to LLM; LLM either retries with different args, or wraps up the task with `working → completed` and the error explained in the artifact text. | -| Tool raises a Python exception | No FSM impact. Caught in orchestrator, fed back to LLM as `{ error: "Tool execution failed: ..." }`. | -| Tool succeeds | No FSM impact during execution; eventual `working → completed` driven by the LLM finishing its loop. | -| Tool times out (network upstream) | No FSM impact; tool layer's timeout returns an error dict. Orchestrator does NOT cause the task to fail — that's the LLM's call. | -| The orchestrator's overall iteration cap is hit (max_iterations=6) | Task `working → completed` with whatever text the LLM had; or `working → failed` if the LLM never produced a final reply. The choice is determined by the orchestrator's existing logic. | - -Critically: a misbehaving MCP tool can never land the task in a state outside the TaskFSM. Every tool error eventually surfaces as a `completed` artifact (with the error described), a `failed` task (with reason), or a `canceled` task (if the caller cancels mid-execution). - -### 8.4 Tool calls that produce A2A traffic of their own - -`message_zynd_agent` (the orchestrator's outbound networking tool) is itself an A2A caller. When invoked inside an external-mode handler, the outbound A2A `message/send` it produces opens a new task **on a different connection** (the destination). That task lives in its own TaskFSM, with its own `contextId` (the other connection's `dm_threads.id`). There is no cross-talk between the two TaskFSMs — the outer task waits on the tool result like any other tool call; the inner task settles independently. - -This compositional property is what makes networking-by-proxy work cleanly: Alice's agent can be in the middle of a task with Bob, where the tool that helps complete that task is itself a task with Carol. All three tasks are independent, all governed by the same FSM, all auditable. - ---- - -## 9. Invariants and Guarantees - -### 9.1 Identity invariants - -* **I-1.** Every cross-agent message that the receiver dispatches has `auth.signed = true`. (Strict mode; missing or invalid auth ⇒ rejected at gate.) -* **I-2.** `entity_id` matches the SHA-256 prefix of `public_key`. The receiver verifies this before any other check. -* **I-3.** `developer_proof`, when present, verifies against the agent's pubkey. The receiver MAY require it on the first message of a context. - -### 9.2 Connection invariants - -* **C-1.** `dm_threads.id` is the `contextId` for every cross-agent message on this connection. There is no other place a contextId comes from. -* **C-2.** The ConnectionFSM is in exactly one of six states at any time; transitions follow §3.1.3 strictly. -* **C-3.** Permission writes are only accepted from the receiver's side via the documented endpoint. Service role bypasses are reserved for backend reconciliation only. - -### 9.3 Task invariants - -* **T-1.** Every task is in exactly one TaskFSM state at any time. Transitions follow §3.2.3 strictly. -* **T-2.** A task always reaches a terminal state. Either via handler completion, timeout, cancel, or connection-level revocation. There is no eternal task. -* **T-3.** A task's permission snapshot is fixed at dispatch and is not mutated for the lifetime of the task. -* **T-4.** A task's `contextId` is fixed at creation and never mutates. - -### 9.4 Message invariants - -* **M-1. No silent drops.** Every accepted inbound message is persisted (in `a2a_tasks.history` AND, for human-readable reflection, in `dm_messages` channel='agent') before the handler is dispatched. If the handler crashes, the message is still on record. -* **M-2. No spurious duplicates from the receiver.** The receiver dedupes on `(contextId, messageId)` within a 5-minute window. A retried send with the same `messageId` is a no-op (returns the existing task state). -* **M-3. Replay protection.** Per-sender LRU cache rejects any message whose nonce is already seen within its expiry window. -* **M-4. Strict ordering within a task.** Handler-side: messages on the same `taskId` are processed in the order they were verified by the gate. The gate is single-threaded per task (the suspended-handler resume model in `task_store` enforces this). Across tasks there is no ordering guarantee. -* **M-5. Signed reflectors.** Every outbound message the receiver sends back (status updates, ask questions, completion artifacts, push notifications) is itself signed by the receiver's keypair. The caller can verify the response exactly the same way. - -### 9.5 Reliability invariants - -* **R-1. At-least-once handler dispatch.** A successfully verified message on an `accepted` connection causes the handler to run at least once (zero-or-more in failure recovery: a crash during dispatch is reconciled to `failed` on restart, and if the caller retries with the same messageId, M-2 dedupes it). -* **R-2. At-most-once observable side effect.** A handler that performs an externally-visible action (sending an email, writing to LinkedIn) MUST do so via an MCP tool that itself implements idempotency (e.g. by hashing the payload). The protocol guarantees the handler runs at-least-once; the tool guarantees the side effect happens at-most-once. -* **R-3. Push delivery is best-effort with bounded retries.** Push notification is a UX optimization; truth lives in the task record reachable via `tasks/get`. -* **R-4. Bounded memory.** Per-sender replay cache caps at 4096 entries. Per-process task store sweeps interrupted-state tasks past TTL and terminal-state tasks past retention. No unbounded growth. - -### 9.6 Predictability and fail-safety - -* **P-1. Determinism.** For every pair (state, event) on both FSMs there is exactly one outcome documented in the transition tables. There are no implementation-defined behaviors. -* **P-2. Conservative defaults.** Every default chosen in this design fails toward refusal rather than action: connections start `requested`, permissions start mostly `false`, auth_mode starts `strict`, mode starts `agent`. A misconfigured persona is unreachable, not over-permissive. -* **P-3. Cryptographic accountability.** Every accepted message and every emitted reply is signed with an Ed25519 key tied to a registered persona. There is no anonymous traffic on this protocol after this redesign. -* **P-4. Schema-bounded payloads.** Every task's input is validated against the persona's `payloadModel` (Zod / Pydantic) before the handler sees it. Validation failure ⇒ `rejected`, never silent acceptance of garbage. - ---- - -## 10. Implementation Notes (non-binding, for the next stage) - -This section is informational — it sketches the changes needed to realize the architecture. It does NOT specify code; the code is for the implementation step. It is included so reviewers can sanity-check that the design is realizable. - -### 10.1 Schema deltas - -* `dm_threads.status` — add values `declined` and `revoked` to the existing CHECK constraint. -* New table `a2a_tasks` — keyed by `task_id` UUID. Columns: `task_id`, `context_id` (= `dm_threads.id`), `state`, `permission_snapshot` JSONB, `created_at`, `updated_at`, `terminal_at`, `idle_until`, `push_url`, `push_token`, `history` JSONB, `artifacts` JSONB, `last_message_id`, `idle_ttl_ms`, `failure_reason`. Indexed on `(context_id, updated_at desc)`. -* `persona_agents` — add `card_path` and `a2a_path` columns documenting the per-persona endpoints (defaulting to `/api/persona/{user_id}/.well-known/agent-card.json` and `/api/persona/{user_id}/a2a/v1` respectively). - -### 10.2 Module deltas inside agent-persona - -* `agent/agent_message.py` — retired. Outbound construction goes through the SDK's `to_a2a_message` + `sign_message`. -* `api/persona.py` — `/webhooks/{user_id}` and `/webhooks/{user_id}/sync` removed. New blueprint mounting the SDK's `A2AServer` per persona at `/api/persona/{user_id}/a2a/v1` and `/api/persona/{user_id}/.well-known/agent-card.json`. Single `handler` registered with the server is the existing `handle_user_message` entry, wrapped to translate A2A `HandlerInput → handle_user_message` args and `handle_user_message return → task.complete(...)`. -* `api/persona.py` — new `POST /api/persona/push/{user_id}` endpoint that receives push wrappers from peers and updates the local task records (the caller side). -* `mcp/tools/zynd_network.py::message_zynd_agent` — rewritten to use the SDK's `A2AClient` rather than raw `requests.post + DB poll`. Reply discovery is the JSON-RPC result of `client.sync(...)`; multi-turn uses `client.stream(...)` or repeated `client.sync(taskId=...)`. -* `agent/orchestrator.py` — unchanged in shape; it gains the per-task `permission_snapshot` parameter (replacing the on-the-fly permission read) and learns to recognize `Task.status.state == "input-required"` as a "the other side asked something — show it to my principal" signal. - -### 10.3 Out of scope explicitly - -* The Ed25519 identity layer, HD derivation, heartbeat manager, registry registration, and developer-proof generation are unchanged. They already work and are not part of the A2A redesign. -* The human channel (`channel='human'` rows on `dm_messages`) is unchanged. People still type to people the way they always did. -* The MCP tool internals are unchanged. Only the layer above MCP changes. - ---- - -## 11. Open questions deferred to implementation - -These are conscious deferrals — they don't block the design, but they need answers when code lands. - -1. **Per-persona vs shared task store.** All personas live in one FastAPI process; do they share one `TaskStore` or get one each? Memory cost is trivial either way; the question is observability. Recommendation: one per persona, keyed by `user_id`, so debug logs cleanly attribute task IDs. -2. **Public endpoint shape for the agent card.** The SDKs default to `/.well-known/agent-card.json` at the persona's base URL. We must decide whether to expose `https://server/.well-known/agent-card.json` (one persona per host, won't work multi-tenant) or `https://server/api/persona/{user_id}/.well-known/agent-card.json` (multi-tenant, non-spec path). Recommendation: the second, and the registry stores the explicit URL in the entity_url so external peers don't depend on the default path. -3. **Card signature key rotation.** Out of scope for v3 but the schema should anticipate it: agent cards already support multiple `signatures[]` entries. -4. **Backpressure on push delivery.** If a peer is slow to ack push POSTs, we shouldn't queue arbitrarily many of them. Cap per-peer in-flight pushes at 16; drop oldest with a logged warning. - ---- - -## 12. Glossary - -| Term | Meaning | -|---|---| -| A2A | Agent-to-Agent protocol, https://a2a-protocol.org/v0.3.0/specification/ | -| Agent card | The signed JSON document at `/.well-known/agent-card.json` describing a persona's capabilities, endpoints, and identity. | -| Connection | The long-lived relationship between two personas, stored on `dm_threads`. | -| Context (`contextId`) | A2A's name for the conversation; we equate it with `dm_threads.id`. | -| Handler | The function the SDK's A2AServer calls to process an inbound message. agent-persona's handler delegates to the existing orchestrator. | -| Initiator | The side that opened the connection. | -| JCS | RFC 8785 JSON Canonicalization Scheme — the deterministic JSON serialization used for signing. | -| Orchestrator | The existing LLM-driven conversation engine in `agent/orchestrator.py`. | -| Persona | A user-deployed AI agent on the Zynd Network. One per principal. | -| Principal | The human who owns the persona. | -| Push notification | A receiver-initiated signed POST that wakes up the caller when a task reaches a terminal state. | -| Receiver | The side that owns the inbox a message is being delivered to. | -| Session | Synonym for context (in conversational terms). | -| Task (`taskId`) | A single discrete request inside a context. | -| `x-zynd-auth` | The per-message Ed25519 envelope embedded in `Message.metadata`. | \ No newline at end of file diff --git a/docs/persona/AUDIT_REPORT.md b/docs/persona/AUDIT_REPORT.md deleted file mode 100644 index b72b0b4..0000000 --- a/docs/persona/AUDIT_REPORT.md +++ /dev/null @@ -1,208 +0,0 @@ -# Agent-Persona — End-to-End Audit Report - -**Date:** 2026-08-05 -**Scope:** Full-stack audit of the live production deployment (`api` on :8000, `web` on :3001, PM2-managed, `persona.zynd.ai`) — feature testing, integration health (Zynd Network, memory layer, Google/LinkedIn/Twitter/Notion), chat UX, privacy/OAuth scoping, and missing-feature research. -**Status:** **Findings only — nothing in this report has been fixed yet.** See "Next Steps" at the bottom. - ---- - -## How this was tested - -This was a production system already serving real traffic, so I deliberately avoided anything that could disrupt it or a real user's data: - -- **No restarts, no live chat traffic injected as the real user, no new persona/agent registered on the public Zynd registry.** I did not mint an auth token to impersonate the account. -- **Static code audit** of the full request path (backend orchestrator → MCP tools → frontend renderers) using the code graph plus direct reads, cross-checked by three focused background passes (log mining, chat-pipeline trace, OAuth scope audit). -- **90 days of real production logs** (`pm2 logs api/web`, full history back to 2026-05-07) — this is genuine evidence of what's actually happening in prod, not synthetic testing. -- **Backend test suite**: `144 passed, 0 failed` (`backend/tests/`, run via the project's own `.venv`). -- **LinkedIn live posting/DM was skipped**, per your instruction (no Apify credits) — the LinkedIn *code path* (scopes, scraper, tool implementations) was still audited. - -Confidence is high throughout — every claim below cites a file:line or a specific log line, not speculation. Where I'm inferring rather than certain, I've said so. - ---- - -## Executive summary - -The core architecture (HD key derivation, A2A protocol, the Zynd search/ranking logic, Google scope minimization on Docs/Gmail) is genuinely well engineered — there's careful, documented thinking throughout. But three things are quietly broken in ways a normal user would never get an error message for, and one advertised feature is completely inert: - -1. **The memory layer / context graph — the personalization feature — does nothing right now.** A required shared secret was never set in production. Every "remember this," "what do you know about me," proactive nudge, daily brief, and digital-twin personalization call silently no-ops. -2. **The backend has been crashing from native memory corruption roughly once a day for the last 90 days**, including hours before this audit. Root cause is unidentified (glibc-level, not a Python traceback). -3. **Realtime updates have been broken since the very first day in the logs (2026-05-07).** A backend bug means meeting/connection notifications never reach the frontend via realtime, which is very likely *why* the dashboard resorts to polling 5 endpoints every 8-20 seconds, continuously, for as long as a tab stays open (one session ran 34 days straight). -4. **The "ugly mixed-format" chat bug you flagged is real and I found the exact cause**: two different search tools are both sanctioned by the system prompt for the same "find people" query, return incompatibly-shaped results, and land in two visually unrelated UI components — plus the prompt tells the model to write full prose descriptions of people *and* a card renders the same people, with no instruction to be brief (unlike the sibling code path, which does have that instruction). - -None of this needed exotic testing to find — it's either currently happening in production logs or directly readable in the code. Details below, worst first. - ---- - -## P0 — Critical - -### 1. Memory layer / context graph is fully disabled in production - -- `backend/config.py:127`: `MEMORY_LAYER_JWT_SECRET: str = os.getenv("MEMORY_LAYER_JWT_SECRET", "")` — **not present in `backend/.env`** (confirmed: zero matches). -- `backend/agent/memory_client.py:61-63`: `_is_enabled()` returns `bool(config.MEMORY_LAYER_JWT_SECRET)` — always `False` right now. -- Every memory call degrades silently: `get_context`, `ingest_turns`, `confirm_fact`, `forget_fact` all short-circuit to empty/no-op (`memory_client.py:116-117, 192-193, 229-230, 249-250`). -- The user-facing tools are explicit about it: `backend/mcp/tools/memory.py:35-36` — if a user says "what do you remember about me," the persona literally replies *"Memory layer is not configured. Ask your admin to set MEMORY_LAYER_JWT_SECRET."* -- **Blast radius is bigger than just memory Q&A** — six subsystems import `memory_client`/`memory_context` and depend on it for personalization: `agent/nudge_engine.py`, `agent/network_intros.py`, `agent/digital_twin.py`, `agent/daily_brief.py`, `agent/proactive_loop.py`, `mcp/tools/twin.py`. All of these are running with zero personal context right now. -- `load_memory_context`/`ingest_conversation` are wired into **every** chat turn (`agent/orchestrator.py:2546, 2640, 2672, 2849, 2936, 3027, 3058, 3222`) — so this isn't a rarely-hit path, it's on the hot path of every single conversation, just quietly doing nothing. - -**Fix direction (not yet applied):** set `MEMORY_LAYER_JWT_SECRET` in `backend/.env` to match whatever the memory-layer service (`api.zynd.ai`) has configured, restart `api`, confirm with a real "remember that I like X" → "what do you remember about me" round trip. - -### 2. Backend is crashing from native memory corruption, ~daily, for 3 months - -From full log history (`api-err.log`, 2026-05-07 → 2026-08-05, i.e. today): - -- **86 glibc heap-corruption aborts** — `malloc(): unsorted double linked list corrupted`, `double free or corruption (!prev / out)`, `corrupted double-linked list`, `free(): corrupted unsorted chunks`, `malloc(): invalid size`, `malloc(): unaligned tcache chunk detected`. Averaging ~1/day, still occurring — most recent at **2026-08-05T19:56:33**, i.e. hours before this audit started. -- Each one kills the process instantly with **no Python traceback and no graceful-shutdown log line** — PM2's `autorestart` brings it back 1-2s later, which is why this has been invisible: users see a brief hiccup, not an error page. -- This is a native-code-level bug (glibc heap corruption doesn't originate in Python), so the usual suspects are C-extension dependencies loaded in-process — `grpc`/`google-api-core`, crypto libs (Ed25519 signing runs constantly for the heartbeat manager), or similar. **Root cause is not identifiable from application logs alone** — this needs a core dump (`ulimit -c unlimited`) or running under `valgrind`/ASAN to pin down which library is corrupting the heap. -- Secondary effect: of **2,936 total API process restarts** in 90 days, only 86 (2.9%) are explained by these crashes and 163 (5.5%) by clean deploy-shutdowns — **the remaining 91% have no explanatory log line at all**, which suggests either very frequent external restarts (CI/deploy hitting the process with a hard signal that bypasses uvicorn's graceful shutdown) or something else not visible in these logs. Worth cross-checking against deploy tooling separately. - -### 3. Realtime broadcasting has been broken since day one — and is very likely *why* the app polls so aggressively - -- `backend/services/meetings.py:88-97` (`_broadcast`) and `backend/mcp/tools/zynd_network.py:1361-1372` (`request_connection`'s new-thread ping) both call `.channel("system_pings").send(...)` on a client returned by `config.get_supabase_anon()`. -- `backend/config.py:170-179`: `get_supabase_anon()` uses `create_client` (the **synchronous** Supabase client), which does not support realtime broadcast sends. Every call throws, is caught, and logged as a warning — never surfaced anywhere else. -- Confirmed in production logs: **27 occurrences of `[meetings] broadcast failed for {event}: ... "This feature isn't available in the sync client. You can use the realtime feature in the async client only."`**, spanning the *entire* log history from **2026-05-07T06:47:50 to 2026-08-05T18:27:30 (today)**. This has never worked. -- Direct consequence: meeting-proposal and new-connection notifications never reach the frontend in real time. Correlates with real `book_failed` states in the logs (e.g. `2026-05-07T06:48:25`, `2026-08-05T18:26:41`). -- The frontend's own realtime subscriptions (`webapp/src/components/MessagesPanel.tsx:205-221`) use the JS Supabase client, which *does* support broadcast — so the client side is fine; this is purely a backend bug (wrong client type for the send side). -- **Likely root cause of the polling load** (see P2 below): `webapp/src/contexts/DashboardActivityContext.tsx:202` polls `/api/todos/`, `/api/approvals/`, `/api/meetings/pending/*`, `/api/groups/invitations/incoming`, and `/api/persona/*/status` every **20 seconds** unconditionally — a design that makes sense as a fallback *if* realtime is assumed unreliable, but given realtime broadcasting has never actually worked on the backend, polling is currently the *only* way these views update at all. - -**Fix direction:** use `create_async_client` (or Supabase's REST broadcast endpoint, which doesn't require the realtime websocket client) for server-side broadcasts in `meetings.py` and `zynd_network.py`. Once broadcasts genuinely work, the 20s poll can likely be relaxed significantly. - -### 4. MCP tool server runs with security disabled in production - -- `backend/mcp/server.py:107,113,233`: `create_mcp_server(disable_security: bool = True)` is called with no override — `mcp_server = create_mcp_server()` at module load. -- `contextaware/ContextAware.py:17-28`: when `disable_security=True`, the server explicitly logs `"Security disabled. NOT RECOMMENDED!"` — confirmed firing on every restart in production logs (`2026-08-05T19:56:33`, `21:11:20`, etc.), alongside a freshly generated (and then unused, since security is off) API key. -- I did not trace how/whether this MCP endpoint is network-reachable beyond localhost — that's the key open question determining real severity — but the library's own wording ("NOT RECOMMENDED") and the fact it's silently true in prod (not an explicit, documented decision anywhere I found) makes this worth a deliberate go/no-go decision rather than an accidental default. - ---- - -## P1 — High priority - -### 5. The chat UX bug you described ("find AI founders" → cards + text + say-hi cards) — full root cause - -This is a **combination bug**, confirmed via code + a prior engineer's own comment describing the exact same symptom: - -- **Two tools, one query, two shapes.** The system prompt (`backend/agent/orchestrator.py:2235, 2263`) tells the LLM it's fine to use *either* `search_zynd_personas` or `search_zynd_network(kind="persona")` for a people-search like "AI founders" ("prefer... " not "always"). They return differently-keyed rows — `{name, agent_id, description, ...}` (`zynd_network.py:1002-1010`) vs. `{name, entity_id, summary, ...}` (`zynd_network.py:645-659`) — and the frontend has a purpose-built card renderer for only one of them. -- **Five independent, tool-name-keyed renderers, no shared "search result" component:** - | Tool | Renderer | Look | - |---|---|---| - | `search_zynd_personas` | `MatchCard` (`webapp/src/components/chat/MatchCard.tsx:26-62`, wired at `ChatInterface.tsx:209-237`) | avatar + name + reason + **"Say hi →"** button — this is almost certainly the "hi card" you're describing | - | `search_zynd_network` | `AgentResultRow`/`ServicesPanel` (live-stream only, `ChatInterface.tsx:673-728`) | plain technical list row with a `type-{kind}` badge and "Call"/"View card" buttons | - | `call_zynd_service`/`call_zynd_agent` | `GenUiResult` (`GenUiResult.tsx:393-497`) | generic shape-classified card (table/list/record/raw) | - | everything else | plain `ReactMarkdown` | plain text | -- **A previous engineer already found and half-fixed this.** `ChatInterface.tsx:660-672` has a comment explaining that `search_zynd_personas` results used to render **three times in one turn — "once as this block, once as prose, once as MatchCard"** — and excludes it from the generic card path. But `search_zynd_network` (the *other* tool the prompt sanctions for the identical query) was never given the same treatment, so the exact bug the comment describes is still fully reachable whenever the model picks that tool instead. -- **The system prompt directly contradicts itself** on whether the model should re-describe results in prose: `orchestrator.py:2139-2152` and `:2266` say "keep it short, the card carries the detail," while `orchestrator.py:2279-2287` ("When presenting PEOPLE results") explicitly instructs 1-2 full sentences *per person* — for the same result set the card is about to render. -- **Bonus finding:** `search_zynd_network` results aren't persisted (`webapp/src/components/chat/types.ts:45-47`, "Local-only, never persisted") — reload the page or switch conversations and those cards vanish entirely, leaving whichever prose the model happened to write. -- **Bonus finding #2:** the tool's own careful docstring guidance (`zynd_network.py:578-587`, explicit "IMPORTANT" advice on when to use `kind="persona"`) never reaches the model at all — `ContextAware.register()` (`contextaware/ContextAware.py:47-54`) replaces it with a much shorter `description=` string passed at registration (`mcp/server.py:170-171`). The good guidance exists; the LLM has never seen it. - -**Fix direction:** pick one tool as the canonical path for people-search (the code comment already argues for `search_zynd_personas`), stop sanctioning the other for the same intent, and remove the "write 1-2 sentences per person" instruction now that a card renders the same info — mirroring the brevity instruction that already exists for the network/service path. - -### 6. Zynd Network heartbeat is failing constantly - -- **3,273 heartbeat failures over 90 days** (~36/day), still occurring: `TimeoutError: timed out while waiting for handshake response` (2,007×, dominant), `InvalidStatus: server rejected WebSocket connection: HTTP 502` (1,022×, stopped 2026-05-27), `ConnectionClosedError` (88×), `InvalidMessage` (64×), `ConnectionRefusedError` (53×), DNS resolution failures (34×, 2026-06-08/10 outage window). -- Heartbeats are what keep a persona showing as "active" on the network for discovery/messaging — this directly affects the reliability of the exact feature you asked me to test ("find AI founders" and the underlying network). One explicit persona-search timeout is also logged (`2026-08-04T12:40:13`). -- This looks like an external dependency issue (connectivity to `dns01.zynd.ai`) more than an app bug, but it's frequent enough to be worth surfacing/monitoring rather than silently retrying forever. - -### 7. Google Calendar: users who connected Gmail-only get silent, confusing 403s - -- Calendar is an **optional, separately-granted** Google feature (`backend/api/oauth_routes.py:285-326`) — a user who only ever clicked "Connect Email" never has the `calendar` scope on their stored token. -- Nothing gates a Calendar API call on that scope actually being present. `mcp/tools/google/calendar.py:39-47` → `common.py:73-74` only checks *whether* a Google token exists, never *which scopes* it has — unlike the equivalent, already-existing pattern for Drive (`backend/api/brief.py:86-104`, `_has_drive_scope()`, which returns a clean, actionable error). -- Confirmed happening in production: `2026-08-05T18:26:41`, `HttpError 403 ... "Request had insufficient authentication scopes"` from the new conflict-detection code (`calendar.py:84-113`, added in the recent "detect scheduling conflicts" commit) → directly caused a real `book_failed` meeting state that day. Same gap affects `agent/smart_scheduling.py:267-273`, `agent/daily_brief.py:142-143`, and `agent/group_calendar.py:81-92`. -- Also in logs: a **second, distinct Calendar bug** — `2026-08-05T16:50:10/16:50:44`, `HttpError 400 "Invalid attendee email"` with 26 invalid-email entries in one response, suggesting a malformed attendee list gets passed through somewhere upstream of the API call. - -**Fix direction:** add a `_has_calendar_scope()` check mirroring the existing Drive one, and return a "reconnect Google Calendar" prompt instead of letting the raw Google error surface. - -### 8. Systemic missing input validation → 500s instead of 404s - -- **27 tracebacks**, `postgrest.exceptions.APIError: invalid input syntax for type uuid`, across `api/groups.py:125`, `agent/persona_manager.py:565`, `services/meetings.py:380/403`, `api/matches.py:75` — none of these endpoints validate that a path param is actually a UUID before querying, so any non-UUID value (`/1`, `/test`, `/abc`, `/%20`, `/bogus-user-000`, `/999999`, …) 500s instead of a clean 400/404. -- This produced real, repeated 500s: `/api/groups/discover` (14×), `/api/groups/auto-join-candidates` (13×), `/api/persona//status` (9× legitimate + many fuzz variants), `/api/meetings/pending/` (24× — the single most common 500 in the whole log). -- Some of this traffic looks like external scanning/fuzzing rather than real users, but real users can trivially hit the same bug (a stale bookmark, a copy-paste error, a client-side bug passing the wrong ID type). - -### 9. OAuth token storage is plaintext - -- `backend/services/token_store.py:46-58` stores `access_token`, `refresh_token`, and `raw_data` (a redundant copy of both) as plain `TEXT`/`JSONB` columns (`backend/db/schema.sql:38-50`). No `pgcrypto`, Supabase Vault, KMS, or app-level encryption anywhere in the schema or migrations. -- Protection today is entirely Row-Level Security (only the owning user or the service role can read a row) — reasonable as a baseline, but it means a leaked service-role key or a SQL-injection-class bug would expose live Gmail/Calendar/Drive/Twitter/LinkedIn/Notion tokens for every connected user in cleartext. - -### 10. Google Sheets integration is completely broken for every user - -- `backend/api/oauth_routes.py:292` still documents `"sheets"` as a valid `features` value, but `feature_map` (`:313-317`) has **no `"sheets"` key** — so the Sheets scope (`.../auth/spreadsheets`) is never requested for anyone, ever. -- `backend/mcp/tools/google/sheets.py` (`create_spreadsheet`, `append_to_sheet`, `read_sheet_values`) will 403 for every single user who tries it. This is a pure functional bug, not a scoping/privacy issue — flagging it here because it surfaced during the scope audit and would otherwise go unnoticed (no user has the scope to even discover it's broken). - ---- - -## P2 — Medium - -- **`httpx`/`httpcore` "Server disconnected" — 633 occurrences** across the log (49 as full tracebacks, the rest as caught/logged warnings in background pollers: `[brief_watcher] Poll failed` ×257, `[a2a poll] list_pending failed` ×195, `[a2a poll] fetch failed` ×13). This is the single largest recurring backend fault pattern — looks like a connection-pool idle-timeout mismatch between the app's httpx pool and Supabase's edge closing idle connections. `config.py:148-155` already added one retry for exactly this class of error; it's not fully covering it. -- **Aggressive polling, now with hard numbers.** `DashboardActivityContext.tsx:202` (`POLL_MS = 20_000`) confirmed against live traffic: sessions polling all 5 endpoints in lockstep every 8-20s, one running **continuously for 34 days** (still active as of today), another for 18 days. At these rates, a single open tab generates on the order of 15,000-20,000 requests/day. No `document.visibilitychange`-based backoff appears consistently applied. This is the direct downstream cost of finding #3 (broken realtime). -- **Twitter `offline.access` scope requested but the refresh flow was never implemented** (`mcp/tools/twitter.py` builds `tweepy.Client(access_token=...)` only, no refresh call anywhere) — access tokens will just go stale and force a full reconnect rather than silently refreshing. Not a privacy over-grant (grants no extra API surface), but a functional gap. -- **Google `openid`/`email`/`profile` scope requested but never consumed** (`oauth_routes.py:307` — no `id_token` decode or userinfo call anywhere for Google, unlike LinkedIn which does use its equivalent). Dead scope — should either be removed or actually used to prefill the user's name/avatar. -- **`page_publisher.py`**: `get_page_public` throws `'NoneType' object has no attribute 'data'` for missing slugs (5× in logs, 2026-07-27/28) — resolves to a correct 404 at the HTTP layer but logs an ugly, uninformative error; a simple null-check away from being clean. -- **A2A payment schema drift**: `persona_agents.pricing` column referenced in code doesn't exist in the DB (`agent/a2a_router.py:486`) — currently papered over by a fallback that treats every request as free (`[a2a payment] pricing column missing — treating addressee as free`, 2026-05-26), not an actual fix. -- **Next.js "Failed to find Server Action" — 55 occurrences** (2026-05-11 through at least 2026-06-03) with a mix of plausible real hash-format IDs (stale client bundle after a deploy) and obviously synthetic ones (`"x"`, `"y"`, 40 zeros — bot fuzzing). Worth a client-side error reporting tool (Sentry or similar) since these server logs can't see actual browser-side runtime errors at all — meaning frontend health here is likely under-observed. - -## P3 — Low / polish - -- Generic 403 error messages: `error_utils.py:75-80` classifies "insufficient scope" the same as generic "forbidden" ("check the permissions for this account") rather than the more specific, more actionable reconnect-hint already used for expired tokens (`:47-52`) — small change, meaningfully better UX for finding #7. -- `web-err.log` has a cosmetic lockfile-location warning on every one of 269 restarts (`package-lock.json` vs `webapp/pnpm-lock.yaml` both present) — one line in `next.config.js` (`outputFileTracingRoot`) silences it. -- Bot/scanner noise hitting `/api/.env`, `/api/v1/.env` … `/api/staging/.env`, `/api/graphql`, `/api/proxy` (~36 hits) — not an app bug, but suggests no WAF/rate-limiting in front of the API; worth considering given real secrets live in that file. -- `APP_SECRET_KEY` (`config.py:134`) is left at its literal placeholder default (`"change-me-in-production"`) — harmless today since nothing in the codebase actually reads this constant (confirmed: only definition, zero uses), but it's dead, confusing config that should either be wired up or removed. -- Two historical bad-deploy incidents (2026-05-15: a `NameError` from a forward-referenced Pydantic model, then an `IndentationError` in `orchestrator.py` that hit PM2's `max_restarts:10` ceiling and took the app down for ~42s) — both were syntax/import-level errors that a basic CI check (`python -m py_compile` or just running the test suite pre-deploy) would have caught before they reached prod. - ---- - -## Privacy / OAuth scope audit (summary) - -Full detail available on request; headline table: - -| Provider | Scope requested | Verdict | -|---|---|---| -| Google — identity | `openid email profile` | **Unused** — requested, never consumed | -| Google — Calendar | full `.../auth/calendar` | **Broader than needed** — only `calendar.events`-level operations are ever performed (events CRUD + freebusy on `primary` only; no calendar/ACL management) | -| Google — Docs/Drive | `documents` + `drive.file` | **Good** — already deliberately minimized (git history shows this was narrowed from a broader Drive scope), matches actual usage exactly | -| Google — Gmail | `gmail.readonly` + `gmail.send` | **Good** — already deliberately minimized from `gmail.modify`, matches actual usage (no delete/label/modify calls exist) | -| Google — Sheets | *(not requested — see P1 #10)* | Functionally broken, not a privacy issue | -| LinkedIn | `openid profile email w_member_social` | **Good** — tightly matched; unimplemented DM features correctly request no scope | -| Twitter/X | `tweet.read/write users.read dm.read/write offline.access` | **Good** — every scope maps to an implemented call (offline.access is unused for refresh but grants no extra surface) | -| Notion | *(page-level grants, no OAuth scope string)* | N/A — Notion's own model enforces least privilege | - -**Token storage:** plaintext (see P1 #9). **Recommendation:** narrow the Calendar scope to `calendar.events`, remove or wire up the unused Google identity scope, and consider encrypting `access_token`/`refresh_token` at rest (e.g. via Supabase Vault or an application-level envelope key) given what's at stake if that table ever leaks. - ---- - -## What's already working well - -Worth naming, since a report like this skews toward problems: - -- **144/144 backend tests pass**, no regressions. -- The **Zynd persona search/ranking logic** (`zynd_network.py`) is genuinely sophisticated — coverage-based multi-concept scoring, compound-word stemming ("cofounder" matching "founders"), honest "matched on X, missing Y" reasoning surfaced to the LLM rather than invented explanations, and a documented pool-floor workaround for a real registry quirk. This is not a "just wire up an API call" integration — real thought went into making "AI founders" actually work. -- **Google Docs/Gmail/Drive scoping is a good example of least-privilege done right**, and the code comments show it was *deliberately* narrowed over time, not accidentally broad. -- **The duplicate-meeting-proposal guardrail works correctly** (rejects a second proposal on a thread that already has one pending, per logs). -- **Published-page visibility defaults to "unlisted"** (link-only), not public, and the public persona-card endpoint explicitly omits `webhook_url`, `public_key`, and brief-document fields (`api/persona.py:200-209`) — sensible default-deny thinking. -- The A2A protocol design doc (`A2A.md`) is unusually rigorous for a "living design doc" — deterministic state machines, cryptographic accountability, explicit permission layering. The gap is between this design and a couple of the implementation details above (e.g. the sync/async client mixup), not the architecture itself. - ---- - -## Missing features / integration opportunities - -You asked me to think about what a broader "AI user" of this product would expect that isn't here yet. Current integration surface: Google (Calendar/Docs/Drive/Gmail/Sheets\*), LinkedIn, Twitter/X, Notion, Telegram, and the Zynd Network itself. Gaps, roughly in order of fit with the existing "professional-networking AI persona" positioning: - -1. **Slack / Discord** — the two places professional and builder communities actually live; a persona that can represent you in a community channel (not just DM-to-DM on Zynd) would be a natural extension of the existing A2A model. -2. **Microsoft 365 / Outlook + Teams** — a large fraction of the target audience (corporate professionals) isn't on Google Workspace at all; right now they can't connect Calendar/Email/Docs to their persona at all. -3. **A public booking/Calendly-style link** — the app already has real scheduling and free/busy logic (`group_calendar.py`, `smart_scheduling.py`); a shareable "book time with my persona" page is a small extension of what's already built, not a new capability. -4. **A standard, opt-in MCP endpoint for external AI agents** — right now the only way another AI can reach a persona is the proprietary Zynd A2A protocol. Exposing a scoped, per-user MCP server (the app already runs `ContextAware`/MCP internally — see P0 #4 on its current security posture) would let Claude, ChatGPT, or any other MCP-speaking agent interact with someone's persona directly, which is a very literal reading of "an AI user could use it." -5. **A lightweight CRM view of Zynd connections** — the product already tracks connections, threads, and match reasons; surfacing that as "people I've met through my persona" with notes/follow-up reminders is a small step from data that already exists. -6. **GitHub** — for a developer-leaning persona, surfacing activity/notifications would fit the existing "brief" concept (`daily_brief.py`) well. - -I'd treat this section as directional product input, not an audited finding — happy to go deeper on any of these if useful. - ---- - -## Next steps - -This report is findings-only, as requested. Suggested order if you want me to start fixing (I'll wait for your go-ahead before touching anything): - -1. **Memory layer secret** (P0 #1) — almost certainly a one-line `.env` change plus a restart; highest personalization impact for the least risk. -2. **Realtime broadcast client bug** (P0 #3) — small, well-isolated code fix (two call sites), and it's the likely root cause of the polling load (P2). -3. **Chat UX triple-format bug** (P1 #5) — the fix is a product decision (which tool wins) more than a hard technical problem; I can propose a specific change once you confirm the direction. -4. **Calendar scope gate + UUID validation** (P1 #7, #8) — mechanical, low-risk, fixes real recurring 500s. -5. **The native memory-corruption crashes** (P0 #2) — this one needs investigation (core dump / ASAN) before a fix is even possible to propose; flagging it as the thing most worth dedicated time given it's a live stability issue, not a UX one. - -Let me know which of these (if any) you'd like me to act on, and in what order. diff --git a/docs/persona/GROUPS.md b/docs/persona/GROUPS.md deleted file mode 100644 index 7715750..0000000 --- a/docs/persona/GROUPS.md +++ /dev/null @@ -1,252 +0,0 @@ -# Persona Groups — rollout notes - -Living checklist of things that must be done outside the code repo before -Persona Groups is production-ready, and the scope of each shipped phase. - -## What's in the repo - -### Phase 1 — human chat MVP (commit `3e3266a`) -- 3 tables in `backend/db/patch_add_persona_groups.sql` -- 14 routes in `backend/api/groups.py`, mounted at `/api/groups` -- Sidebar nav + 4 pages (`/dashboard/groups`, `/dashboard/groups/[id]`, - `/dashboard/groups/[id]/settings`, `/g/[slug]/[token]`) -- Supabase Realtime subscription for `persona_group_messages` filtered - by `group_id` - -### Phase 2 — @-mention persona dispatch (commit `8c77aaa`) -- `backend/agent/group_dispatch.py` — `extract_mentions`, - `resolve_mentions_to_members`, `dispatch_group_mention`. Routes - through `handle_user_message(is_external=True, …)` with group - permissions translated to the existing `external_permissions` shape. -- `POST /api/groups/{id}/messages` parses `@DisplayName` mentions on the - way in, fires a background `asyncio.create_task` for each resolved - member, and returns `mentioned_user_ids` for the thinking indicator. -- Replies written back as `channel='agent'` rows with - `metadata.reason='group_mention'`; realtime delivers them. -- Composer: `@`-typeahead picker (↑↓/Enter/Tab/Esc). Sent messages - render `@Name` as styled chips; "@you" gets its own highlight. -- "X's persona is thinking…" indicator above the composer, cleared - when the agent row arrives or after 60 s. - -### Phase 5 — discovery, domain auto-join, audit receipts -- `backend/db/patch_add_persona_group_discovery_audit.sql`: - - `persona_groups.join_domain` column (open groups only). - - `persona_group_audit_events` table for privacy-sensitive reads. - - RLS: affected users see their own receipts; owner/admin sees the - whole-group feed. -- 3 new routes: - - `GET /api/groups/discover` — open, non-archived groups the caller - isn't already in. Optional `query` for keyword filtering. - - `GET /api/groups/auto-join-candidates` — open groups whose - `join_domain` matches the caller's email domain. Returns the - invite token so the existing `/by-invite/{token}/join` path - handles the actual join. - - `GET /api/groups/{id}/activity?scope=me|all` — audit feed. - "me" returns the caller's own receipts (any member); "all" - requires owner/admin. -- Audit logger fires `brief_shared` events from the @-mention - dispatch path (only when brief content actually crossed the - boundary) and `calendar_queried` events from - `GET /availability` for each member whose calendar was read. - Self-checks are filtered out. -- Frontend: - - `/dashboard/groups` gains a Discover panel below the user's - groups, with an "Open to you via @yourdomain" auto-join section - when applicable. Search + one-click join. - - Group settings: `Auto-join domain` field appears when visibility - is open. - - Group settings: new `Activity` section showing the caller's - receipts. Owner/admin gets a toggle to view the whole-group feed. - -### Phase 4 — group memory / shared constraints (commit `4aa7f83`) -- New table `persona_group_constraints` with three kinds: - - `fact` — positive context the team has agreed on - - `rule` — guardrails to avoid (do NOT do X) - - `voice` — style/tone guidance -- `backend/db/patch_add_persona_group_constraints.sql` — table, index, - RLS (members read, service-role full). Soft-archive via - `archived_at` so removed rules stay queryable for audit. -- 4 new routes (`/{id}/constraints` GET / POST and - `/{id}/constraints/{cid}` PATCH / DELETE). Writes are owner/admin - only. `MAX_CONSTRAINTS_PER_GROUP = 20` enforced on POST — keeps - the LLM's working set focused. -- Dispatcher: `dispatch_group_mention` accepts a `group_constraints` - list. `_format_constraints_block` groups by kind and emits a - "Group rules — must-follow guardrails" prompt section. Rules - listed first (model honors earlier-listed instructions more - reliably), then facts, then voice. Unlike the brief, constraints - are NOT gated by `can_see_brief` — they're team-wide guardrails - that apply regardless of asker permissions. -- Right rail gains a Memory tab. Members read; owner/admin can add - one-liners (with inline help text per kind) and remove existing - rules. Optimistic remove with rollback on failure. - -### Phase 3b — calendar overlay + meeting proposals (commit `7742820`) -- `backend/agent/group_calendar.py`: - - Concurrent free/busy fan-out via Google Calendar's - `freebusy.query` (no event titles/attendees cross member - boundaries — only "busy from X to Y"). - - `find_common_slots` walks the window in N-minute steps with - business-hours + weekday gating, projected into the viewer's local - TZ via a passed offset. -- 2 new routes: - - `GET /api/groups/{id}/availability?start&end&duration_minutes&tz_offset_minutes` - — gated by the asker's `can_query_calendar`. Returns per-member - busy blocks and a list of `common_slots` (capped at 12). - - `POST /api/groups/{id}/meetings` — creates an event on the - asker's calendar with other members as attendees (Google handles - invite emails), then posts a `channel='system'` message to the - group with the meeting metadata. -- Right rail gains a `Schedule` tab: range + duration picker → "Find - slots" → list of "everyone is free" slots → click to propose → - modal asks for title/description/location, then sends invites. -- Members without a connected calendar are flagged in the UI and - excluded from the common-slot intersection so we never claim - "everyone is free" based on missing data. - -### Phase 3a — shared group brief (commit `45bbb88`) -- New columns on `persona_groups`: `brief_doc_id`, `brief_doc_url` - (`backend/db/patch_add_persona_group_brief.sql`). -- 3 new routes: - - `POST /api/groups/{id}/brief/init` (owner-only) — creates a Google - Doc in the owner's Drive, seeded with the group description. - - `GET /api/groups/{id}/brief` (members) — live-fetched body. - - `PATCH /api/groups/{id}/brief` (owner/admin) — replaces the body. -- Dispatcher injection: when `can_see_brief` is on for the asker and - the group has a brief, the doc body is pre-fetched once per turn - (in `_spawn_mention_dispatch`) and threaded into each target's - prompt via `dispatch_group_mention(group_brief_content=…)`. The - prefix labels it explicitly as "shared group brief" so the LLM - doesn't confuse it with the persona's own per-user brief. -- UI: chat right rail is now a tabbed pane (`People` / `Brief`). - Brief tab shows the doc body, an "Open in Docs" link, a Reload - button, and an Edit-in-place flow for owner/admin. - -### Phase 2 polish — follow-ups completed -- **Brief gating is now hard, not just behavioral.** `_format_user_brief` - takes a new `redact_brief` flag. When the asker doesn't have - `can_see_brief`, the brief Google Doc body is stripped from the - system prompt entirely — it never enters the LLM's context window - for that turn. The dispatch prefix also threads an explicit hint so - the LLM doesn't leak equivalents from memory. -- **Per-member permission toggles** in the group settings page. Each - non-owner row has an expandable Permissions panel with three - switches: see briefs of mentioned members, check their calendars, - and post in the group (mute toggle). Auto-saves optimistically with - rollback on PATCH failure. -- **Member cap** (`MAX_GROUP_MEMBERS = 15`) enforced in `add_member` - and `join_via_invite` — keeps the @-mention dispatch fan-out and - future calendar overlay scope bounded. -- **Owner transfer**: `POST /api/groups/{id}/transfer-owner` plus a - "Make owner" action on each admin's row in settings. The previous - owner is demoted to admin in the same call. -- **Group keypair derivation foundation**: `derive_group_seed` and - `derive_group_keypair` in `zynd_identity.py` (domain-separated from - agent derivation), plus `build_group_context_claim` / - `verify_group_context_claim` in `group_dispatch.py`. The - cryptographic primitives are ready for cross-instance dispatch; the - routing layer that uses them is still pending (see below). - ---- - -## Pending end-to-end (must do before users see groups) - -These are operational items that the code can't do for itself. - -### 1. Apply the schema via Prisma -The canonical schema is now `prisma/schema.prisma` — the `backend/db/*.sql` -patches are historical and shouldn't be applied directly on a Prisma-managed -project (they'd duplicate what Prisma generates). See `prisma/README.md` -for the full flow. - -```bash -cd webapp -npm install # picks up prisma devDep -# Edit webapp/.env with DATABASE_URL + DIRECT_URL (see prisma/.env.example) - -npm run prisma:validate # sanity-check the schema -npm run prisma:migrate:deploy # tables + indexes + enums -npm run prisma:policies # RLS + realtime publication -``` - -If the existing DB already has these tables (from the legacy .sql patches), -adopt under Prisma without re-creating — see `prisma/README.md` → -"Adopting on an existing database". - -What it creates (all 19 public tables): -- Core: `api_tokens`, `chat_messages`, `persona_agents`, `dm_threads`, - `dm_messages`, `agent_tasks`, `a2a_tasks`, `pending_approvals`, - `telegram_links`, `telegram_chat_history`, `linkedin_profiles`, - `brief_todos`, `outbound_callbacks`, `callback_results` -- Groups (phases 1–5): `persona_groups`, `persona_group_members`, - `persona_group_messages`, `persona_group_constraints`, - `persona_group_audit_events` - -### 2. Enable Realtime on `persona_group_messages` -The chat view subscribes to `postgres_changes` on this table. Until -realtime is enabled, new messages only appear on page refresh. - -Supabase Studio → **Database** → **Replication** → toggle on for -`persona_group_messages`. (`persona_group_members` and `persona_groups` -don't need realtime for phase 1 or 2.) - -### 3. Confirm RLS works for your service-role key -The API uses the service-role key, which bypasses RLS. The RLS policies -are defense-in-depth for any direct frontend reads (notably the realtime -channel). Sanity-check that: -- A signed-in member CAN read messages via the realtime channel -- An anonymous client CANNOT subscribe to a group they're not a member of - ---- - -## Still pending (code, lower priority) - -### Cross-instance `group_context` dispatch -Phase 2 dispatch only routes through the in-process orchestrator. When -persona members live on different Zynd backends, dispatch needs to: -- Build a signed `group_context` claim via the helpers already shipped - (`build_group_context_claim`). -- Send it as part of the A2A v3 envelope to the target's host. -- The receiving host verifies via `verify_group_context_claim`, - resolves `asker_agent_id` against its own group roster - (defense-in-depth — the signature alone doesn't prove membership), - and then calls the same `dispatch_group_mention` logic. - -The orchestrator/A2A wrapping is the missing piece. Crypto + permission -mapping are in place. - -### Asker- vs target-side permission model -The current model is "asker permissions": the group owner decides which -members are allowed to see briefs / calendars of anyone they @mention. -A more conservative model is "target permissions": each member toggles -whether their *own* brief/calendar is sharable inside the group. We can -add a second permission key (e.g. `share_my_brief`) on top of the -existing one and intersect at dispatch time. Worth revisiting if real -users push back on the asker-only model. - ---- - -## Beyond phase 5 — possible next-up - -Phases 1–5 are in main. Things worth doing next if the feature gets -traction: - -- **Cross-instance dispatch** — wire the existing - `build_group_context_claim` / `verify_group_context_claim` helpers - into the outbound A2A v3 envelope so personas hosted on different - Zynd backends can answer @-mentions in shared groups. -- **Target-side privacy preferences** — today permissions live on the - ASKER ("X is trusted to see briefs"). Adding `share_my_brief` / - `share_my_calendar` on each member would let users opt out of being - read regardless of the asker's permissions. -- **Brief redaction layers** — same brief, different views per group. - Useful when one person is in multiple groups and wants engineering - to see different details than HR. -- **Domain auto-join: automatic** — currently surfaces as a CTA on - the discover page. A signup-time hook could auto-add new users - matching `join_domain` rules to their team's group. -- **Activity export** — let users download their audit log as JSON or - CSV for a self-managed privacy record. -- **Group-level approvals** — meetings booked through `POST /meetings` - could optionally require N member acknowledgments before the calendar - event is actually created (reuses the existing approvals indicator). diff --git a/docs/persona/IMPLEMENTATION_MEMORY_BRIEF.md b/docs/persona/IMPLEMENTATION_MEMORY_BRIEF.md deleted file mode 100644 index 9a04f56..0000000 --- a/docs/persona/IMPLEMENTATION_MEMORY_BRIEF.md +++ /dev/null @@ -1,685 +0,0 @@ -# Implementation Plan — Consolidate user-info, retire the Google-Doc Brief - -Status: **approved architecture**. Implement in phases. Each phase is -independently deployable; stop and verify at the end of each phase. - -## Goal & boundary (do not re-litigate during implementation) - -- **Conversation transcript stays in `agent-persona` (`chat_messages`).** Do NOT move it to the memory layer. -- **Memory layer stays a separate fact/recall/matching engine** (`/home/ubuntu/memory-layer`). -- **The Brief becomes a plain text field** on `persona_agents.brief_content` (column already exists, migration `db/migrations/0001_persona_fts/migration.sql`). Google Docs is retired for the brief. -- **Memory layer gains one new endpoint**: private fact "declare" (Phase 3). - -Out of scope (do NOT touch): group briefs (`persona_groups.brief_doc_*`, `api/groups.py`, `agent/group_dispatch.py`). Group briefs remain Google Docs for now. - ---- - -## Current state map (why we're changing this) - -| Path | Endpoint / code | Store | -|---|---|---| -| Onboarding brief (Google Doc "My brief") | `webapp/.../onboarding/brief/page.tsx` → `POST /api/brief/create` (`api/brief.py`) | `auth.users.user_metadata.brief_doc` | -| Dashboard brief editor (Google Doc "Brief — {name}") | `BriefPanel.tsx` → `GET/PATCH /api/persona/{id}/brief`, `POST .../brief/init` (`api/persona.py`) | `persona_agents.brief_doc_id/url/revision_id` | -| Agent reads brief for prompt | `orchestrator._format_user_brief` → `_fetch_brief_doc_content` | reads `persona_agents.brief_doc_id` | -| MCP tools `read/append/replace/clear_my_brief` | `mcp/tools/brief.py` → `persona_manager` | `persona_agents.brief_doc_id` + Google Docs | -| Todo extraction from brief | `brief_watcher.py`, `api/todos.py` | `persona_agents.brief_doc_id` + Drive polling | -| Duplicate brief code in memory-layer | `memory-layer/app/services/persona_brief.py` | same `persona_agents.brief_doc_id` | - -The `brief_content` TEXT column already exists and is written by -`save_brief_content`/`brief_watcher` as a *mirror* of the Google Doc body. We -promote it to the single source of truth. - ---- - -## Phase 0 — Backfill `brief_content` (RUN FIRST, before deleting Google-Docs read paths) - -Because Phase 1 removes the Google-Docs *read* path, you must copy any existing -brief text into `brief_content` while `read_document` still works. - -1. Create `backend/scripts/backfill_brief_content.py` (sibling of existing `scripts/`). It must: - - - Load Supabase via `config.get_supabase()`. - - For each `persona_agents` row where `brief_doc_id` is not null and `brief_content` is null/empty: call `mcp.tools.google.docs.read_document(user_id, document_id=brief_doc_id)`; on `success`, write `content` into `persona_agents.brief_content`. - - For each `auth.users` row with `user_metadata.brief_doc.doc_id` set (via the Supabase admin API, same pattern as `api/brief.py:_read_user_metadata`): read that doc and write into `persona_agents.brief_content` for the matching `supabase_user_id` (join on `persona_agents.user_id`). Use the admin `GET /auth/v1/admin/users` listing (or iterate known users) — simplest: read `user_metadata.brief_doc` the same way `api/brief.py:my_brief` does, keyed off `persona_agents.user_id`. - - Log every user migrated and every skip/failure; never raise. - - Minimal skeleton (fill in the Supabase admin-user listing to match how your - env exposes it): - - ```python - # backend/scripts/backfill_brief_content.py - import config - from mcp.tools.google.docs import read_document - - def main(): - sb = config.get_supabase() - # 1. persona_agents.brief_doc_id -> brief_content - rows = sb.table("persona_agents").select("user_id,brief_doc_id,brief_content").not_.is_("brief_doc_id", "null").execute() - for r in rows.data or []: - if (r.get("brief_content") or "").strip(): - continue - try: - got = read_document(user_id=r["user_id"], document_id=r["brief_doc_id"]) - if got.get("success") and (got.get("content") or "").strip(): - sb.table("persona_agents").update({"brief_content": got["content"].strip()}).eq("user_id", r["user_id"]).execute() - print(f"[backfill] persona {r['user_id']}: migrated {len(got['content'])} chars") - else: - print(f"[backfill] persona {r['user_id']}: read failed: {got.get('error')}") - except Exception as e: - print(f"[backfill] persona {r['user_id']}: ERROR {e}") - # 2. user_metadata.brief_doc -> brief_content (only for users with a persona row) - # Use the admin API as in api/brief.py:_read_user_metadata to get user_metadata.brief_doc - # for each persona_agents.user_id, read the doc, and upsert brief_content if still empty. - print("[backfill] done") - - if __name__ == "__main__": - main() - ``` - -2. Run it once in the prod backend env: `python -m scripts.backfill_brief_content`. -3. Verify: `SELECT count(*) FROM persona_agents WHERE brief_content IS NOT NULL AND brief_content <> '';` - -Do NOT null out `brief_doc_id` / `user_metadata.brief_doc` yet — keep them until -Phase 1 is deployed and verified (rollback path). - ---- - -## Phase 1 — Backend: brief becomes `persona_agents.brief_content` - -### 1a. `backend/agent/persona_manager.py` - -**(i)** In `get_persona_status` (currently lines ~554–576), add one key to the -returned dict: - -```python - "brief_content": persona.get("brief_content"), -``` - -**(ii)** Replace `get_brief` (lines ~670–708) with: - -```python -def get_brief(user_id: str) -> dict: - """Return the persona's brief text (single source of truth: persona_agents.brief_content). - - The brief is now a plain text field, not a Google Doc. `exists` is True - whenever the persona is deployed (the field always exists; it may be empty). - """ - persona = get_persona_status(user_id) - if not persona.get("deployed"): - raise ValueError("No active persona.") - - content = (persona.get("brief_content") or "").strip() - return { - "exists": True, - "content": content, - "fallback_description": persona.get("description") or "", - } -``` - -**(iii)** Replace `save_brief_content` (lines ~710–732) with: - -```python -def save_brief_content(user_id: str, content: str) -> dict: - """Store the persona's brief body as plain text on persona_agents.brief_content.""" - persona = get_persona_status(user_id) - if not persona.get("deployed"): - raise ValueError("No active persona.") - - sb = _get_supabase() - sb.table("persona_agents").update({ - "brief_content": content or None, - }).eq("user_id", user_id).execute() - - logger.info(f"[persona] saved brief for {user_id} ({len(content or '')} chars)") - return {"success": True, "content": content} -``` - -**(iv)** Replace `init_brief_doc` (lines ~621–668) with a no-op ensure (keeps the -name so callers don't churn; the brief needs no creation step anymore): - -```python -def init_brief_doc(user_id: str) -> dict: - """Legacy shim. The brief is a text field now — nothing to create. - - Returns the current brief state so existing callers (the old /brief/init - endpoint, MCP _ensure_brief_doc) behave as a read instead of a Google call. - """ - persona = get_persona_status(user_id) - if not persona.get("deployed"): - raise ValueError("No active persona — create a persona before using the brief.") - return { - "doc_id": None, - "url": "", - "created": False, - "exists": True, - "content": (persona.get("brief_content") or "").strip(), - } -``` - -Remove the `from mcp.tools.google.docs import ...` inside these functions (they no -longer touch Google). - -### 1b. `backend/agent/orchestrator.py` - -**(i)** Delete `_BRIEF_DOC_CACHE`, `_BRIEF_DOC_CACHE_TTL_SECONDS`, and -`_fetch_brief_doc_content` (lines ~1671–1695). - -**(ii)** Replace `_format_user_brief` (lines ~1697–1764) — change only the source -of the description text; keep the `redact_profile` / `redact_brief` logic and the -profile-field rendering intact: - -```python -def _format_user_brief( - persona: dict, - redact_profile: bool = False, - redact_brief: bool = False, - user_id: str | None = None, -) -> str: - # Source priority: brief_content (plain text) -> persona.description. - # `user_id` is kept for call-signature compatibility; it is unused now that - # the brief is a field rather than a Google Doc fetch. - brief_text = None - if not redact_brief: - brief_text = (persona.get("brief_content") or "").strip() - - desc = (brief_text or persona.get("description") or "").strip() - profile = persona.get("profile") or {} - - lines = [] - if desc: - lines.append(desc) - - if redact_profile: - return "\n".join(lines) if lines else "(no profile details set yet)" - - # ... (existing profile_lines block unchanged: title/organization/location/ - # interests/socials) ... - - return "\n".join(lines) if lines else "(no profile details set yet)" -``` - -(The `_build_system_prompt` caller already passes `persona = get_persona_status(...)` -which now includes `brief_content`, so no change there.) - -### 1c. `backend/api/persona.py` - -**(i)** Remove `init_brief_doc` from the import list (lines ~19–28) if it's no longer -referenced by an endpoint; keep `get_brief`, `save_brief_content`, `get_persona_status`. - -**(ii)** Replace the brief endpoints (lines ~373–440): - -```python -# ── Brief (plain text) ───────────────────────────────────────────── -@router.get("/{user_id}/brief") -async def read_brief(user_id: str): - """Return the persona's brief text (single field, no Google Docs).""" - try: - return get_brief(user_id) - except ValueError as e: - raise HTTPException(status_code=400, detail=str(e)) - except Exception as e: - raise HTTPException(status_code=500, detail=str(e)) - - -class BriefSave(BaseModel): - content: str - - -@router.patch("/{user_id}/brief") -async def save_brief(user_id: str, req: BriefSave): - """Store the persona's brief body as plain text.""" - try: - result = save_brief_content(user_id, req.content) - if not result.get("success"): - raise HTTPException(status_code=502, detail=result.get("error") or "Couldn't save the brief.") - return result - except ValueError as e: - raise HTTPException(status_code=400, detail=str(e)) - except Exception as e: - raise HTTPException(status_code=500, detail=str(e)) -``` - -Delete the `POST /{user_id}/brief/init` endpoint and `_classify_google_error` -(no longer used). - -### 1d. Delete `backend/api/brief.py` and unregister it - -- Delete the file `backend/api/brief.py`. -- In `backend/main.py`: remove `from api.brief import router as brief_router` (line 26) and `app.include_router(brief_router, prefix="/api/brief", tags=["Brief"])` (line 132). - -### 1e. `backend/mcp/tools/brief.py` (the chat tools the LLM calls) - -These are the durable-fact write tools. Point them at `brief_content`: - -**(i)** `_ensure_brief_doc` (lines ~32–101) — simplify to a persona check (no Google): - -```python -def _ensure_brief_doc(user_id: str) -> dict: - """Return {"ok": True} if the user has a deployed persona, else a no_persona error.""" - from agent import persona_manager - try: - persona = persona_manager.get_persona_status(user_id) - if persona.get("deployed"): - return {"ok": True} - return { - "ok": False, - "code": "no_persona", - "message": "You haven't deployed a persona yet. Open the Zynd dashboard and finish onboarding.", - } - except Exception as e: - logger.warning(f"[brief] ensure failed: {e}") - return {"ok": False, "code": "no_persona", "message": "Couldn't confirm your persona is ready. Try again."} -``` - -**(ii)** `read_my_brief` (lines ~104–149) — return the field content: - -```python -def read_my_brief(user_id: str) -> dict: - from agent import persona_manager - try: - result = persona_manager.get_brief(user_id) - except ValueError as e: - return friendly_error("read your brief", e) - except Exception as e: - logger.exception(f"[brief] read_my_brief failed: {e}") - return friendly_error("read your brief", e) - - return { - "success": True, - "exists": result.get("exists", True), - "content": result.get("content") or "", - "fallback_description": result.get("fallback_description") or "", - } -``` - -**(iii)** `append_to_my_brief` (lines ~152–206) — append to the field: - -```python -def append_to_my_brief(user_id: str, text: str) -> dict: - if not isinstance(text, str) or not text.strip(): - return friendly_error_message( - "add to your brief", "Nothing to append — `text` was empty.", - hint="Tell me what you'd like me to add.", - ) - ensured = _ensure_brief_doc(user_id) - if not ensured.get("ok"): - return {"success": False, "error": ensured["message"], "code": ensured["code"]} - - from agent import persona_manager - try: - current = persona_manager.get_brief(user_id) - except ValueError as e: - return friendly_error("add to your brief", e) - - existing = current.get("content") or "" - body = text if text.endswith("\n") else text + "\n" - new_content = (existing.rstrip() + "\n\n" + body.rstrip() + "\n") if existing else body - result = persona_manager.save_brief_content(user_id, new_content) - if not result.get("success"): - return friendly_error_message("add to your brief", result.get("error") or "Append failed.") - return {"success": True, "appended": text.strip()} -``` - -**(iv)** `replace_my_brief` (lines ~209–249) — drop the `doc_id`/`url` in the return; -keep calling `persona_manager.save_brief_content`. Remove the `get_persona_status` -lookup for `brief_doc_url`: - -```python -def replace_my_brief(user_id: str, content: str) -> dict: - if content is None: - content = "" - ensured = _ensure_brief_doc(user_id) - if not ensured.get("ok"): - return {"success": False, "error": ensured["message"], "code": ensured["code"]} - - from agent import persona_manager - try: - result = persona_manager.save_brief_content(user_id, content) - except ValueError as e: - return friendly_error("replace your brief", e) - except Exception as e: - logger.exception(f"[brief] replace_my_brief failed: {e}") - return friendly_error("replace your brief", e) - - if not result.get("success"): - return friendly_error_message("replace your brief", result.get("error") or "Replace failed.") - return {"success": True, "content": content} -``` - -`clear_my_brief` and `add_todo` are unchanged. - -### 1f. `backend/agent/brief_watcher.py` (todo extraction from the brief) - -Rewire `_poll_all` / `_poll_one` to read `brief_content` instead of Drive. - -Replace the class body polling (lines ~218–293) with a hash-based sweep: - -```python - def _poll_all(self): - sb = self._supabase() - rows = ( - sb.table("persona_agents") - .select("user_id,brief_content") - .eq("active", True) - .not_.is_("brief_content", "null") - .execute() - ) - for row in rows.data or []: - try: - self._poll_one(sb, row) - except Exception as e: - logger.warning(f"[brief_watcher] Poll failed for user {row.get('user_id')}: {e}") - - def _poll_one(self, sb, row: dict): - user_id = row["user_id"] - content = (row.get("brief_content") or "").strip() - if not content: - return - digest = hashlib.sha256(content.encode()).hexdigest() - if self._last_seen.get(user_id) == digest: - return # unchanged since last extraction - - llm_titles = extract_todo_titles_llm(content) - if llm_titles is None: - titles = extract_todo_titles(content) - extractor = "regex_fallback" - else: - titles = llm_titles - extractor = "llm" - if titles: - self._upsert_todos(sb, user_id, titles) - self._last_seen[user_id] = digest - logger.info(f"[brief_watcher] {user_id}: {len(titles)} todos via {extractor}") -``` - -Add `import hashlib` and an instance dict `self._last_seen: dict[str, str] = {}` in -`__init__`. Keep `extract_todo_titles`, `extract_todo_titles_llm`, `_upsert_todos` -unchanged. Remove the `brief_doc_revision_id` write and the Google `build`/`get_google_creds` imports. - -### 1g. `backend/api/todos.py` — `/extract` endpoint - -Replace the doc-fetch block (lines ~86–103) with a field read: - -```python - persona = get_persona_status(user_id) - if not persona.get("deployed"): - raise HTTPException(status_code=404, detail="No active persona.") - content = (persona.get("brief_content") or "").strip() - if not content: - raise HTTPException( - status_code=400, - detail="Your brief is empty — add some text first.", - ) - - titles = extract_todo_titles_llm(content) - extractor = "llm" - if titles is None: - titles = extract_todo_titles(content) - extractor = "regex_fallback" -``` - -Remove the `read_document` import in that function. - -### Phase 1 verification - -```bash -cd /home/ubuntu/agent-persona/backend -python -m pytest tests/test_brief_tools.py -q # after updating tests (Phase 6) -# Manual: GET /api/persona/{id}/brief returns {exists, content, fallback_description} -# PATCH /api/persona/{id}/brief with {content} persists; GET reflects it -# Chat "add to my brief: X" -> read_my_brief shows X; /api/brief/* now 404s -``` - ---- - -## Phase 2 — Frontend - -### 2a. `webapp/src/components/BriefPanel.tsx` - -- Remove the `exists`-gating branch (`if (!brief?.exists)`) and `handleCreate`, plus the `BriefState.doc_id/url/title/error` fields and the "Open in Google Docs ↗" link. -- `BriefState` becomes `{ content?: string; fallback_description?: string }`. -- `initialIsSynced` in `` becomes `!!brief.content && brief.content.length > 1`; change the label "Synced from Google Docs" (line 369) to "Saved". -- Update the empty-state copy: no mention of Google Docs ("Your brief is the long-form context your agent uses to represent you. Edit here any time."). -- Keep the `apiPatch('/api/persona/{id}/brief', { content })` save path unchanged. -- Remove the now-dead `friendlySaveError`/`summarizeRawGoogleError` Google-mapping (optional; harmless to keep). - -### 2b. `webapp/src/app/onboarding/brief/page.tsx` - -- Remove `postCreateBrief` and the `POST /api/brief/create` + Google OAuth redirect logic. -- Replace `handleCreate` with a no-op that just marks onboarding complete (no doc to create), or seed `brief_content` via `PATCH /api/persona/{id}/brief` with a starter template. Simplest: keep the screen as an informational step and make the button set `brief_created: true` then continue. Keep `handleSkip`. -- Remove the `created` view that links to Google Docs. - -### 2c. `webapp/src/app/dashboard/settings/accounts/page.tsx` - -- The "brief" connector (lines ~516–530 and the `handleConnect("brief")` branch) only triggered Google `docs,calendar` OAuth. Remove the "brief" connector card; keep `calendar` and `email` Google connectors. Remove `brief` from `ConnId` and the `googleSiblingsNote` self-type (`brief`), so brief is no longer listed as a Google-scope sibling. - -### Phase 2 verification - -```bash -cd /home/ubuntu/agent-persona/webapp -npm run build # REQUIRED before restart (AGENTS.md) -``` - -Manual: Dashboard → "Your brief" loads an editable text box (no create step); save persists; onboarding brief step no longer calls `/api/brief/create`. - ---- - -## Phase 3 — Memory layer: private fact "declare" endpoint (`/home/ubuntu/memory-layer`) - -### 3a. `app/services/findability.py` — generalize declare for private memory - -Add a predicate→entity-type map for private predicates and a `declare_private` -function (alongside the existing `declare`): - -```python -# Private-memory declarable predicates -> entity family. Literal/enum predicates -# (has_age, is_seeking, open_to) are intentionally excluded — they need literal -# storage, not entity resolution. -PRIVATE_DECLARE_ENTITY_TYPE: dict[str, str] = { - "is_building": "project_venture", - "is_working_on": "project_venture", - "is_creating": "artifact_creative", - "wants_to_preserve": "concept_topic", - "is_learning": "skill_domain", - "has_expertise_in": "skill_domain", - "has_skill": "skill_technical", - "intends_to": "intent_project", - "is_preparing_for": "intent_project", - "fears": "concept_topic", - "believes": "belief_opinion", - "values": "belief_value", - "recently_changed_stance_on": "belief_opinion", - "has_aesthetic": "belief_opinion", - "is_navigating": "concept_topic", - "is_constrained_by": "concept_topic", - "is_frustrated_by": "concept_topic", - "has_been_wronged": "concept_topic", - "is_transitioning": "concept_topic", - "is_experiencing": "concept_topic", - "is_processing": "concept_topic", - "is_rediscovering": "concept_topic", - "has_unsolved_problem": "concept_topic", - "has_collaborator": "collaborator", - "is_responsible_for": "project_assignment", - "is_advocating_for": "concept_topic", - "is_in_conflict_with": "adversary", - "is_inspired_by": "influence", - "is_located_in": "place_physical", - "is_affiliated_with": "place_institutional", - "has_language_context": "concept_topic", - "is_motivated_by": "concept_topic", -} - -PRIVATE_DECLARED_CONFIDENCE = 0.97 - - -async def declare_private(pool: asyncpg.Pool, user_id: str, predicate: str, value: str) -> None: - """User explicitly adds a PRIVATE memory fact (never matched/public).""" - if predicate not in PRIVATE_DECLARE_ENTITY_TYPE: - raise ValueError(f"{predicate!r} is not declarable as private memory") - value = (value or "").strip() - if not value: - raise ValueError("value is required") - - entity_type = PRIVATE_DECLARE_ENTITY_TYPE[predicate] - async with pool.acquire() as conn: - async with conn.transaction(): - entity_id = await resolve_entity(conn, user_id, value, entity_type) - existing = await conn.fetchrow( - """SELECT id FROM assertions WHERE user_id = $1 AND predicate = $2 - AND object_entity_id = $3 AND valid_until IS NULL LIMIT 1""", - user_id, predicate, entity_id) - if existing: - await conn.execute( - """UPDATE assertions SET source = 'declared', - confidence = $2, version = version + 1 WHERE id = $1""", - existing["id"], PRIVATE_DECLARED_CONFIDENCE) - else: - await conn.execute( - """INSERT INTO assertions - (user_id, predicate, object_entity_id, confidence, source_system, - source, is_public, decay_fn) - VALUES ($1, $2, $3, $4, 'user_confirmed', 'declared', false, $5)""", - user_id, predicate, entity_id, PRIVATE_DECLARED_CONFIDENCE, - decay_fn_for(predicate)) - # Private memory is not matched, so no recompute_user_embeddings is needed. -``` - -Note: `decay_fn_for` is already imported in this module; `resolve_entity` too. - -### 3b. `app/main.py` — add the route - -```python -@app.post("/me/memory/declare") -async def declare_memory_fact(req: DeclareRequest, user_id: str = Depends(current_user)) -> dict: - """User explicitly adds a PRIVATE memory fact (stays private, never matched).""" - from app.services.findability import declare_private - try: - await declare_private(get_pool(), user_id, req.predicate, req.value) - except ValueError as exc: - raise HTTPException(status_code=400, detail=str(exc)) from exc - return {"status": "declared", "predicate": req.predicate, "value": req.value} -``` - -`DeclareRequest` already exists in `app/models.py` and is imported in `main.py`. - -### 3c. Delete `app/services/persona_brief.py` (duplicate brief code) — after confirming -nothing imports it (`grep -rn persona_brief app/`). It is a slim port of the -Google-Docs brief and is now dead. - -### 3d. (Optional, Phase 5) `app/services/persona_ingest.py` — no change needed; it -already ingests the persona profile into memory. If you also want the brief prose as -facts, add a `brief_seeded` ingest in Phase 5, not here. - -### Phase 3 verification (memory-layer repo) - -```bash -cd /home/ubuntu/memory-layer -uv run pytest -q tests/test_findability.py tests/test_pipeline.py -# Manual: POST /me/memory/declare {"predicate":"is_working_on","value":"Acme launch"} -# then GET /me/graph shows the new private fact with source=declared, is_public=false -``` - ---- - -## Phase 4 — `agent-persona` memory client: `declare_fact` - -Add to `backend/agent/memory_client.py`: - -```python -async def declare_fact(user_id: str, predicate: str, value: str) -> bool: - """Write a user-authored PRIVATE memory fact directly (structured predicate/value). - - Distinct from ingest_turns (which runs async extraction) — this is an - explicit, high-confidence declaration for the editable memory surface. - """ - if not is_enabled(): - return False - token = _make_jwt(user_id) - try: - async with _client() as client: - resp = await client.post( - "/me/memory/declare", - json={"predicate": predicate, "value": value}, - headers={"Authorization": f"Bearer {token}"}, - ) - resp.raise_for_status() - return True - except Exception as exc: - logger.debug("[memory] /me/memory/declare failed for %s: %s", user_id, exc) - return False -``` - -Optionally expose an MCP tool `remember_this_structured(user_id, predicate, value)` -in `backend/mcp/tools/memory.py` for the LLM to write discrete private facts. Do NOT -change the existing free-text `remember_this` (ingest-based) — both are valid. - -### Phase 4 verification - -`python -m pytest tests/test_memory_fact_ref.py -q` (still green), plus a manual -`declare_fact` round-trip if the memory layer (Phase 3) is deployed. - ---- - -## Phase 5 — Optional: seed brief prose into memory facts - -In `backend/agent/persona_manager.py:save_brief_content`, after the DB write, -fire-and-forget an ingest so the brief's substance is recallable via `/context`: - -```python - # Best-effort: make the brief's substance semantically recallable in the - # memory layer (narrative stays verbatim in the prompt). - import asyncio - from agent.memory_client import ingest_turns, is_enabled - if is_enabled() and content and len(content.strip()) >= 40: - async def _seed(): - await ingest_turns( - user_id=user_id, - turns=[{"role": "user", "content": content.strip()}], - source_system="brief_seeded", - ) - try: - asyncio.get_running_loop() - except RuntimeError: - asyncio.run(_seed()) - else: - asyncio.create_task(_seed()) -``` - -(None of this blocks the save. `brief_seeded` mirrors the existing -`persona_seeded` tag in the memory layer.) - ---- - -## Phase 6 — Tests - -- `backend/tests/test_brief_tools.py` — update the monkeypatched `persona_manager` - stubs to the new shapes: `get_brief` returns `{exists, content, fallback_description}`, - `save_brief_content` returns `{success, content}`, `init_brief_doc` returns the shim dict. - Remove `brief_doc_id` expectations. Keep `_ensure_brief_doc` no_persona / google_unavailable - tests but drop the google path (now `no_persona` only). -- `backend/tests/test_action_summary.py` already patches `ingest_conversation` — unaffected. -- Add a test for the new `_format_user_brief` reading `brief_content` (assert description - from `brief_content` wins over `persona.description`, and empty `brief_content` falls back). - -### Phase 6 verification - -```bash -cd /home/ubuntu/agent-persona/backend && python -m pytest -q -``` - ---- - -## Deploy order (prod + dev channels per AGENTS.md) - -1. Run Phase 0 backfill in **prod** env first (before any code deploy that removes read_document). -2. Deploy code changes: commit + push `main` → dev copy `git pull`, `pip install -r requirements.txt` (only if changed), **`npm run build`** (always after webapp changes), `pm2 restart api-dev web-dev` → smoke test dev. -3. Deploy memory-layer (Phase 3) to its own service (`api.zynd.ai`) independently. -4. Prod copy: `git pull`, `npm run build`, `pm2 restart api web`. -5. After prod verified, run cleanup (optional): null `brief_doc_id/url/revision_id` on `persona_agents` and clear `user_metadata.brief_doc`, and drop the now-unused columns in a later migration. - -## Rollback - -- `persona_agents.brief_doc_id` / `brief_doc_url` columns are left in place through - Phases 1–5, so reverting to the Google-Docs path is a code revert only. Only the - final cleanup step removes them. diff --git a/docs/persona/architecture.md b/docs/persona/architecture.md deleted file mode 100644 index 5b081b1..0000000 --- a/docs/persona/architecture.md +++ /dev/null @@ -1,375 +0,0 @@ -# Zynd AI Persona Platform — Technical Architecture (v2) - -## Overview - -Zynd is a multi-tenant AI agent platform where users create autonomous "personas" that live on the Zynd AI Network. Each persona can be discovered by other agents, receive messages, and take actions on behalf of its owner (posting tweets, scheduling calendar events, querying Notion, etc.). - -This document covers the v2 architecture after migrating from the legacy DID/PolygonID system to the Ed25519/agdns identity system. - ---- - -## Identity Architecture - -### Hierarchical Deterministic (HD) Key Derivation - -All persona identities on the platform are derived from a single **developer keypair** using HD derivation: - -``` -Developer Key (admin-created, stored at ~/.zynd/developer.json) - │ - ├── Index 0 → User A's persona keypair → agdns:a1b2c3... - ├── Index 1 → User B's persona keypair → agdns:d4e5f6... - ├── Index 2 → User C's persona keypair → agdns:g7h8i9... - └── ... -``` - -**Derivation algorithm:** -``` -agent_seed = SHA-512(developer_seed || "agdns:agent:" || index_as_4_byte_big_endian)[:32] -agent_id = "agdns:" + SHA-256(public_key_bytes).hex()[:32] -``` - -**Why HD derivation?** -- No private keys stored in the database — only the derivation index -- Deterministic reconstruction: given the developer key + index, the exact same keypair is reproduced -- Server restarts don't lose identity — all personas are rehydrated from DB indexes -- Single root of trust for the entire platform - -### Identity Flow - -``` -Admin Setup (one-time): - zynd init → creates ~/.zynd/developer.json (Ed25519 keypair) - -User Creates Persona: - 1. Backend allocates next derivation_index from persona_agents table - 2. Derives Ed25519 keypair: SHA-512(dev_seed || "agdns:agent:" || index)[:32] - 3. Computes agent_id: "agdns:" + SHA-256(pubkey).hex()[:32] - 4. Registers on zns01.zynd.ai with signature auth + "persona" tag - 5. Saves to persona_agents table (user_id, agent_id, index, public_key, ...) - 6. Starts heartbeat for this agent -``` - ---- - -## Database Schema - -### Tables - -| Table | Purpose | -|-------|---------| -| `api_tokens` | OAuth access/refresh tokens per provider per user | -| `chat_messages` | Conversation persistence (future use, currently in-memory) | -| `persona_agents` | Maps users to their Zynd Network agent identities | -| `dm_threads` | Direct messaging threads between agents | -| `dm_messages` | Individual messages within DM threads | - -### persona_agents (core identity table) - -```sql -persona_agents ( - user_id UUID PK → auth.users(id) - agent_id TEXT UNIQUE -- agdns:... format - derivation_index INTEGER UNIQUE -- HD derivation index - public_key TEXT -- ed25519:... format - name TEXT - description TEXT - capabilities JSONB -- ["calendar_management", "social_media", ...] - webhook_url TEXT - active BOOLEAN - created_at TIMESTAMPTZ - updated_at TIMESTAMPTZ -) -``` - -### RLS Policy Strategy - -DM threads and messages use TEXT columns for participant IDs (not UUID foreign keys) to support both Supabase user UUIDs and `agdns:` agent IDs. RLS policies check both formats: - -```sql --- User can see threads where they participate via UUID or agent_id -auth.uid()::text = initiator_id -OR EXISTS (SELECT 1 FROM persona_agents WHERE user_id = auth.uid() AND agent_id = initiator_id) -``` - ---- - -## Heartbeat Architecture (Scalable Design) - -### The Problem - -The ZyndAI SDK spawns one WebSocket thread per agent for heartbeats. At scale: -- 10,000 users = 10,000 Python threads = ~80GB thread stack memory -- GIL contention destroys throughput -- OS file descriptor limits on sockets - -### The Solution: Batched Heartbeat Manager - -A single asyncio task manages all heartbeats: - -``` -┌─────────────────────────────────────────────┐ -│ HeartbeatManager (singleton) │ -│ │ -│ agents: { agent_id → (agent_id, priv_key) } │ -│ │ -│ Loop (every 30s): │ -│ 1. Snapshot all agents │ -│ 2. Split into batches of 50 │ -│ 3. For each batch: │ -│ a. Open ONE WebSocket to registry │ -│ b. Sign + send heartbeat for each agent │ -│ c. Close connection │ -│ 4. Stagger batches across the 30s window │ -│ 5. Sleep remaining time │ -└─────────────────────────────────────────────┘ -``` - -**Performance characteristics:** - -| Metric | Per-Thread (old) | Batched Manager (new) | -|--------|-----------------|----------------------| -| 10K users: threads | 10,000 | 1 asyncio task | -| 10K users: WebSocket connections | 10,000 persistent | ~200 transient per cycle | -| 10K users: memory | ~80GB stack | ~50MB dict | -| Signatures/sec | N/A (per thread) | ~333/sec (trivial for Ed25519) | -| Startup time (rehydration) | Sequential | Parallel batch | - -The registry's 5-minute inactivity timeout provides ample buffer — even if a full cycle takes a few extra seconds, no agent goes stale. - -**Theoretical capacity:** 100,000+ agents per server instance. - ---- - -## Request Flow - -### User Chat (internal) - -``` -Browser → POST /api/chat/message (JWT auth) - → Orchestrator - → Build system prompt (persona config from DB) - → LLM (OpenAI/Gemini/Custom) with MCP tools - → Execute tool calls (Twitter, Calendar, Notion, etc.) - → Return response + actions_taken -``` - -### Persona Registration - -``` -Browser → POST /api/persona/register - → persona_manager.create_persona() - → Allocate derivation_index - → Derive Ed25519 keypair from developer key - → POST zns01.zynd.ai/v1/agents (signed registration) - → INSERT persona_agents - → heartbeat_manager.add_agent() - → Return { agent_id, webhook_url } -``` - -### Incoming Network Message (async) - -``` -External Agent → POST /api/persona/webhooks/{user_id} - → Parse AgentMessage - → Log to dm_messages (inbound) - → Background task: - → Orchestrator (is_external=True, security boundary) - → Log to dm_messages (outbound) - → HTTP POST reply to sender's webhook -``` - -### Incoming Network Message (sync) - -``` -External Agent → POST /api/persona/webhooks/{user_id}/sync - → Parse AgentMessage - → Log to dm_messages (inbound) - → Orchestrator (blocking) - → Log to dm_messages (outbound) - → Return { response } (within 30s timeout) -``` - -### Agent Discovery - -``` -Browser → GET /api/persona/search?query=Alice - → POST zns01.zynd.ai/v1/search { query, tags: ["persona"] } - → Filter results to persona-tagged agents only - → Return { results: [...] } -``` - -### Server Startup - -``` -uvicorn main:app - → lifespan startup: - 1. start_zynd_agent() — global developer agent on port 5050 - 2. persona_manager.startup() — rehydrate all active personas: - a. Load all active=true from persona_agents - b. For each: derive keypair from dev_key + index - c. Register with heartbeat_manager - d. Start heartbeat loop -``` - -### Server Shutdown - -``` -SIGTERM/SIGINT - → lifespan shutdown: - 1. persona_manager.shutdown() - → heartbeat_manager.stop() — cancel asyncio task - → All WebSocket connections close gracefully -``` - ---- - -## Webhook Architecture - -All user personas share a single FastAPI server. Webhooks are differentiated by user_id in the URL path: - -``` -https://your-server.com/api/persona/webhooks/{user_id} (async) -https://your-server.com/api/persona/webhooks/{user_id}/sync (sync) -``` - -The webhook URL is registered with the Zynd registry during persona creation, so other agents know where to send messages. - -**Why single-port, path-based routing?** -- No per-user ports or processes needed -- Standard reverse proxy (nginx/Caddy) works out of the box -- Scales horizontally — add more FastAPI workers, not more ports -- user_id in the path makes routing deterministic and debuggable - ---- - -## Registry Integration (zns01.zynd.ai) - -### Endpoints Used - -| Operation | Endpoint | When | -|-----------|----------|------| -| Register agent | `POST /v1/agents` | Persona creation | -| Update agent | `PUT /v1/agents/{id}` | Persona update | -| Delete agent | `DELETE /v1/agents/{id}` | Persona deletion | -| Search agents | `POST /v1/search` | Discovery (with `tags: ["persona"]`) | -| Get agent card | `GET /v1/agents/{id}/card` | Webhook URL lookup | -| Heartbeat | `WSS /v1/heartbeat` | Continuous liveness (30s interval) | - -### Authentication - -All registry operations use Ed25519 signature-based auth: -``` -message = "{agent_id}:{timestamp}" -signature = Ed25519.sign(message, private_key) -header: "ed25519:{base64(signature)}" -``` - -No API keys needed — identity is proven cryptographically. - ---- - -## Security Model - -### Internal vs External Requests - -The orchestrator builds different system prompts based on request origin: - -**Internal (user chat):** Full access to all connected tools. Conversational, helpful. - -**External (network webhook):** Strict security boundary: -- Only capabilities the user explicitly granted are available -- No destructive actions -- No private data leakage -- Brief, professional responses -- External agent identity logged - -### Authentication Layers - -| Layer | Mechanism | -|-------|-----------| -| Frontend → Backend | Supabase JWT (Bearer token) | -| Backend → Registry | Ed25519 signatures | -| Backend → Supabase | Service role key (bypasses RLS) | -| Agent → Agent | AgentMessage with sender_id verification | -| Registry → Agent | WebSocket heartbeat with signed messages | - ---- - -## File Structure - -``` -backend/ -├── main.py # FastAPI app with lifespan -├── config.py # Environment config (v2) -├── requirements.txt # Dependencies (zyndai-agent>=0.3.2) -├── agent/ -│ ├── zynd_core.py # Global developer agent (port 5050) -│ ├── persona_manager.py # HD derivation, registration, lifecycle -│ ├── heartbeat_manager.py # Batched async heartbeat for all personas -│ └── orchestrator.py # LLM orchestration loop -├── api/ -│ ├── auth.py # Supabase JWT validation -│ ├── persona.py # Persona CRUD + webhooks + search proxy -│ ├── chat.py # User chat endpoint -│ ├── connections.py # OAuth connection management -│ ├── oauth_routes.py # OAuth flows -│ └── telegram.py # Telegram bot webhook -├── mcp/ -│ ├── server.py # ContextAware MCP tool registry -│ └── tools/ -│ ├── zynd_network.py # Registry search + agent messaging (v2) -│ ├── twitter.py # X/Twitter tools -│ ├── linkedin.py # LinkedIn tools -│ ├── notion.py # Notion tools -│ └── google/ # Calendar, Docs, Gmail, Sheets, Drive -├── services/ -│ └── token_store.py # OAuth token CRUD -└── db/ - ├── schema.sql # Complete v2 schema (fresh install) - └── migrate_v2.sql # Migration from v1 → v2 - -webapp/src/ -├── components/ -│ ├── PersonaBuilder.tsx # Identity creation (agent_id, not DID) -│ ├── MessagesPanel.tsx # Network DMs (zns01.zynd.ai search) -│ ├── ChatInterface.tsx # User-to-agent chat -│ └── ConnectionsPanel.tsx # OAuth integrations -├── contexts/ -│ └── DashboardContext.tsx # Auth state -└── app/dashboard/ - ├── layout.tsx # Sidebar navigation - ├── page.tsx # Redirect logic - ├── chat/page.tsx - ├── identity/page.tsx - ├── messages/page.tsx - └── connections/page.tsx -``` - ---- - -## Migration Checklist (v1 → v2) - -- [x] Replace `did:polygon:...` identifiers with `agdns:...` agent IDs -- [x] Replace `registry.zynd.ai` with `zns01.zynd.ai` -- [x] Replace API key auth with Ed25519 signature auth -- [x] Replace `ConfigManager.create_agent()` with HD key derivation -- [x] Replace filesystem config (`.agent-{user_id}/`) with `persona_agents` DB table -- [x] Replace per-agent heartbeat threads with batched async manager -- [x] Drop `persona_dids` table, create `persona_agents` table -- [x] Truncate `dm_threads` and `dm_messages` (fresh start) -- [x] Update RLS policies to use `persona_agents` instead of `persona_dids` -- [x] Update frontend to use `agent_id` instead of `did`/`didIdentifier` -- [x] Update frontend search to use `zns01.zynd.ai/v1/search` (via backend proxy) -- [x] Add persona deletion endpoint (`DELETE /api/persona/{user_id}`) -- [x] Add search proxy endpoint (`GET /api/persona/search`) -- [x] Add profile update endpoint (`PUT /api/persona/{user_id}/profile`) -- [x] Remove old SQL migration files -- [x] Remove `.well-known/` directory from backend -- [x] Update `.env.example` with new config variables -- [x] Use `contextlib.asynccontextmanager` lifespan instead of deprecated `@app.on_event` -- [x] Add `profile` JSONB column to `persona_agents` (social links, title, org, interests) -- [x] Add networking MCP tools (get_persona_profile, list_my_connections, request_connection, check_connection_status) -- [x] Rewrite system prompt to be networking-first (search Zynd Network before internet) -- [x] Redesign identity page with full profile display, social links, and edit mode -- [x] Gate dashboard pages behind persona deployment (mandatory identity) diff --git a/docs/persona/linkedin-portability-pre-apply.md b/docs/persona/linkedin-portability-pre-apply.md deleted file mode 100644 index bd11f34..0000000 --- a/docs/persona/linkedin-portability-pre-apply.md +++ /dev/null @@ -1,270 +0,0 @@ -# Zynd — Google + LinkedIn Auth Compliance & Data Controls - -**Entity:** Zynd AI Inc (Delaware) -**Scope:** This repo only (`persona.zynd.ai` — `backend/` + `webapp/`). The ZYND memory layer (`api.zynd.ai`) is a separate deployment and is **out of scope**. -**Last updated:** 2026-08-22 - ---- - -## 1. Objective - -Reach a credible, unified compliance baseline covering **both Google OAuth and LinkedIn OAuth**, so that a new, isolated LinkedIn Developer App can request **Member Data Portability** without being rejected for missing consent / storage / deletion / disclosure infrastructure. - -Key LinkedIn facts driving this work: - -- Member Data Portability is **not a paid product** (LinkedIn "Costs & Fees" §8.2 does not apply; each party bears its own costs). -- It must live on its **own app** — not the current `Zynd Persona` app (which has Share on LinkedIn, Sign-In with LinkedIn, Advertising API request). -- Access is **member-authorized** and region-scoped (DMA program → EU/EEA/CH members). -- LinkedIn API terms require LinkedIn-sourced data to be **identifiable, segregable, and selectively deletable**, deleted on member request / account closure, and **never mixed with scraped/crawled data** (e.g. Apify). -- Storage is allowed with member consent + legal basis. - -## 2. Current state (audited 2026-08-22) - -### Already present - -| Capability | Location | -|---|---| -| Privacy Policy page (Google-only) | `webapp/src/app/privacy/page.tsx` | -| Terms of Service page (Google-only) | `webapp/src/app/terms/page.tsx` | -| Full account purge (delete everything) | UI `webapp/src/components/settings/DeleteAccountModal.tsx` + `webapp/src/app/dashboard/settings/you/page.tsx` → `DELETE /api/persona/{id}/account` → `backend/agent/persona_manager.py:422` `purge_user_account` | -| Google disconnect | `webapp/src/app/dashboard/settings/accounts/page.tsx` → `DELETE /api/connections/google` (`backend/api/connections.py:54`) | -| LinkedIn disconnect (wipes scrape + token) | `DELETE /api/linkedin/me` (`backend/api/linkedin.py:130`) + `DELETE /api/connections/linkedin` | -| OAuth flows | `backend/api/oauth_routes.py` — Google `:285`, LinkedIn `:98` | -| Segregated token storage | `api_tokens` table (`backend/services/token_store.py`) | -| Segregated LinkedIn data | `linkedin_profiles` table (`db/migrations/0000_baseline/migration.sql:224`) | -| RLS on all sensitive tables | `db/sql/policies.sql` | -| Privacy audit log (group discovery) | `persona_group_discovery_audit` table | - -### Gaps - -1. Legal pages are **Google-only** — no LinkedIn disclosure anywhere. -2. Terms governing law = **California** (`webapp/src/app/terms/page.tsx:247`) — entity is Delaware. -3. No `/data-deletion`, `/security`, `/contact` pages; no cookie/analytics note (GA `G-1L9YDBKGRT` is loaded). -4. No granular **"Delete [provider] data"** (keep-account delete) — only full disconnect. -5. `purge_user_account` does **not** explicitly delete `linkedin_profiles` / `telegram_links` / `telegram_chat_history` (relies on `ON DELETE CASCADE` from `auth.users`). No deletion audit event. -6. No **data export** ("download my data") and no **retention policy**. -7. OAuth consent does not explicitly disclose storage/processing (relies on provider consent screens only). -8. LinkedIn data is currently **Apify-scraped** (`backend/services/linkedin_scraper.py`) — no `source` lineage; cannot be mixed with future Portability data. - ---- - -## 3. How to test (baseline) - -- **Backend tests:** `cd backend && python -m pytest tests/ -q` (no pytest config file; `conftest.py` injects `backend/` into `sys.path`). -- **Webapp build (REQUIRED before deploy):** `cd webapp && npm run build`. -- **Webapp lint:** `cd webapp && npm run lint`. -- **Deploy flow (dev → prod):** see `AGENTS.md`. `npm run build` is mandatory after every webapp change; backend needs `pip install -r requirements.txt` only if deps changed. - -Run backend tests **before** any change, then again after each change group. - -> **Known baseline (2026-08-22):** `backend/tests/test_linkedin_people_search.py::test_maps_real_actor_output_shape` fails (`count == 1` expected, got `0`). This is **pre-existing and unrelated** to this plan — the scraper mapping changed but the test wasn't updated. Treat "153 passed, 1 failed (this one)" as the green baseline. Do not fix it as part of this plan unless separately approved. - ---- - -## 4. Change catalog - -Legend: -- **Type** — `additive` (new file/route/field, no existing behavior changed) vs `modifies-existing` (changes current output/behavior → **requires approval**). -- **🔒 Approval** — item changes an existing workflow; do not implement until explicitly approved. - ---- - -### Group A — Public legal pages (Phase 1) - -#### A1. Add "LinkedIn Data" section to Privacy Policy — `modifies-existing` 🔒 -- **File:** `webapp/src/app/privacy/page.tsx` -- **Change:** Insert a new `
` (after the existing Google scopes section) titled **"LinkedIn Data"** covering: what Zynd accesses (public profile fields, posts — as authorized), the OAuth scopes (`openid profile email w_member_social`), purpose (persona context / profile enrichment), storage (encrypted, segregated `linkedin_profiles` table), consent basis, deletion mechanism, and retention. -- **Also:** add a **"Data Retention"** section covering both providers; update entity string `ZyndAI` → **"Zynd AI Inc"** in the intro; update "Last updated". -- **Approval:** yes — changes existing published page text. -- **Test:** `npm run build`; visually verify `/privacy`; confirm existing Google sections still render unchanged. - -#### A2. Add "LinkedIn API Services" section to Terms + fix governing law — `modifies-existing` 🔒 -- **File:** `webapp/src/app/terms/page.tsx` -- **Change:** Add a LinkedIn section mirroring the existing Google API section (scopes + limited-use language). Change §11 governing law from **"State of California"** → **"State of Delaware"**. Entity → **"Zynd AI Inc"**. Update "Last updated". -- **Approval:** yes — changes existing Terms (governing-law change is a legal decision; confirm before shipping). -- **Test:** `npm run build`; verify `/terms`. - -#### A3. Create `/data-deletion` page — `additive` -- **File (new):** `webapp/src/app/data-deletion/page.tsx` -- **Change:** Public page describing how to delete data for **both** providers: disconnect steps, "delete provider data" (in-app), full account deletion, retention window, and contact email. -- **Approval:** no (new page). -- **Test:** `npm run build`; visit `/data-deletion`. - -#### A4. Create `/security` page — `additive` -- **File (new):** `webapp/src/app/security/page.tsx` -- **Change:** Describe encryption at rest (Supabase), RLS, OAuth token handling, and scope minimization for both providers. -- **Approval:** no. -- **Test:** `npm run build`; visit `/security`. - -#### A5. Create `/contact` page — `additive` -- **File (new):** `webapp/src/app/contact/page.tsx` -- **Change:** Contact page listing support + privacy/DPO contact email (see **Decision D1**). -- **Approval:** no. -- **Test:** `npm run build`; visit `/contact`. - -#### A6. Cookie / analytics note — `additive` -- **File:** `webapp/src/app/privacy/page.tsx` (new short section) or a small `/cookies` page. -- **Change:** Disclose GA `G-1L9YDBKGRT` usage. Minimal note in Privacy Policy is sufficient for now. -- **Approval:** no (additive section). -- **Test:** `npm run build`. - -#### A7. Wire new pages into nav/footers + crawler metadata — `modifies-existing` 🔒 -- **Files:** - - `webapp/src/app/LandingClientWrapper.tsx` (footer ~line 230) - - `webapp/src/app/privacy/page.tsx` (footer ~line 216) - - `webapp/src/app/terms/page.tsx` (footer ~line 259) - - `webapp/src/app/llms.txt/route.ts` (`STATIC_PAGES` array) - - `docs/seo-plan.md` (sitemap static list) and, if a real `webapp/public/sitemap.xml` exists, add the new paths. -- **Change:** Add footer links for `/data-deletion`, `/security`, `/contact` (and `/cookies` if created). Add them to `llms.txt` static list. -- **Approval:** yes — edits existing footers/nav (low visual risk, but touches published surfaces). -- **Test:** `npm run build`; click each new footer link on `/`, `/terms`, `/privacy`. - ---- - -### Group B — Backend data controls (Phase 2) - -#### B1. Granular "delete LinkedIn data" endpoint — `additive` -- **File:** `backend/api/linkedin.py` -- **Change:** Add `DELETE /api/linkedin/data` that deletes the `linkedin_profiles` row **and** the `api_tokens` `linkedin` row, while **keeping the account/persona** (distinct from full disconnect — functionally the same DB writes as today's disconnect, but exposed as an explicit "delete my LinkedIn data" action for compliance). -- **Approval:** no (additive route; does not alter existing `DELETE /me`). -- **Test:** add `backend/tests/test_linkedin_delete_data.py` — mock Supabase, assert both tables hit with `eq("user_id", ...)`. Run `python -m pytest tests/test_linkedin_delete_data.py`. - -#### B2. Generalize provider data deletion — `additive` -- **File:** `backend/api/connections.py` -- **Change:** Add `DELETE /api/connections/{provider}/data` that deletes provider tokens (and for `linkedin`, the profile data) without removing the account. Keep the existing `DELETE /api/connections/{provider}` (disconnect) unchanged. -- **Approval:** no (additive route). -- **Test:** extend `backend/tests/` with a connections data-delete test (mirror existing disconnect tests if any). - -#### B3. Harden `purge_user_account` to explicitly delete provider data + audit — `modifies-existing` 🔒 -- **File:** `backend/agent/persona_manager.py` (`purge_user_account`, line 422) -- **Change:** Add explicit `linkedin_profiles`, `telegram_links`, `telegram_chat_history` deletes **before** the `auth.users` delete (defence-in-depth; today these rely on FK cascade and would linger if the final step fails). Append a `data_deleted` audit entry (reuse `persona_group_audit_events` pattern or a lightweight audit table). -- **Approval:** yes — modifies the account-deletion path (a production-critical workflow). -- **Test:** add `backend/tests/test_purge_account.py` — mock Supabase, assert each table is deleted; assert behavior is unchanged when `auth.users` delete succeeds. -- **Note:** do **not** change the existing cascade semantics; this is strictly additive hardening. - -#### B4. Deletion audit logging — `additive` -- **File:** `backend/services/` (small helper) or reuse `persona_group_audit_events`. -- **Change:** Record provider-data deletions (user_id, provider, timestamp, actor=user). -- **Approval:** no. -- **Test:** unit-test the helper; verify no audit row is written for failed deletes. - -#### B5. Data export endpoint — `additive` -- **File:** `backend/api/persona.py` (new `GET /api/persona/{user_id}/export`) -- **Change:** Return a JSON bundle of the user's own data (persona profile, brief, chat history, connected-provider list). Read-only; does not expose tokens. -- **Approval:** no (additive, read-only). -- **Test:** `backend/tests/test_persona_export.py` — assert response shape + no token fields. - ---- - -### Group C — Frontend data controls (Phase 2) - -#### C1. "Delete data" buttons on Accounts page — `modifies-existing` 🔒 -- **File:** `webapp/src/app/dashboard/settings/accounts/page.tsx` -- **Change:** For LinkedIn **and** Google cards, add a secondary **"Delete data"** action (calls B1/B2) distinct from the existing **"Disconnect"** (which already exists). Surface the deletion result via the existing `oauthFlash`/notice pattern. Add an inline confirm note describing what is removed. -- **Approval:** yes — adds controls to an existing, working settings screen (should not alter existing Disconnect/Connect behavior). -- **Test:** `npm run build`; manual: connect → delete data → confirm the card returns to "Not connected" state and that Disconnect/Connect still work. - -#### C2. "Download my data" button — `additive` -- **File:** `webapp/src/app/dashboard/settings/you/page.tsx` (Account card) + `webapp/src/components/settings/` if a modal is needed. -- **Change:** Add a "Download my data" button that fetches B5 and triggers a client-side download. -- **Approval:** no. -- **Test:** `npm run build`; manual download + JSON validity. - -#### C3. OAuth consent disclosure text — `modifies-existing` 🔒 -- **File:** `webapp/src/app/dashboard/settings/accounts/page.tsx` -- **Change:** Add a one-line disclosure near the LinkedIn/Google connect buttons, e.g. *"By connecting, you authorize Zynd to store and process the data described in our Privacy Policy."* -- **Approval:** yes — changes existing connect-flow copy. -- **Test:** `npm run build`; visual check on both cards. - ---- - -### Group D — Data lineage + migration (Phase 3) - -#### D1. Add `source` field to `linkedin_profiles` — `modifies-existing` 🔒 (DB migration) -- **File:** new migration `db/migrations/XXXX_add_linkedin_source.sql` (or `db/sql/` patch, following existing convention — see `backend/db/patch_*.sql`) -- **Change:** `ALTER TABLE public.linkedin_profiles ADD COLUMN source TEXT NOT NULL DEFAULT 'apify_scrape';` with a CHECK (`source IN ('apify_scrape','linkedin_api','portability')`). Backfill existing rows to `'apify_scrape'`. -- **Approval:** yes — schema change on a live table (requires running the migration on the shared Supabase project; both prod + dev copies). -- **Test:** run the migration via `psql "$DIRECT_URL" -f ` (see `webapp/package.json` `db:policies` for the connection pattern); verify default + CHECK enforcement; ensure existing scrape upserts still succeed. - -#### D2. Write `source` on scrape vs OIDC upsert — `modifies-existing` 🔒 -- **File:** `backend/services/linkedin_scraper.py` (scrape upsert) and `backend/api/oauth_routes.py` (OIDC placeholder upsert, ~line 199) -- **Change:** Set `source='apify_scrape'` on scrape writes and `source='linkedin_api'` on OIDC placeholder writes. Any future Portability import MUST write `source='portability'` to a **separate table** (never `linkedin_profiles`). -- **Approval:** yes — modifies existing write paths (must not alter current scrape behavior; only add the column value). -- **Test:** update `backend/tests/test_linkedin_scraper_search_payload.py` / `test_linkedin_post_fields.py` if they assert on the upsert payload; add assertion that `source` is set. - -#### D3. Document separation rule — `additive` -- **File:** `docs/` note or this file's Phase 3 section. -- **Change:** State that Portability-derived connection data must never be co-mingled with `apify_scrape` data. -- **Approval:** no. - ---- - -### Group E — LinkedIn submission (Phase 4) - -#### E1. Create isolated LinkedIn Developer App — `external` (no code) -- New app with **only** Member Data Portability product. Register with **Zynd AI Inc** (Delaware) details, privacy policy URL, contact email. - -#### E2. Verify Portability schema includes 1st-degree connections — `external` (no code) -- Confirm the 2026 Portability API actually exposes connections before building any import feature. - ---- - -## 5. Regression risk matrix - -| Change | Existing flow touched | Blast radius | Mitigation | -|---|---|---|---| -| A1/A2 legal copy | `/privacy`, `/terms` | Low (static) | build + visual | -| A7 footer/nav links | `/`, `/terms`, `/privacy`, `llms.txt` | Low | build + click-through | -| B3 purge hardening | `DELETE /api/persona/{id}/account` | **High** (account deletion) | additive-only, explicit test | -| B1/B2 new delete endpoints | none (new routes) | Low | new tests | -| C1 delete buttons | Accounts page | Medium (existing UI) | keep Disconnect/Connect untouched | -| C3 consent copy | connect flow | Low (text only) | build + visual | -| D1/D2 source field | scrape + OIDC upsert, schema | **High** (live table + write paths) | backfill default, keep upserts compatible | - ---- - -## 6. Implementation order (safe sequencing) - -1. **Group A** (legal pages) — static, low risk. -2. **Group D1/D2** (source field) — schema first so future writes are tagged. *(needs approval)* -3. **Group B1/B2 + B4/B5** (new endpoints) — additive backend. -4. **Group C1/C2/C3** (frontend controls) — after backend endpoints exist. -5. **Group B3** (purge hardening) — last among backend changes, with full test. *(needs approval)* -6. **Group E** (external LinkedIn steps) — after all code is live + `npm run build` on dev/prod. - -## 7. Deploy checklist - -- [ ] `cd backend && python -m pytest tests/ -q` green. -- [ ] `cd webapp && npm run build` green, `npm run lint` green. -- [ ] Run D1 migration on Supabase (prod project shared by both copies). -- [ ] Deploy **dev** copy first (`git pull`, restart `api-dev web-dev`, build webapp), smoke-test. -- [ ] Deploy **prod** copy (`git pull`, restart `api web`, build webapp), smoke-test. -- [ ] Verify new pages reachable at `persona.zynd.ai/data-deletion`, `/security`, `/contact`. - ---- - -## 8. Resolved decisions - -- **D1 (contact email):** `contact@zynd.ai` (privacy/DPO + all "Contact Us" sections). -- **D2 (entity string):** `Zynd AI Inc` — registered address: **Zynd AI Inc, 8 The Green STE A, Dover, DE 19901**. -- **D3 (governing law):** **Delaware** (Terms §11). -- **D4 (cookie policy):** a note inside the Privacy Policy (A6), not a separate page. - ---- - -## 9. Permission summary - -The following items change **existing behavior/workflows** and will **not** be implemented until you approve each: - -| ID | What it changes | Why approval is needed | -|---|---|---| -| A1 | Privacy Policy content | published page text | -| A2 | Terms content + governing law | legal text + jurisdiction | -| A7 | footers / llms.txt | published surfaces | -| B3 | account-deletion path | production-critical workflow | -| C1 | Accounts settings UI | existing working screen | -| C3 | connect-flow copy | existing flow text | -| D1 | `linkedin_profiles` schema | live table migration | -| D2 | scrape/OIDC write paths | existing write behavior | - -All other items are **additive** (new pages, new routes, new endpoints, new field default) and are safe to implement without touching existing behavior. diff --git a/docs/persona/seo-plan.md b/docs/persona/seo-plan.md deleted file mode 100644 index 9886ce4..0000000 --- a/docs/persona/seo-plan.md +++ /dev/null @@ -1,102 +0,0 @@ -# SEO & Crawlability Improvement Plan — ZyndAI Persona - -**GA Measurement ID:** `G-1L9YDBKGRT` - ---- - -## Task 1: robots.txt - -**File:** `webapp/public/robots.txt` - -``` -User-agent: * -Allow: / -Disallow: /onboarding/ -Disallow: /dashboard/ -Sitemap: https://persona.zynd.ai/sitemap.xml -``` - ---- - -## Task 2: sitemap.xml - -**File:** `webapp/public/sitemap.xml` - -Static XML listing `/`, `/terms`, `/privacy`, `/data-deletion`, `/security`, `/contact` with `` and ``. - ---- - -## Task 3: llms.txt Route Handler - -**File:** `webapp/src/app/llms.txt/route.ts` - -Dynamically generates Markdown-formatted list of: -- Static pages: `/`, `/terms`, `/privacy` -- Public persona pages (from backend API) -- Published pages (from backend API) - -Output: `text/plain; charset=utf-8` - ---- - -## Task 4: Google Analytics - -- Install `@next/third-parties` package -- Add `` in root layout -- Add `NEXT_PUBLIC_GA_ID=G-1L9YDBKGRT` to `.env.local` - ---- - -## Task 5: Root Metadata Enhancement - -**File:** `webapp/src/app/layout.tsx` - -Add to existing metadata: -- `metadataBase: new URL("https://persona.zynd.ai")` -- `robots: { index: true, follow: true }` -- `openGraph` with site_name, image, locale -- `twitter: { card: "summary_large_image" }` -- `alternates: { canonical: "/" }` -- `keywords` - ---- - -## Task 6: JSON-LD Structured Data - -**File:** `webapp/src/app/layout.tsx` - -Add `WebSite` schema via `