Skip to content

Re-deriving the skills catalog every turn invalidates the Anthropic prompt cache #5569

Description

@practical-tools-lab

Area

Multiple areas

What are you trying to accomplish?

I want a long-running Codex session to keep its Anthropic prompt cache intact across turns, so that ordinary work on disk — editing a SKILL.md, installing or removing a skill — does not silently re-bill the entire conversation prefix on the next turn.

PROBLEM. The <skills_instructions> block (~24,500 characters, ~7,000 tokens) is injected ahead of the conversation on every request. It is re-derived per turn rather than snapshotted per session, so any edit to any SKILL.md on disk - or any change to the registered skill set - produces a different block mid-session. Because Anthropic prompt caching keys on the prefix, a single changed byte near the front invalidates the entire cached context behind it.

What prevents this today?

The catalog is rebuilt from disk for each request instead of being frozen for the lifetime of the session, so there is no stable byte sequence to cache against. Nothing in the current behaviour lets a session say "use the skills catalog as it was when I started".

MEASUREMENT (ours, 2026-09-23, 71,191 requests across 872 Codex rollout files over 7 days):

  • Splitting requests by position within their turn: the FIRST request of a turn has a 17.22% full-cache-re-creation rate (572 of 3,322) versus 0.55% for later requests in the same turn (372 of 67,809). That is 31x.
  • Those first-of-turn requests are 4.7% of traffic but carry 29.8% of all cache-write tokens (103,845,984 of 347,741,169).
  • Hashing every developer/system message body by its opening text shows <skills_instructions> appearing with 13 distinct bodies inside one session, 12 in another, 8 and 6 in others. A diff of two bodies from the same session shows an ordinary one-line skill-description edit.
  • Concrete cost: editing one SKILL.md while a 500,000-token session is live costs that session 500,000 x 1.25 = 625,000 token-equivalents on its next turn, multiplied by every session running at that moment.

What should OpenCodex do?

REQUEST. Snapshot the skills catalog once at session start and reuse those exact bytes for the lifetime of the session, instead of re-deriving it each turn. New or edited skills would then take effect for new sessions, which is the normal expectation anyway. If a live refresh must stay possible, gate it behind an explicit opt-in or an explicit user action rather than making it the default.

NOTE ALSO. The same reasoning applies to other front-of-prompt blocks we measured as volatile within a single session: <permissions instructions>, <app-context>, <model_switch>, and any injected block carrying a live timestamp or index count.

Example usage or interface

Before (today), within one live session:

turn 1  -> <skills_instructions> body sha256 a1b2c3...   cache write: full prefix
        (user edits one line of a SKILL.md description)
turn 2  -> <skills_instructions> body sha256 d4e5f6...   cache write: full prefix again

After (requested):

turn 1  -> <skills_instructions> body sha256 a1b2c3...   cache write: full prefix
        (user edits one line of a SKILL.md description)
turn 2  -> <skills_instructions> body sha256 a1b2c3...   cache read: hit
        (the edit applies to the next session started)

Optional escape hatch, if a live refresh must remain available:

# opt in explicitly; default stays snapshotted
ocx config set skills.catalog_refresh per_session   # default
ocx config set skills.catalog_refresh per_turn      # previous behaviour, opt-in

Alternatives or workarounds

The only workaround available to us is operational: never touch any SKILL.md while any session is live, and never install or remove a skill during the working day. That is not enforceable across multiple concurrent sessions and multiple machines, and it fails silently — the cost shows up as cache-write tokens on the next turn with no signal that anything happened.

Trimming the catalog is not a fix either. Shrinking the block reduces the size of the volatile region but not its position; the invalidation cost is driven by everything cached behind it, not by the block's own length.

Additional context

Related prefix-stability work in this repository: #4148 (mid-conversation system messages aggregated into top-level instructions, destroying the prefix), #4052 / #4347 (stabilizing Responses instructions for prompt cache), #4532 (dynamic image downscaling on append busting the prefix cache), and #4546 / #4580 (pool routing destroying prompt cache). This report is the same class of defect at a different layer: a front-of-prompt block that is derived from mutable on-disk state rather than pinned at session start.

Upstream basis: Anthropic prompt caching matches on an exact prefix, so the first differing byte invalidates everything after it.

Checks

  • I searched existing issues and documentation.
  • This request describes a concrete OpenCodex workflow rather than merely naming a desired technology.
  • I removed secrets and personal data.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    catalogModel catalog, slugs, visibility, routed entriesenhancementNew feature or requestpriority: P2Medium: provider/client-specific bug with a workaround, bounded enhancement tied to a tracked issue,

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions