Skip to content

Attribute handler outcomes and pool waits - #465

Merged
valthon merged 1 commit into
codex/mixed-application-capacityfrom
codex/workbench-operation-attribution
Sep 20, 2026
Merged

valthon merged 1 commit into
codex/mixed-application-capacityfrom
codex/workbench-operation-attribution

Conversation

@valthon

@valthon valthon commented Sep 20, 2026

Copy link
Copy Markdown
Owner

Summary

A slow handler can spend time outside SQL, and a fast HTTP response can leave substantial background work behind. The opt-in workbench now records returned HTTP status classes, handler errors, pool mutex acquisition, and declared scheduled/durable job attempts. All scopes and query shapes remain bounded; no tenant identifiers, job payloads, or dynamic submitted names are captured, and instrumentation compiles out by default.

The capacity runner gains --workbench, retaining initial and per-phase cumulative operator snapshots outside timed load. This connects the mixed-workload report to route/job evidence without treating inclusive durations as exclusive CPU time.

Documentation & examples sync

  • Framework/API/testing docs, skill references, site overview/configuration, changelog, and backlog updated.
  • Runnable workbench fixture, SQLite/PostgreSQL live coverage, compile-out checks, and capacity integration CI updated.
  • Public-site generation, validation, build, doctor, static/generated checks, and browser smoke passed.
  • N/A: SDK behavior, screenshots, and example application behavior.

Verification

  • Enabled unit suite: 2,244 passed, 48 skipped; default-off suite: 2,239 passed, 52 skipped.
  • SQLite/PostgreSQL/enabled/disabled live workbench: 9 passed; 8 compile-out pattern checks passed.
  • Combined capacity suite: 14 passed, 16 subtests, including cumulative job and HTTP observations.
  • Independent review of dispatch cleanup, scope identity, bounds, privacy, and report integration found no remaining findings.
  • Clean-checkout chapter gates verify both commits independently.

Pool timing measures mutex acquisition, excluding connection creation, SQL/database locks, and ownership time. Job timing covers handler attempts, excluding queue residence, retry backoff, claim/acknowledgment work, memory jobs, and app.submit. HTTP and job scopes share a bounded table with separate dropped counters.

@valthon
valthon added this pull request to stack #466 September 20, 2026 00:39
Copilot AI lite review requested due to automatic review settings September 20, 2026 01:23
@valthon
valthon force-pushed the codex/workbench-operation-attribution branch from 90edd5d to d86294e Compare September 20, 2026 01:23
@valthon

valthon commented Sep 20, 2026

Copy link
Copy Markdown
Owner Author

Fixed the admission CI failure. The generic admin login helper is imported by durable-capacity tests; adding the workbench boot-job wait there made those default-off fixtures request an unavailable diagnostics endpoint and fail with KeyError: jobs. The wait now lives in a workbench-specific helper, while shared login only authenticates. All five admission tests and all six workbench live tests pass locally against the relevant rebuilt fixtures. Amended the feature commit and preserved native stack #466; hosted CI is rerunning.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Durable-job attribution remains active through claim completion and can misattribute acknowledgment and subsequent job work.

Review effort: Lite
Findings: None

What changed in this PR

Adds opt-in workbench attribution for HTTP outcomes, pool waits, scheduled/durable jobs, and capacity-runner snapshots.

Changes:

  • Adds bounded metrics and operator-report fields.
  • Instruments routes, jobs, schedulers, and database pools.
  • Expands tests, fixtures, CI, documentation, and capacity reporting.
File Summary
tools/​application_capacity.py Adds workbench snapshots and CLI support.
tests/​tools/​test_application_capacity.py Tests capacity integration.
tests/​admin/​test_query_workbench.py Tests HTTP, pool, and job metrics.
tests/​admin/​test_query_workbench_postgres.py Tests PostgreSQL pool metrics.
src/​server.zig Attributes custom-route outcomes and errors.
src/​scheduler.zig Instruments scheduled jobs.
src/​router.zig Records route outcomes.
src/​queue/​memory.zig Masks inline memory-job attribution.
src/​queue/​durable.zig Instruments durable job attempts.
src/​query_workbench.zig Adds bounded attribution and metric storage.
src/​backend/​sqlite/​db.zig Measures SQLite pool acquisition.
src/​backend/​postgres/​pool.zig Measures PostgreSQL pool acquisition.
src/​api/​query_workbench.zig Exposes expanded workbench reports.
skills/​zigbase-app-genesis/​references/​testing.md Updates testing guidance.
site/​sources/​docs/​overview.md Updates the site overview.
site/​sources/​docs/​configuration.md Documents new diagnostics.
scripts/​check-query-workbench-gating.sh Extends compile-out checks.
fixtures/​query-workbench/​main.zig Adds pool-wait and job fixtures.
docs/​testing.md Documents workbench runner usage.
docs/​framework.md Documents workbench scope and metrics.
docs/​api.md Documents diagnostic API additions.
changelog.d/​workbench-operation-attribution.md Adds the changelog entry.
BACKLOG.md Updates the workbench backlog.
.github/​workflows/​ci.yml Adds capacity workbench CI coverage.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Distinguish returned HTTP status classes, handler errors, and declared
scheduled/durable attempts in bounded workbench aggregates. Measure pool
mutex acquisition separately from query time without capturing payloads
or dynamic tenant/job labels; retain default-off compilation gates.

Let the mixed capacity runner capture cumulative operator snapshots
outside timed phases. Validate contention, live SQLite/PostgreSQL routes,
job attribution, and enabled/disabled configurations independently.
Copilot AI review requested due to automatic review settings September 20, 2026 01:56
@valthon
valthon force-pushed the codex/workbench-operation-attribution branch from d86294e to 254605c Compare September 20, 2026 01:56
@valthon

valthon commented Sep 20, 2026

Copy link
Copy Markdown
Owner Author

Audited the review summary as well as all inline/general feedback. The claim in review #5258788270 about durable attribution extending through acknowledgment is incorrect: in src/queue/durable.zig, the deferred measured.leave() at line 491 runs when the registered-handler if block closes at line 496, before the acknowledgment writer checkout at line 500.

Amended 254605c strengthens the regression at line 1321 with two distinct consecutive job kinds (success and retry). Each retains exactly one handler SQL execution and writer acquisition; acknowledgment state queries and subsequent caller SQL do not enter handler aggregates. No production scope change was needed. Enabled Zig suite: 2,244 passed, 48 skipped. Restacked capacity/workbench plus documentation checks: 85 passed, 19 subtests passed. The earlier admission CI repair remains included. Both previous heads had green CI; the amended heads are rerunning in native stack #466.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

Final review found no unresolved blocking issues.

Review effort: Lite
Findings: None

@valthon
valthon merged commit ccb37d3 into main Sep 20, 2026
31 of 55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants