The open-source core for agentic work environments. Catamorphic brings projects, agents, documents, browser tabs, terminals, apps, and automations into products whose identity and deployment belong to their host.
Work is the desktop, company brain, and mobile product built on Catamorphic. It is designed for anyone with work to do. This repository contains the framework and Work's reference applications.
Work supports Apple silicon Macs running macOS 12 or newer. The current preview is the download on work.software, or:
brew install --cask opencx-labs/tap/work@alphaAfter the first Work Stable is published:
brew install --cask opencx-labs/tap/workYou can also download the signed and notarized DMG from
GitHub Releases and drag
Work into Applications. Installed builds let you choose when to download and
restart. Switch channels under Help > Update Channel, or use brew upgrade
with the same cask you installed. Work keeps a pre-migration database backup.
Work uses a fresh application identity and storage directory. Alpha data and credentials from the previous app are not migrated automatically. Project folders remain ordinary files you can import.
The desktop app (apps/desktop) is a local-first workspace and
the framework's reference implementation. Claude Code, Codex, and API models
can work side by side in durable sessions. Agents use the browser, terminals,
editors, project files, apps, and connectors you can also use yourself. You can
take over a surface at any moment, inspect diffs when a change deserves your
eyes, or trust routine work to continue.
Projects are ordinary folders and git repositories. Every agent turn that changes files creates a checkpoint commit, and any turn can be undone with its files. A session is a durable log of turns that every client folds the same way, so a turn whose machine stopped continues elsewhere on the agent's own conversation. Local execution uses an embedded database and local sandboxes, so the desktop does not depend on a hosted Catamorphic service.
The Work server (apps/server) gives a personal or company
brain an always-on home. Every release publishes it as a multi-architecture
work-server image to the GitHub Container Registry, next to the desktop
release. It serves the mobile PWA, remote agents, project apps,
and MCP endpoints. Desktop and mobile clients sign in through OAuth with PKCE;
the Work server supports local credentials and configured OAuth or OIDC
providers. Invitations grant admission after sign-in, and committed project
roles decide what each member and agent may do.
The zero-service setup keeps PGlite, git origins, credentials, and project data
under one data directory. A deployment can opt into real Postgres as it grows.
The accepted multi-machine architecture uses server instances sharing one
Postgres and one authority, with machines exposed as permitted Environments.
Postgres mode shares origins, artifacts, credentials, auth, and leased execution.
Each Environment can choose its sandbox image, give agents their own Docker,
and limit which hosts they reach; an agent's sandboxing decides what leaves
its sandbox, apart from its harness's own permission mode. See
the machine setup and recovery model.
The Work server is single-tenant because its local-process execution can access
the host machine. Run one trusted organization or household per deployment.
See the server guide or give an agent
skills/setup-work-server
to provision it.
The packages under packages/ let a host application mount
Catamorphic in-process without adopting Catamorphic's identity, deployment, or
design choices. A host supplies auth, users and organizations, database,
storage, execution providers, credentials, telemetry, skills, and doctrine.
The framework supplies general-purpose projects, git-native work tracking,
multi-harness coding agents, durable TypeScript workflows, sandboxed
user-built apps, generated API clients, headless React state, and composable
UI. See INTEGRATION.md.
- A personal brain. Keep notes, research, code, recurring work, and small tools in projects on your Mac. Add a Work server when you want the same projects and conversations from your phone or while the Mac is asleep.
- A company brain. Put shared knowledge, automations, apps, and tuned agent roles in one reviewable project program. Members sign in with their own identity (a company's Google Workspace, with access revoked within minutes when someone leaves) and receive only the roles and project store paths they need. Customers get sign-in links to exactly the documents, folders, or apps shared with them. The same brain is available from desktop, the hosted PWA, an MCP client, or an embedded product surface.
- A daily-driver dev shell. Import a monorepo with its existing agent instructions. Worktrees, terminals, browser tabs, diffs, pull requests, and multiple coding harnesses live in one window.
- An embedded copilot. Mount the libraries in a SaaS backend, add the React surfaces you want, and give the agent the host's skills, tools, trigger kinds, look, and doctrine.
- Per-customer tools over MCP. Give each customer a project where agents build typed workflows and apps, then serve the approved tools to the MCP client they already use.
Desktop and server are designed to compose. A linked project's files sync over git after settled turns. Desktop conversations mirror to the server, so the mobile PWA can open the same chat and continue with the server's always-on agent. If someone continues remotely, the server owns that conversation fork instead of pretending two writers have one history. Incognito desktop chats stay local and are never mirrored.
Projects can shape the shared experience without hard-coding company personas
into the app. Committed role files grant scoped artifacts and project
permissions such as program:write or sessions:read; the shared sidebar and
up to six New Tab starting actions can target the caller's resolved
permissions. If a project does
not configure an action, the desktop adds no placeholder or empty surface.
(ADR 0092)
Code is the source of truth. Everything is stored as plain files (TypeScript, markdown, whatever the work is) in a git repository, never a proprietary DSL or an opaque store. When a project holds workflows, the parser renders the workflow code as an intuitive visual graph for non-technical users, while technical users and AI agents work directly with the code. A project is just a git repo.
Six capabilities, co-equal. Each is shipped and verifiable in this repo; the ADR column is the settled design record.
A project is a folder that can hold documents, notes, data, plans, code,
automations, and apps, in any mix. A blank project is a git repository, a
.work/project.json manifest, and hidden seed skills; nothing in
the visible tree claims the project is about code. The workflow/app
workspace (an independent Bun workspace under .work/: contracts/, workflows/, apps/*) is
scaffolded on demand, the first time someone asks for an automation or
app. Imported repositories are adopted as-is; existing files are never
overwritten. (ADRs 0032,
0043)
Work is tracked without anyone performing git. Every agent turn that
changed files ends in a checkpoint commit, its sha stamped on the chat
message, so every reply's diff is addressable forever. Work never updates
the default branch of a repository it did not create, or any branch there
it did not make: sync fetches and fast-forwards, and local work reaches a
shared repository as a work/ branch plus a pull request. A repository
Work created itself syncs automatically (fetch, fast-forward, merge, push),
and a conflicting divergence lands on a rescue branch instead of a stuck
state. A code host acts through an ordinary connection: the person's own
account, or the organization's service connection (a GitHub App
installation), with tokens minted per call. GitHub is the first code host
(repository import, publish, pull requests, reviews). Agents get explicit
git verbs (sync_project, create_pull_request) rather than raw git.
(ADRs 0044,
0170,
0177)
One agent-session engine over pluggable harnesses: the built-in
@catamorphic/ai-sdk tool loop (any API model), Claude Code
(@catamorphic/claude-code), and Codex (@catamorphic/codex), selectable
per session and switchable mid-session, with a normalized effort scale.
Agents run either in a sandbox or directly on the host filesystem.
ask_user questions, durable sessions that survive restarts, and
per-profile agent rosters work across harnesses. MCP goes both directions:
agents consume MCP connectors (registry search, plugin marketplaces,
elicitation via @catamorphic/mcp), and a project's workflows are served
as MCP tools to any MCP client. In the desktop, agents also drive the
workspace itself: browser tabs (real input, uploads, downloads, and a page's
console and network for web development), terminals, open_surface,
point_at. Local harnesses get the person's login-shell PATH, so they reach
the same tools as their terminal.
(ADRs 0038,
0042)
Concurrent sessions start in the same visible project checkout. Before a
turn, each agent sees the other active sessions in that project and can read
their bounded transcripts. It can keep sharing the folder, wait, or use the
same harness-neutral tools to create or adopt a Git worktree. Per-agent
coordination doctrine ranges from shared-first to isolation-required, so
non-technical roles can stay in one familiar folder while engineering agents
isolate only when needed. Shared sessions deliberately share files, Git
state, checkpoint commits, and rollback; Catamorphic does not pretend those
changes belong to one agent. (ADR
0063)
Project agent definitions: an agent can be a work product. Committed
.work/agents/<slug>.json files (plus an optional .work/agents/<slug>.md persona)
version with the project and appear in every collaborator's picker. A
committed definition never runs on your personal credentials until you
consent, and consent is bound to a hash of what you approved; definitions
using a project secret need no personal consent and work headlessly.
(ADR 0050)
Delegation is also first-class. A subagent works in an ordinary durable child session with explicit delegation authority, its own transcript, and the same message, interrupt, attention, and policy machinery as any other session. Its result reaches the parent like a subagent's: during the parent's turn while it still works, or as a new turn after it. Archive recursively stops active work only after reporting its impact, while keeping the session tree restorable and searchable. Session source records whether a conversation began on desktop, mobile, Slack, Claude, MCP, or API. (ADRs 0089, 0090)
TypeScript automations in git, rendered as a visual graph for
non-technical users, executed durably on Postgres. One model: every
workflow is an exported defineWorkflow value, every run executes a
deployed commit. Boundaries (atomic retry scopes), batch scopes, pauses
and signals, correlation keys, shared rate budgets, retention, and
triggers, including host-defined trigger kinds with typed payloads and
sync-until-first-wait firing, project-defined kinds, and declarative where
filters. Full authoring model below.
Role access is separate from unattended consent. Each member previews and enables an exact deployed workflow with its trigger, Environment, agent, and connection requirements. Connecting an account may complete that chosen flow; it never silently enables every compatible automation.
Real frontends users (and agents) build on top of workflows: sandboxed
React bundles wired through a typed contract that cannot drift from the
workflows it calls (same repo, same commit). An app's callable workflow
set is frozen per published version and re-authorized on every call. Apps
interoperate with MCP Apps in both directions: the desktop renders MCP
Apps from connectors, and /projects/:id/apps-mcp serves your apps to MCP
hosts like Claude, unchanged. Apps get persistent app-local storage per
(app, user), and the @catamorphic/app/ui kit gives agent-built apps
polished, accessible components with zero CSS.
(ADRs 0035 through
0037,
0048)
Import a real monorepo and use the desktop as your daily driver. The
Claude Code harness runs at full fidelity: the SDK's own preset system
prompt, and the repo's CLAUDE.md, .claude/ skills, agents, commands, and
settings load exactly as in the CLI. Worktrees are first-class: discovered,
listed, diffed, and assignable by agents when concurrent work needs
isolation. Checkout management is harness-neutral, so Claude Code, Codex,
and the built-in agent follow the same policy. Diff tabs render in Monaco;
the sidebar has Changes and Pull Requests sections; PR review opens per-file
diffs through the CodeHost seam. Terminals are real PTYs with shell
integration (OSC 133); the embedded browser, which installs Chrome
extensions from the Chrome Web Store, and the command palette round out the
shell. All of it degrades quietly for non-technical users.
(ADRs 0045,
0063,
0203)
The signed app ships the audited Claude Code and Codex adapters, then downloads each large, platform-specific executable only on first use. Every component is versioned and SHA-512 pinned by the app release, installed atomically, and reused offline. (ADR 0091)
Apps and agents inside an embedder's product are unmistakably the
embedder's. Feel: the app kit ships structure and behavior only; every
aesthetic decision flows from host theme tokens with neutral defaults,
plus hostCss and kit: false for total control. Doctrine: two
createCatamorphic hooks (projectSeeds, standingAgentPrompt) receive
the framework defaults and return the host-final set, so seeded skills
and the agents' standing prompt are both replaceable. The desktop
consumes the same hooks and passes nothing: the proof the defaults are
real defaults.
(ADRs 0048,
0049)
Catamorphic's engine ships as libraries a host application mounts
in-process. The desktop app is itself such a host. The host provides auth,
the user/org model, the database, and the deployment surface. There is no
default identity or tenant: every request carries identity from the host's
auth context. See INTEGRATION.md for the host
integration flow.
Durable agent-and-workflow infrastructure that does not assume an external server. Every dependency is an axis with a heavy and a light end. Pick per axis:
| Axis | Heavy end | Light end |
|---|---|---|
| Database | Network Postgres ({ pool } / { connectionString }) |
Embedded pglite (Kysely PGliteDialect via database: { db }; migrations run statement-by-statement so single-connection dialects just work) |
| Execution | Cloud sandboxes: @catamorphic/cloudflare, @catamorphic/daytona |
Local sandboxes (@catamorphic/microsandbox), plain local processes (@catamorphic/local-process, trusted single-tenant hosts only), or none (read-only embed) |
| Code storage | S3-compatible bucket (@catamorphic/s3) or Cloudflare Artifacts |
Two writable directories |
| Identity | Host org/user per request | One fixed tenant/user |
| Surface | HTTP API + React UI | In-process SDK calls, or migrations-only |
The desktop app is the proof: it runs the lightest column end to end
(pglite, local sandboxes, filesystem storage, no external server) by design
(apps/desktop/src/main/server/boot.ts).
No durable-execution vendor can run entirely inside a desktop app; this one
does, and the same substrate is what offline-first agents need. Full matrix
and host shapes: INTEGRATION.md.
Also worth knowing, because it's easy to miss from the package list:
- The product teaches agents from the inside. Every project is seeded
with hidden skills (
.work/skills/): the project model and on-demand workspace scaffold (work-projects), workflow authoring (writing-workflows,batch-workflows,durable-workflows), and app building split into mechanics (building-apps) and replaceable design doctrine (designing-apps). Coding agents learn Catamorphic's authoring model at the moment they need it. The publicskills/setup-work-serverskill extends the same idea to stock-server setup and integrating Catamorphic into an existing app without replacing its auth or deployment. - One Run model. Boundary and batch workflows all share the same Runs API, hooks, and UI: capabilities, not categories. Every run executes a deployed commit.
- Observability is free for hosts. Everything instruments against
@opentelemetry/api; register your SDK and Catamorphic's spans appear in your traces.
| Package | What it is |
|---|---|
@catamorphic/server-sdk |
The core SDK for your Node/Bun backend. Takes a Postgres connection (or pg.Pool), manages its own schema-scoped tables and migrations, and exposes projects, workflows, files, runs, triggers, agent sessions, connections, and code hosts (githubCodeHost over the github connection provider). |
@catamorphic/fastify-plugin |
A mountable Fastify plugin (app.register(catamorphicPlugin, { core, prefix: "/api" })) exposing the standard HTTP API for frontends, plus the per-project MCP endpoints. Also exports a standalone createApp factory for sidecar deployments. |
@catamorphic/react |
Headless React bindings: CatamorphicProvider, TanStack Query hooks, and jotai atoms. Build a fully custom UI on top of these. |
@catamorphic/ui |
Ready-made components: the React Flow workflow canvas, member review and consent, the Runs panel, and AppMount (the sandboxed app iframe host). Every piece is opt-in. |
@catamorphic/registry |
shadcn-style copy-paste components for hosts that want to own and customize the component source (project browser, git panel, runs panel, agent chat, Monaco editor). |
@catamorphic/api-client |
Generated OpenAPI types + openapi-fetch client for the HTTP API. |
@catamorphic/workflow |
Typed workflow-authoring primitives. Projects opt in directly, or a SaaS can wrap it and re-export only its approved surface. |
@catamorphic/app |
The guest-side app runtime bundled into every user-built app: typed workflow client, persistent app-local storage shim, dual-dialect MCP Apps support, and the @catamorphic/app/ui component kit styled entirely by host theme tokens. |
Supporting packages (consumed through the surface above, importable directly for advanced wiring):
| Package | What it is |
|---|---|
@catamorphic/core |
Framework-agnostic service layer: projects, workflows, runs, deployments, triggers, apps, app storage, plugins, secrets, agent sessions, agent definitions, remote sync, and the CodeHost seam. The kernel behind server-sdk and fastify-plugin. |
@catamorphic/db |
Kysely + Postgres. Schema-scoped (default schema catamorphic), raw SQL migrations, programmatic migrateToLatest. |
@catamorphic/git |
Git-backed project storage (isomorphic-git): members' drafts as refs in the project origin (ADR 0191), local folder checkouts, pluggable origin remotes (RemoteBackend), and the remote sync engine (syncWithNetworkRemote: fetch and fast-forward; merge, push, and rescue branches only for repositories Work created), with the one push guard every network push passes (work/ branches only on attached repositories). |
@catamorphic/github |
GitHub mechanics: OAuth + device-flow helpers, GitHub App auth (app JWTs, installation tokens, manifest registration), and the REST API client. The server SDK builds the github connection provider and code host on it. |
@catamorphic/parser |
ts-morph AST → WorkflowGraph parser + dagre layout; also powers the seeded project check script. |
@catamorphic/sandbox |
Vendor-neutral sandbox contracts (SandboxProvider, SandboxManager, RunExecutor), the stdio supervisor transport, OTel instrumentation, and plugin-doc staging and tool-policy helpers for harnesses. |
@catamorphic/agent-protocol |
The agent session log (turns, attempts, items, requests, provider threads), its events, commands and shared reducer, and the runner protocol every harness adapter speaks (HarnessAdapter). |
@catamorphic/agent-runner |
The harness-agnostic runner that drives one attempt of a turn on an adapter, in-process or over stdio. |
@catamorphic/runner-bundle |
The runner and its Claude Code and Codex adapters as one hash-addressed file a sandbox runs with Bun or Node. |
@catamorphic/microsandbox |
Local sandbox provider over the microsandbox SDK: the desktop's default execution. |
@catamorphic/local-process |
Sandboxless execution as plain subprocesses with an explicit env. Trusted single-tenant hosts only (ADR 0047). |
@catamorphic/cloudflare |
Cloudflare backend plugin: CloudflareSandboxProvider (execution via Bridge Worker) + ArtifactsRemoteBackend (Cloudflare-native code storage when available). |
@catamorphic/s3 |
S3-compatible git origin backend for Cloudflare R2, AWS S3, MinIO, and similar stores. |
@catamorphic/daytona |
Daytona backend plugin: DaytonaSandboxProvider + experimental Daytona git storage. |
@catamorphic/ai-sdk |
Built-in harness adapter: Vercel AI SDK tool loop on any API model, running in the host and driving the session's sandbox through its tools. |
@catamorphic/claude-code |
Harness adapter backed by the Claude Code (Claude Agent SDK) CLI, with per-turn MCP servers, portable native state and full settings-source fidelity. |
@catamorphic/codex |
Harness adapter backed by the pinned OpenAI Codex app-server protocol. |
@catamorphic/mcp |
MCP client infrastructure: both protocol generations with auto-negotiation, elicitation, the official MCP registry search, and plugin-marketplace install. |
@catamorphic/otel |
Tiny OpenTelemetry helpers (@opentelemetry/api only: the host owns the SDK/exporters). |
@catamorphic/runtime |
Execution harness that runs inside the sandbox and reports step results. |
@catamorphic/plugins |
Plugin manifest contract + resolvers for host-provided packages and secrets. |
@catamorphic/cloudflare-sandbox-bridge |
Deployable Cloudflare Worker exposing Cloudflare Sandbox over HTTP. |
The in-repo reference host is the Catamorphic desktop app (apps/desktop): an Electron app that embeds the server in-process (src/main/server/boot.ts).
- Workflows are regular code. User-defined workflows run like normal apps: full IO, real npm dependencies, no crippled JS runtime. Execution happens inside a sandbox (or a local process, where the host shape allows it) using Bun to run and bundle.
- Code stays simple. Both AI agents and humans must be able to write, edit, and understand workflows, and the parser must render them intuitively for non-technical users. See the code format below.
- Host-injectable everything. Database connections, storage backends, sandbox credentials, LLM credentials, telemetry, trigger kinds, seeds, and doctrine are all injected by the host: nothing is hard-coded.
- Postgres is authoritative for runs, retries, pauses, batch-item state, queues, and scheduling via
SKIP LOCKED. Cloudflare Sandbox is the default cloud execution provider; backends ship as vendor plugin packages so hosts install only what they use. - OpenTelemetry throughout. Libraries instrument against
@opentelemetry/apionly; the host registers the SDK and exporters and gets full traces for free.
Settled design decisions are recorded as ADRs in docs/decisions/.
import { CloudflareSandboxProvider } from "@catamorphic/cloudflare";
import {
createCatamorphic,
defineStaticEnvironments,
} from "@catamorphic/server-sdk";
const sandboxProvider = new CloudflareSandboxProvider({
apiUrl: process.env.CLOUDFLARE_SANDBOX_API_URL!,
apiKey: process.env.CLOUDFLARE_SANDBOX_API_KEY,
});
const environmentProvider = defineStaticEnvironments([
{
descriptor: {
id: "local",
label: "Managed execution",
trust: "managed",
isolation: "sandbox",
workloads: ["agent", "workflow"],
agentTopologies: ["controller"],
capabilities: ["network.egress"],
resources: {},
},
sandboxProvider,
},
]);
// Boot once per process
const catamorphic = createCatamorphic({
database: { connectionString: process.env.DATABASE_URL! }, // or { pool }
storage: {
projectsPath: process.env.CATAMORPHIC_PROJECTS_PATH!,
remotesPath: process.env.CATAMORPHIC_REMOTES_PATH!,
},
// Backend plugins: @catamorphic/cloudflare, @catamorphic/daytona,
// @catamorphic/microsandbox, or @catamorphic/local-process
sandboxProvider,
environmentProvider,
});
await catamorphic.migrate(); // idempotent, schema-scoped
// Worker startup is explicit and host-owned.
const executionWorker = catamorphic.startExecutionWorker({ concurrency: 4 });
// Per request: bind the host's org + user
const client = catamorphic
.forTenant({ tenantId: orgId })
.forUser({ externalUserId: userId });
const project = await client.projects.create({ name: "Onboarding" });
const run = await client.runs.triggerProduction({
projectId: project.id,
workflowName: "welcomeUser",
input: { email: "ada@example.com" },
});To expose the HTTP API to your frontend:
import { catamorphicPlugin } from "@catamorphic/fastify-plugin";
app.register(catamorphicPlugin, { core: catamorphic.core, prefix: "/api" });And on the frontend, wrap your tree with CatamorphicProvider from @catamorphic/react and drop in WorkflowEditor from @catamorphic/ui (or build your own UI from the hooks). See INTEGRATION.md.
There is one public Workflow model and one public Run model: every workflow is
an exported defineWorkflow(({ defineBoundary, defineBatch }) => ({ steps: [...] })) value from @catamorphic/workflow, and every run executes a
deployed commit.
A boundary is one atomic retry scope: if its callback fails, all operations in
that callback retry together. Orchestration code lives in boundary run
bodies; IO and business operations live in "use step" functions: plain
async functions with the exact "use step" directive, called from boundary
bodies. All functions take one destructured object parameter and carry JSDoc
display metadata.
import { type BoundaryContext, defineWorkflow } from "@catamorphic/workflow";
/**
* @displayname Welcome New User
* @description Onboard a new user
*/
export const welcomeUser = defineWorkflow(({ defineBoundary }) => ({
steps: [
defineBoundary({
run: async ({
input,
}: BoundaryContext<{ email: string; name: string }>) => {
const user = await createUser({ email: input.email, name: input.name });
await sendWelcomeEmail({ to: user.email, name: user.name });
if (user.plan === "premium") {
await assignPremiumBenefits({ userId: user.id });
}
await sendFollowUpEmail({ to: user.email });
return { status: "complete", userId: user.id };
},
}),
],
}));A batch scope is paged per-item processing with an optional sink. defineBatchStep is a
physical coalescing primitive for compatible calls inside defineBatch.process;
it does not define a Workflow or a separate logical step scope.
import { defineWorkflow } from "@catamorphic/workflow";
export const processAccount = defineWorkflow(
({ defineBoundary, defineBatch }) => ({
steps: [
defineBoundary<{ accountId: string }, { accountId: string }>({
retry: { maxAttempts: 3 },
run: async ({ input }) => prepareAccount({ accountId: input.accountId }),
}),
defineBatch({
source: async ({ input }: { input: { accountId: string } }) => ({
source: recordsSource,
config: { accountId: input.accountId },
}),
process: async ({ item }: { item: AccountRecord }) =>
processRecord({ record: item }),
sink: resultSink,
}),
],
}),
);The parser and UI expose capabilities such as batch processing and cancellation. There is no public stage concept, category switch, or separate Run family for these capabilities. A workflow with no pause, retry backoff, rate limit, batch, or child call settles inline when triggered synchronously: inline request-response is a property of execution, not a separate authoring form.
Workflows subscribe to trigger kinds in code (triggers: [trigger("kind", config)]). Hosts define their own kinds with
defineTriggerKind (typed payloads and configs via zod, generated
work-triggers.d.ts per project) and fire them sync or async; a kind
whose payload varies per workflow uses typed holes, and kinds declared via
mcpToolKinds are served as MCP tools from POST /projects/:id/mcp.
Projects define their own kinds on top of the host's in .work/triggers/
(defineTrigger({ name, from: trigger("webhook", { ... }), where })), and
every binding may add a where filter the host checks before a run starts,
so a GitHub or Slack integration is project code on the webhook kind, whose
verification schemes and handshakes are declared in its config. The host
skill slack is the worked example: a chat per Slack thread, replies posted
back through a gateway connection whose named operations are its
capabilities (Connect Slack).
The host skill reviewing-pull-requests is the larger one: a code and
security review of every pull request, in a chat per pull request on a
dedicated review pool, verified by running the change and posted as a review
and a check run
(Review pull requests).
See INTEGRATION.md and ADRs
0039 /
0042 /
0171 /
0179 /
0181.
A run may carry a correlation key: a host-meaningful identity for its
subject (a contact, an account, a subscription). It is unique among live runs of
the same workflow, which makes it at once an enrollment idempotency key and the
address external events use to reach the run. A boundary declares the shared
third-party budgets it draws on; buckets are keyed per tenant, so every workflow
naming the same globalKey draws on one budget.
export const nurtureContact = defineWorkflow(({ defineBoundary }) => ({
controls: { cancel: true },
steps: [
defineBoundary({
rateLimits: [
{ globalKey: "email", capacity: 500, refillRatePerSecond: 100 },
],
run: async ({ input, pause }: BoundaryContext<{ contactId: string }>) => {
await sendWelcomeEmail({ contactId: input.contactId });
return pause<{ clicked: boolean }, { contactId: string }>({
signal: "reply",
timeout: "72h",
state: { contactId: input.contactId },
});
},
}),
defineBoundary({
// A different step, a different provider, a different budget.
rateLimits: [
{ globalKey: "whatsapp", partitionKey: "sender-1", capacity: 80, refillRatePerSecond: 20 },
],
run: async ({ input }) => sendWhatsApp({ contactId: input.state.contactId }),
}),
],
}));// Enrolling twice for the same contact is a no-op, so webhook redelivery is safe.
await client.runs.triggerProduction({
projectId, workflowName: "nurtureContact",
input: { contactId: "contact-42" },
correlationKey: "contact-42", // onConflict: "ignore" | "error" | "restart"
});
// External events address the contact, not a run id.
await client.runs.signalByKey({
projectId, workflowName: "nurtureContact",
correlationKey: "contact-42",
signal: "reply", idempotencyKey: replyEventId, value: { clicked: true },
});
// Opting out ends the journey wherever it sits. Resolves null if none was live.
await client.runs.cancelByKey({
projectId, workflowName: "nurtureContact",
correlationKey: "contact-42", reason: "unsubscribed",
});Waiting for capacity is not failure: a boundary that cannot reserve is
rescheduled without consuming a retry and without holding a sandbox. When a
provider answers with a 429, report it with rateLimited({ retryAfterMs }):
that blocks every workflow sharing the account, not just the run that hit it.
Hosts bound what any one tenant may consume from these shared resources via
catamorphic.tenantPolicies. It is SDK-only and never exposed over HTTP, so a
tenant cannot raise its own limits:
await catamorphic.tenantPolicies.upsert({
tenantId,
maxConcurrentJobs: 32, // ceiling on simultaneously leased jobs
maxActiveRuns: 50_000, // ceiling on live runs; bounds enrollment fan-out
queueWeight: 4, // relative share of each claim batch
retentionDays: 365, // overrides the installation retention window
rateLimitOverrides: { // can tighten an author's bucket, never loosen it
whatsapp: { capacity: 40, refillRatePerSecond: 10 },
},
});Finished runs are purged after 90 days by default, along with everything hanging off them: jobs, events, step attempts, batch items and their steps. Without this, a daily 100k-item batch adds on the order of 1.5M rows a day and never gives any back.
The sweep runs inside the execution worker, so it needs no wiring. Change the window, or turn it off entirely, at construction:
const catamorphic = createCatamorphic({
database,
storage,
environmentProvider,
retention: { runRetentionDays: 30 }, // or { enabled: false } to keep forever
});Individual tenants can be given a longer or shorter window with
retentionDays above.
See ADR 0027, ADR 0028, and ADR 0030.
Catamorphic is issue-first. If you do not have repository write access, open a GitHub issue with the problem, concrete use case, constraints, and desired outcome. Public pull request creation is restricted to the core team and collaborators with write access.
AI makes producing code cheap, but it does not make a large generated change cheap to review. We review the compact source of intent first, then maintainers own the implementation with the repository's full context. Read CONTRIBUTING.md for the reasoning and issue guidance.
bun install
# Only for an explicitly provisioned host database. These are not part of the
# normal local app workflow; never point them at a database another session uses.
DATABASE_URL="<host-owned database URL>" bun run db:migrate
DATABASE_URL="<host-owned database URL>" bun run db:codegen
# Build everything
bun run build
# Regenerate the OpenAPI spec + typed api-client after route/DTO changes
cd packages/fastify-plugin && bun run generate-spec
cd ../api-client && bun run generateOptional shared OTel, ClickHouse, and sandbox-bridge services run persistently. Start them in a separate terminal only when you need them:
bun run dev:infraThe framework packages are embed-only. In production, run either the stock
server host or an application that boots Catamorphic in-process. For local
development, bun run dev starts the combined desktop and stock-server manual
environment. bun run dev:desktop and bun run dev:server are focused
variants of the same orchestrator, which assigns each worktree its own data
directories and loopback ports. To iterate on Catamorphic alongside your own
host, link the packages via file: (see
.agents/skills/using-catamorphic/SKILL.md and its "Local dev linking"
section).
bun install # Install dependencies in this worktree
bun run dev # Combined desktop and stock-server manual environment
bun run dev:desktop # Desktop-focused development environment
bun run dev:server # Stock-server-focused development environment
bun run dev:infra # Optional OTel, ClickHouse, and sandbox-bridge services
bun run build # Build all packages
bun run test # Deterministic root and Postgres-complete workspace tests
bun run test:external # Explicit opt-in for credentialed external integrations
bun run check # 11-phase merge gate
bun run typecheck # Typecheck all packages with tsgo
bun run lint # Lint with Biome
bun run lint:fix # Auto-fix lint issues
bun run db:migrate # Apply migrations to the DB pointed at by DATABASE_URL
bun run db:codegen # Regenerate Kysely types from the `catamorphic` schema
bun run db:reset # Drop + recreate the catamorphic schema (dev only)
bun run db:status # Show applied / pending migrationsManual development profiles live under ~/.catamorphic/dev/<worktree-instance>.
The runner assigns separate desktop/server profiles and ports per worktree.
Legacy temporary profiles are copied on first launch, with the originals retained.
Do not put manual project databases in the operating system temporary directory.
Tests use Vitest, orchestrated by Turborepo. Docker must be running:
bun run test and bun run check create an isolated disposable Postgres
database for every invocation. They do not need bun run dev:infra; that
command is optional shared observability and sandbox-bridge infrastructure,
not test infrastructure.
bun run test # deterministic root and Postgres-complete workspace tests
bun run check # 11 phases: lint, root/workspace types, build, migrations, root/workspace tests, and E2E
bun run test:external # explicit authority for credentialed integrations
bun run --filter @catamorphic/parser test # one package through repository Node
bun run --cwd packages/parser test src/__tests__/parser.test.ts # one file through repository NodeEach test file uses fresh temporary state and every root test invocation gets
its own Postgres database and temporary caches, so parallel worktrees do not
share test output. Install dependencies in each worktree and never copy a
credentialed .env wholesale. Add only the individual settings a task needs.
Root orchestration tests under scripts/**/*.test.ts run before the package
graph. Package-local Vitest commands also select the repository-pinned Node,
including when the ambient node is Node 20.
External integration tests never run merely because credentials are present.
They require the explicit bun run test:external authority, which sets
CATAMORPHIC_EXTERNAL_INTEGRATIONS=1, and skip when their service-specific
configuration is absent:
- Daytona (
packages/daytona): needsDAYTONA_API_KEY. - Cloudflare Sandbox (
packages/cloudflare,packages/core): needsCLOUDFLARE_SANDBOX_API_URL,CLOUDFLARE_SANDBOX_API_KEY, andCF_SANDBOX_INTEGRATION=1. Start the sandbox bridge separately if the test needs it. - Cloudflare Artifacts (
packages/cloudflare): needsCLOUDFLARE_ACCOUNT_ID,CLOUDFLARE_API_TOKEN, andCLOUDFLARE_ARTIFACTS_NAMESPACE; it skips with a warning while feature-gated. - S3-compatible storage (
packages/s3): needsS3_BUCKET,S3_ACCESS_KEY_ID, andS3_SECRET_ACCESS_KEY. SetS3_REGIONwhen the provider requires it,S3_ENDPOINTfor R2, MinIO, or another non-AWS endpoint, andS3_FORCE_PATH_STYLE=truewhen that endpoint requires path-style requests.
Unit tests run with no setup. bun run --cwd apps/desktop test:e2e runs the
complete desktop suite in Docker, using real Electron on a private Xvfb display
with Openbox and a deterministic fake agent. Windows use normal focus and
rendering without reaching your desktop. The first run builds a cached Linux
image; subsequent runs reuse dependencies. File filters and Vitest shard flags
pass through. Logs, screenshots, and JUnit results go to test-results/.
bun run check includes this isolated suite. Native macOS coverage runs in CI.
See desktop testing for runner and sharding details.
PR checks reuse successful package and shard tasks by content hash. Changing a package invalidates its downstream consumers; local commands and main still execute tests fresh. See test caching and selective execution for cache inputs, failure handling, and execution summaries.
The local development runner uses POSIX process groups and supports macOS and Linux. Windows development orchestration is not supported.
- Bun: runtime, package manager, bundler (also inside sandboxes)
- TypeScript: tsgo for typechecking
- Postgres: all state, schema-scoped; run queues, retries, pauses, and scheduling use the same DB
- Electron + electron-vite: the desktop reference implementation
- Fastify: HTTP surface with Zod + OpenAPI (mounted by the host)
- React Flow: workflow visualization; Jotai + TanStack Query: frontend state
- ts-morph: TypeScript AST parsing
- Kysely: type-safe SQL
- isomorphic-git: portable git, no native CLI dependency (system git is a soft dependency for the desktop's read surfaces)
- OpenTelemetry:
@opentelemetry/apiinstrumentation throughout - Turborepo: build orchestration
Direction, not shipped. Tracked in TODO.md:
- More desktop platforms: Intel Mac validation, then signed Windows and Linux packages.
- ACP harness: project agent definitions already accept
kind: "acp"; the Agent Client Protocol client (local command and remote endpoint transports) is the planned harness behind it. - TS
defineAgent: a typed authoring layer that compiles to the committed.work/agents/<slug>.jsonsubstrate. - Review-mode collaboration: PR-first sync for shared projects, invites, native PR review depth (comments, approvals, merge).
- Remote blob storage: user-connected stores (S3/R2/Drive-style) for large binaries beside git-tracked text.
