Skip to content

Latest commit

 

History

564 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Catamorphic: a really good place to get work done

Catamorphic

The open-source core for agentic work environments. Catamorphic brings projects, agents, documents, browser tabs, terminals, apps, and automations into products whose identity and deployment belong to their host.

Work is the desktop, company brain, and mobile product built on Catamorphic. It is designed for anyone with work to do. This repository contains the framework and Work's reference applications.

Install Work

Work supports Apple silicon Macs running macOS 12 or newer. The current preview is the download on work.software, or:

brew install --cask opencx-labs/tap/work@alpha

After the first Work Stable is published:

brew install --cask opencx-labs/tap/work

You can also download the signed and notarized DMG from GitHub Releases and drag Work into Applications. Installed builds let you choose when to download and restart. Switch channels under Help > Update Channel, or use brew upgrade with the same cask you installed. Work keeps a pre-migration database backup.

Work uses a fresh application identity and storage directory. Alpha data and credentials from the previous app are not migrated automatically. Project folders remain ordinary files you can import.

One system, three ways to use it

Desktop app

The desktop app (apps/desktop) is a local-first workspace and the framework's reference implementation. Claude Code, Codex, and API models can work side by side in durable sessions. Agents use the browser, terminals, editors, project files, apps, and connectors you can also use yourself. You can take over a surface at any moment, inspect diffs when a change deserves your eyes, or trust routine work to continue.

Projects are ordinary folders and git repositories. Every agent turn that changes files creates a checkpoint commit, and any turn can be undone with its files. A session is a durable log of turns that every client folds the same way, so a turn whose machine stopped continues elsewhere on the agent's own conversation. Local execution uses an embedded database and local sandboxes, so the desktop does not depend on a hosted Catamorphic service.

Self-hosted server

The Work server (apps/server) gives a personal or company brain an always-on home. Every release publishes it as a multi-architecture work-server image to the GitHub Container Registry, next to the desktop release. It serves the mobile PWA, remote agents, project apps, and MCP endpoints. Desktop and mobile clients sign in through OAuth with PKCE; the Work server supports local credentials and configured OAuth or OIDC providers. Invitations grant admission after sign-in, and committed project roles decide what each member and agent may do.

The zero-service setup keeps PGlite, git origins, credentials, and project data under one data directory. A deployment can opt into real Postgres as it grows. The accepted multi-machine architecture uses server instances sharing one Postgres and one authority, with machines exposed as permitted Environments. Postgres mode shares origins, artifacts, credentials, auth, and leased execution. Each Environment can choose its sandbox image, give agents their own Docker, and limit which hosts they reach; an agent's sandboxing decides what leaves its sandbox, apart from its harness's own permission mode. See the machine setup and recovery model. The Work server is single-tenant because its local-process execution can access the host machine. Run one trusted organization or household per deployment. See the server guide or give an agent skills/setup-work-server to provision it.

Embeddable framework

The packages under packages/ let a host application mount Catamorphic in-process without adopting Catamorphic's identity, deployment, or design choices. A host supplies auth, users and organizations, database, storage, execution providers, credentials, telemetry, skills, and doctrine. The framework supplies general-purpose projects, git-native work tracking, multi-harness coding agents, durable TypeScript workflows, sandboxed user-built apps, generated API clients, headless React state, and composable UI. See INTEGRATION.md.

Personal and company brains

  • A personal brain. Keep notes, research, code, recurring work, and small tools in projects on your Mac. Add a Work server when you want the same projects and conversations from your phone or while the Mac is asleep.
  • A company brain. Put shared knowledge, automations, apps, and tuned agent roles in one reviewable project program. Members sign in with their own identity (a company's Google Workspace, with access revoked within minutes when someone leaves) and receive only the roles and project store paths they need. Customers get sign-in links to exactly the documents, folders, or apps shared with them. The same brain is available from desktop, the hosted PWA, an MCP client, or an embedded product surface.
  • A daily-driver dev shell. Import a monorepo with its existing agent instructions. Worktrees, terminals, browser tabs, diffs, pull requests, and multiple coding harnesses live in one window.
  • An embedded copilot. Mount the libraries in a SaaS backend, add the React surfaces you want, and give the agent the host's skills, tools, trigger kinds, look, and doctrine.
  • Per-customer tools over MCP. Give each customer a project where agents build typed workflows and apps, then serve the approved tools to the MCP client they already use.

Desktop and server are designed to compose. A linked project's files sync over git after settled turns. Desktop conversations mirror to the server, so the mobile PWA can open the same chat and continue with the server's always-on agent. If someone continues remotely, the server owns that conversation fork instead of pretending two writers have one history. Incognito desktop chats stay local and are never mirrored.

Projects can shape the shared experience without hard-coding company personas into the app. Committed role files grant scoped artifacts and project permissions such as program:write or sessions:read; the shared sidebar and up to six New Tab starting actions can target the caller's resolved permissions. If a project does not configure an action, the desktop adds no placeholder or empty surface. (ADR 0092)

Code is the source of truth. Everything is stored as plain files (TypeScript, markdown, whatever the work is) in a git repository, never a proprietary DSL or an opaque store. When a project holds workflows, the parser renders the workflow code as an intuitive visual graph for non-technical users, while technical users and AI agents work directly with the code. A project is just a git repo.


What's inside

Six capabilities, co-equal. Each is shipped and verifiable in this repo; the ADR column is the settled design record.

Projects hold any work

A project is a folder that can hold documents, notes, data, plans, code, automations, and apps, in any mix. A blank project is a git repository, a .work/project.json manifest, and hidden seed skills; nothing in the visible tree claims the project is about code. The workflow/app workspace (an independent Bun workspace under .work/: contracts/, workflows/, apps/*) is scaffolded on demand, the first time someone asks for an automation or app. Imported repositories are adopted as-is; existing files are never overwritten. (ADRs 0032, 0043)

Git-native work tracking

Work is tracked without anyone performing git. Every agent turn that changed files ends in a checkpoint commit, its sha stamped on the chat message, so every reply's diff is addressable forever. Work never updates the default branch of a repository it did not create, or any branch there it did not make: sync fetches and fast-forwards, and local work reaches a shared repository as a work/ branch plus a pull request. A repository Work created itself syncs automatically (fetch, fast-forward, merge, push), and a conflicting divergence lands on a rescue branch instead of a stuck state. A code host acts through an ordinary connection: the person's own account, or the organization's service connection (a GitHub App installation), with tokens minted per call. GitHub is the first code host (repository import, publish, pull requests, reviews). Agents get explicit git verbs (sync_project, create_pull_request) rather than raw git. (ADRs 0044, 0170, 0177)

Coding agents, multi-harness

One agent-session engine over pluggable harnesses: the built-in @catamorphic/ai-sdk tool loop (any API model), Claude Code (@catamorphic/claude-code), and Codex (@catamorphic/codex), selectable per session and switchable mid-session, with a normalized effort scale. Agents run either in a sandbox or directly on the host filesystem. ask_user questions, durable sessions that survive restarts, and per-profile agent rosters work across harnesses. MCP goes both directions: agents consume MCP connectors (registry search, plugin marketplaces, elicitation via @catamorphic/mcp), and a project's workflows are served as MCP tools to any MCP client. In the desktop, agents also drive the workspace itself: browser tabs (real input, uploads, downloads, and a page's console and network for web development), terminals, open_surface, point_at. Local harnesses get the person's login-shell PATH, so they reach the same tools as their terminal. (ADRs 0038, 0042)

Concurrent sessions start in the same visible project checkout. Before a turn, each agent sees the other active sessions in that project and can read their bounded transcripts. It can keep sharing the folder, wait, or use the same harness-neutral tools to create or adopt a Git worktree. Per-agent coordination doctrine ranges from shared-first to isolation-required, so non-technical roles can stay in one familiar folder while engineering agents isolate only when needed. Shared sessions deliberately share files, Git state, checkpoint commits, and rollback; Catamorphic does not pretend those changes belong to one agent. (ADR 0063)

Project agent definitions: an agent can be a work product. Committed .work/agents/<slug>.json files (plus an optional .work/agents/<slug>.md persona) version with the project and appear in every collaborator's picker. A committed definition never runs on your personal credentials until you consent, and consent is bound to a hash of what you approved; definitions using a project secret need no personal consent and work headlessly. (ADR 0050)

Delegation is also first-class. A subagent works in an ordinary durable child session with explicit delegation authority, its own transcript, and the same message, interrupt, attention, and policy machinery as any other session. Its result reaches the parent like a subagent's: during the parent's turn while it still works, or as a new turn after it. Archive recursively stops active work only after reporting its impact, while keeping the session tree restorable and searchable. Session source records whether a conversation began on desktop, mobile, Slack, Claude, MCP, or API. (ADRs 0089, 0090)

Durable workflows

TypeScript automations in git, rendered as a visual graph for non-technical users, executed durably on Postgres. One model: every workflow is an exported defineWorkflow value, every run executes a deployed commit. Boundaries (atomic retry scopes), batch scopes, pauses and signals, correlation keys, shared rate budgets, retention, and triggers, including host-defined trigger kinds with typed payloads and sync-until-first-wait firing, project-defined kinds, and declarative where filters. Full authoring model below.

Role access is separate from unattended consent. Each member previews and enables an exact deployed workflow with its trigger, Environment, agent, and connection requirements. Connecting an account may complete that chosen flow; it never silently enables every compatible automation.

Apps

Real frontends users (and agents) build on top of workflows: sandboxed React bundles wired through a typed contract that cannot drift from the workflows it calls (same repo, same commit). An app's callable workflow set is frozen per published version and re-authorized on every call. Apps interoperate with MCP Apps in both directions: the desktop renders MCP Apps from connectors, and /projects/:id/apps-mcp serves your apps to MCP hosts like Claude, unchanged. Apps get persistent app-local storage per (app, user), and the @catamorphic/app/ui kit gives agent-built apps polished, accessible components with zero CSS. (ADRs 0035 through 0037, 0048)

The dev shell (desktop)

Import a real monorepo and use the desktop as your daily driver. The Claude Code harness runs at full fidelity: the SDK's own preset system prompt, and the repo's CLAUDE.md, .claude/ skills, agents, commands, and settings load exactly as in the CLI. Worktrees are first-class: discovered, listed, diffed, and assignable by agents when concurrent work needs isolation. Checkout management is harness-neutral, so Claude Code, Codex, and the built-in agent follow the same policy. Diff tabs render in Monaco; the sidebar has Changes and Pull Requests sections; PR review opens per-file diffs through the CodeHost seam. Terminals are real PTYs with shell integration (OSC 133); the embedded browser, which installs Chrome extensions from the Chrome Web Store, and the command palette round out the shell. All of it degrades quietly for non-technical users. (ADRs 0045, 0063, 0203)

The signed app ships the audited Claude Code and Codex adapters, then downloads each large, platform-specific executable only on first use. Every component is versioned and SHA-512 pinned by the app release, installed atomically, and reused offline. (ADR 0091)

Your product, your feel and doctrine (embedding)

Apps and agents inside an embedder's product are unmistakably the embedder's. Feel: the app kit ships structure and behavior only; every aesthetic decision flows from host theme tokens with neutral defaults, plus hostCss and kit: false for total control. Doctrine: two createCatamorphic hooks (projectSeeds, standingAgentPrompt) receive the framework defaults and return the host-final set, so seeded skills and the agents' standing prompt are both replaceable. The desktop consumes the same hooks and passes nothing: the proof the defaults are real defaults. (ADRs 0048, 0049)

The framework

Catamorphic's engine ships as libraries a host application mounts in-process. The desktop app is itself such a host. The host provides auth, the user/org model, the database, and the deployment surface. There is no default identity or tenant: every request carries identity from the host's auth context. See INTEGRATION.md for the host integration flow.

Runs anywhere

Durable agent-and-workflow infrastructure that does not assume an external server. Every dependency is an axis with a heavy and a light end. Pick per axis:

Axis Heavy end Light end
Database Network Postgres ({ pool } / { connectionString }) Embedded pglite (Kysely PGliteDialect via database: { db }; migrations run statement-by-statement so single-connection dialects just work)
Execution Cloud sandboxes: @catamorphic/cloudflare, @catamorphic/daytona Local sandboxes (@catamorphic/microsandbox), plain local processes (@catamorphic/local-process, trusted single-tenant hosts only), or none (read-only embed)
Code storage S3-compatible bucket (@catamorphic/s3) or Cloudflare Artifacts Two writable directories
Identity Host org/user per request One fixed tenant/user
Surface HTTP API + React UI In-process SDK calls, or migrations-only

The desktop app is the proof: it runs the lightest column end to end (pglite, local sandboxes, filesystem storage, no external server) by design (apps/desktop/src/main/server/boot.ts). No durable-execution vendor can run entirely inside a desktop app; this one does, and the same substrate is what offline-first agents need. Full matrix and host shapes: INTEGRATION.md.

Also worth knowing, because it's easy to miss from the package list:

  • The product teaches agents from the inside. Every project is seeded with hidden skills (.work/skills/): the project model and on-demand workspace scaffold (work-projects), workflow authoring (writing-workflows, batch-workflows, durable-workflows), and app building split into mechanics (building-apps) and replaceable design doctrine (designing-apps). Coding agents learn Catamorphic's authoring model at the moment they need it. The public skills/setup-work-server skill extends the same idea to stock-server setup and integrating Catamorphic into an existing app without replacing its auth or deployment.
  • One Run model. Boundary and batch workflows all share the same Runs API, hooks, and UI: capabilities, not categories. Every run executes a deployed commit.
  • Observability is free for hosts. Everything instruments against @opentelemetry/api; register your SDK and Catamorphic's spans appear in your traces.

The developer surface

Package What it is
@catamorphic/server-sdk The core SDK for your Node/Bun backend. Takes a Postgres connection (or pg.Pool), manages its own schema-scoped tables and migrations, and exposes projects, workflows, files, runs, triggers, agent sessions, connections, and code hosts (githubCodeHost over the github connection provider).
@catamorphic/fastify-plugin A mountable Fastify plugin (app.register(catamorphicPlugin, { core, prefix: "/api" })) exposing the standard HTTP API for frontends, plus the per-project MCP endpoints. Also exports a standalone createApp factory for sidecar deployments.
@catamorphic/react Headless React bindings: CatamorphicProvider, TanStack Query hooks, and jotai atoms. Build a fully custom UI on top of these.
@catamorphic/ui Ready-made components: the React Flow workflow canvas, member review and consent, the Runs panel, and AppMount (the sandboxed app iframe host). Every piece is opt-in.
@catamorphic/registry shadcn-style copy-paste components for hosts that want to own and customize the component source (project browser, git panel, runs panel, agent chat, Monaco editor).
@catamorphic/api-client Generated OpenAPI types + openapi-fetch client for the HTTP API.
@catamorphic/workflow Typed workflow-authoring primitives. Projects opt in directly, or a SaaS can wrap it and re-export only its approved surface.
@catamorphic/app The guest-side app runtime bundled into every user-built app: typed workflow client, persistent app-local storage shim, dual-dialect MCP Apps support, and the @catamorphic/app/ui component kit styled entirely by host theme tokens.

Supporting packages (consumed through the surface above, importable directly for advanced wiring):

Package What it is
@catamorphic/core Framework-agnostic service layer: projects, workflows, runs, deployments, triggers, apps, app storage, plugins, secrets, agent sessions, agent definitions, remote sync, and the CodeHost seam. The kernel behind server-sdk and fastify-plugin.
@catamorphic/db Kysely + Postgres. Schema-scoped (default schema catamorphic), raw SQL migrations, programmatic migrateToLatest.
@catamorphic/git Git-backed project storage (isomorphic-git): members' drafts as refs in the project origin (ADR 0191), local folder checkouts, pluggable origin remotes (RemoteBackend), and the remote sync engine (syncWithNetworkRemote: fetch and fast-forward; merge, push, and rescue branches only for repositories Work created), with the one push guard every network push passes (work/ branches only on attached repositories).
@catamorphic/github GitHub mechanics: OAuth + device-flow helpers, GitHub App auth (app JWTs, installation tokens, manifest registration), and the REST API client. The server SDK builds the github connection provider and code host on it.
@catamorphic/parser ts-morph AST → WorkflowGraph parser + dagre layout; also powers the seeded project check script.
@catamorphic/sandbox Vendor-neutral sandbox contracts (SandboxProvider, SandboxManager, RunExecutor), the stdio supervisor transport, OTel instrumentation, and plugin-doc staging and tool-policy helpers for harnesses.
@catamorphic/agent-protocol The agent session log (turns, attempts, items, requests, provider threads), its events, commands and shared reducer, and the runner protocol every harness adapter speaks (HarnessAdapter).
@catamorphic/agent-runner The harness-agnostic runner that drives one attempt of a turn on an adapter, in-process or over stdio.
@catamorphic/runner-bundle The runner and its Claude Code and Codex adapters as one hash-addressed file a sandbox runs with Bun or Node.
@catamorphic/microsandbox Local sandbox provider over the microsandbox SDK: the desktop's default execution.
@catamorphic/local-process Sandboxless execution as plain subprocesses with an explicit env. Trusted single-tenant hosts only (ADR 0047).
@catamorphic/cloudflare Cloudflare backend plugin: CloudflareSandboxProvider (execution via Bridge Worker) + ArtifactsRemoteBackend (Cloudflare-native code storage when available).
@catamorphic/s3 S3-compatible git origin backend for Cloudflare R2, AWS S3, MinIO, and similar stores.
@catamorphic/daytona Daytona backend plugin: DaytonaSandboxProvider + experimental Daytona git storage.
@catamorphic/ai-sdk Built-in harness adapter: Vercel AI SDK tool loop on any API model, running in the host and driving the session's sandbox through its tools.
@catamorphic/claude-code Harness adapter backed by the Claude Code (Claude Agent SDK) CLI, with per-turn MCP servers, portable native state and full settings-source fidelity.
@catamorphic/codex Harness adapter backed by the pinned OpenAI Codex app-server protocol.
@catamorphic/mcp MCP client infrastructure: both protocol generations with auto-negotiation, elicitation, the official MCP registry search, and plugin-marketplace install.
@catamorphic/otel Tiny OpenTelemetry helpers (@opentelemetry/api only: the host owns the SDK/exporters).
@catamorphic/runtime Execution harness that runs inside the sandbox and reports step results.
@catamorphic/plugins Plugin manifest contract + resolvers for host-provided packages and secrets.
@catamorphic/cloudflare-sandbox-bridge Deployable Cloudflare Worker exposing Cloudflare Sandbox over HTTP.

The in-repo reference host is the Catamorphic desktop app (apps/desktop): an Electron app that embeds the server in-process (src/main/server/boot.ts).

Design principles

  • Workflows are regular code. User-defined workflows run like normal apps: full IO, real npm dependencies, no crippled JS runtime. Execution happens inside a sandbox (or a local process, where the host shape allows it) using Bun to run and bundle.
  • Code stays simple. Both AI agents and humans must be able to write, edit, and understand workflows, and the parser must render them intuitively for non-technical users. See the code format below.
  • Host-injectable everything. Database connections, storage backends, sandbox credentials, LLM credentials, telemetry, trigger kinds, seeds, and doctrine are all injected by the host: nothing is hard-coded.
  • Postgres is authoritative for runs, retries, pauses, batch-item state, queues, and scheduling via SKIP LOCKED. Cloudflare Sandbox is the default cloud execution provider; backends ship as vendor plugin packages so hosts install only what they use.
  • OpenTelemetry throughout. Libraries instrument against @opentelemetry/api only; the host registers the SDK and exporters and gets full traces for free.

Settled design decisions are recorded as ADRs in docs/decisions/.

Quick start (embedding)

import { CloudflareSandboxProvider } from "@catamorphic/cloudflare";
import {
  createCatamorphic,
  defineStaticEnvironments,
} from "@catamorphic/server-sdk";

const sandboxProvider = new CloudflareSandboxProvider({
  apiUrl: process.env.CLOUDFLARE_SANDBOX_API_URL!,
  apiKey: process.env.CLOUDFLARE_SANDBOX_API_KEY,
});
const environmentProvider = defineStaticEnvironments([
  {
    descriptor: {
      id: "local",
      label: "Managed execution",
      trust: "managed",
      isolation: "sandbox",
      workloads: ["agent", "workflow"],
      agentTopologies: ["controller"],
      capabilities: ["network.egress"],
      resources: {},
    },
    sandboxProvider,
  },
]);

// Boot once per process
const catamorphic = createCatamorphic({
  database: { connectionString: process.env.DATABASE_URL! }, // or { pool }
  storage: {
    projectsPath: process.env.CATAMORPHIC_PROJECTS_PATH!,
    remotesPath: process.env.CATAMORPHIC_REMOTES_PATH!,
  },
  // Backend plugins: @catamorphic/cloudflare, @catamorphic/daytona,
  // @catamorphic/microsandbox, or @catamorphic/local-process
  sandboxProvider,
  environmentProvider,
});
await catamorphic.migrate(); // idempotent, schema-scoped

// Worker startup is explicit and host-owned.
const executionWorker = catamorphic.startExecutionWorker({ concurrency: 4 });

// Per request: bind the host's org + user
const client = catamorphic
  .forTenant({ tenantId: orgId })
  .forUser({ externalUserId: userId });
const project = await client.projects.create({ name: "Onboarding" });
const run = await client.runs.triggerProduction({
  projectId: project.id,
  workflowName: "welcomeUser",
  input: { email: "ada@example.com" },
});

To expose the HTTP API to your frontend:

import { catamorphicPlugin } from "@catamorphic/fastify-plugin";

app.register(catamorphicPlugin, { core: catamorphic.core, prefix: "/api" });

And on the frontend, wrap your tree with CatamorphicProvider from @catamorphic/react and drop in WorkflowEditor from @catamorphic/ui (or build your own UI from the hooks). See INTEGRATION.md.

Workflow code format

There is one public Workflow model and one public Run model: every workflow is an exported defineWorkflow(({ defineBoundary, defineBatch }) => ({ steps: [...] })) value from @catamorphic/workflow, and every run executes a deployed commit.

Boundaries and steps

A boundary is one atomic retry scope: if its callback fails, all operations in that callback retry together. Orchestration code lives in boundary run bodies; IO and business operations live in "use step" functions: plain async functions with the exact "use step" directive, called from boundary bodies. All functions take one destructured object parameter and carry JSDoc display metadata.

import { type BoundaryContext, defineWorkflow } from "@catamorphic/workflow";

/**
 * @displayname Welcome New User
 * @description Onboard a new user
 */
export const welcomeUser = defineWorkflow(({ defineBoundary }) => ({
  steps: [
    defineBoundary({
      run: async ({
        input,
      }: BoundaryContext<{ email: string; name: string }>) => {
        const user = await createUser({ email: input.email, name: input.name });
        await sendWelcomeEmail({ to: user.email, name: user.name });

        if (user.plan === "premium") {
          await assignPremiumBenefits({ userId: user.id });
        }

        await sendFollowUpEmail({ to: user.email });

        return { status: "complete", userId: user.id };
      },
    }),
  ],
}));

Batch scopes

A batch scope is paged per-item processing with an optional sink. defineBatchStep is a physical coalescing primitive for compatible calls inside defineBatch.process; it does not define a Workflow or a separate logical step scope.

import { defineWorkflow } from "@catamorphic/workflow";

export const processAccount = defineWorkflow(
  ({ defineBoundary, defineBatch }) => ({
    steps: [
      defineBoundary<{ accountId: string }, { accountId: string }>({
        retry: { maxAttempts: 3 },
        run: async ({ input }) => prepareAccount({ accountId: input.accountId }),
      }),
      defineBatch({
        source: async ({ input }: { input: { accountId: string } }) => ({
          source: recordsSource,
          config: { accountId: input.accountId },
        }),
        process: async ({ item }: { item: AccountRecord }) =>
          processRecord({ record: item }),
        sink: resultSink,
      }),
    ],
  }),
);

The parser and UI expose capabilities such as batch processing and cancellation. There is no public stage concept, category switch, or separate Run family for these capabilities. A workflow with no pause, retry backoff, rate limit, batch, or child call settles inline when triggered synchronously: inline request-response is a property of execution, not a separate authoring form.

Triggers

Workflows subscribe to trigger kinds in code (triggers: [trigger("kind", config)]). Hosts define their own kinds with defineTriggerKind (typed payloads and configs via zod, generated work-triggers.d.ts per project) and fire them sync or async; a kind whose payload varies per workflow uses typed holes, and kinds declared via mcpToolKinds are served as MCP tools from POST /projects/:id/mcp. Projects define their own kinds on top of the host's in .work/triggers/ (defineTrigger({ name, from: trigger("webhook", { ... }), where })), and every binding may add a where filter the host checks before a run starts, so a GitHub or Slack integration is project code on the webhook kind, whose verification schemes and handshakes are declared in its config. The host skill slack is the worked example: a chat per Slack thread, replies posted back through a gateway connection whose named operations are its capabilities (Connect Slack). The host skill reviewing-pull-requests is the larger one: a code and security review of every pull request, in a chat per pull request on a dedicated review pool, verified by running the change and posted as a review and a check run (Review pull requests). See INTEGRATION.md and ADRs 0039 / 0042 / 0171 / 0179 / 0181.

Long-lived journeys: correlation keys, signals, shared rate budgets

A run may carry a correlation key: a host-meaningful identity for its subject (a contact, an account, a subscription). It is unique among live runs of the same workflow, which makes it at once an enrollment idempotency key and the address external events use to reach the run. A boundary declares the shared third-party budgets it draws on; buckets are keyed per tenant, so every workflow naming the same globalKey draws on one budget.

export const nurtureContact = defineWorkflow(({ defineBoundary }) => ({
  controls: { cancel: true },
  steps: [
    defineBoundary({
      rateLimits: [
        { globalKey: "email", capacity: 500, refillRatePerSecond: 100 },
      ],
      run: async ({ input, pause }: BoundaryContext<{ contactId: string }>) => {
        await sendWelcomeEmail({ contactId: input.contactId });
        return pause<{ clicked: boolean }, { contactId: string }>({
          signal: "reply",
          timeout: "72h",
          state: { contactId: input.contactId },
        });
      },
    }),
    defineBoundary({
      // A different step, a different provider, a different budget.
      rateLimits: [
        { globalKey: "whatsapp", partitionKey: "sender-1", capacity: 80, refillRatePerSecond: 20 },
      ],
      run: async ({ input }) => sendWhatsApp({ contactId: input.state.contactId }),
    }),
  ],
}));
// Enrolling twice for the same contact is a no-op, so webhook redelivery is safe.
await client.runs.triggerProduction({
  projectId, workflowName: "nurtureContact",
  input: { contactId: "contact-42" },
  correlationKey: "contact-42",           // onConflict: "ignore" | "error" | "restart"
});

// External events address the contact, not a run id.
await client.runs.signalByKey({
  projectId, workflowName: "nurtureContact",
  correlationKey: "contact-42",
  signal: "reply", idempotencyKey: replyEventId, value: { clicked: true },
});

// Opting out ends the journey wherever it sits. Resolves null if none was live.
await client.runs.cancelByKey({
  projectId, workflowName: "nurtureContact",
  correlationKey: "contact-42", reason: "unsubscribed",
});

Waiting for capacity is not failure: a boundary that cannot reserve is rescheduled without consuming a retry and without holding a sandbox. When a provider answers with a 429, report it with rateLimited({ retryAfterMs }): that blocks every workflow sharing the account, not just the run that hit it.

Hosts bound what any one tenant may consume from these shared resources via catamorphic.tenantPolicies. It is SDK-only and never exposed over HTTP, so a tenant cannot raise its own limits:

await catamorphic.tenantPolicies.upsert({
  tenantId,
  maxConcurrentJobs: 32,     // ceiling on simultaneously leased jobs
  maxActiveRuns: 50_000,     // ceiling on live runs; bounds enrollment fan-out
  queueWeight: 4,            // relative share of each claim batch
  retentionDays: 365,        // overrides the installation retention window
  rateLimitOverrides: {      // can tighten an author's bucket, never loosen it
    whatsapp: { capacity: 40, refillRatePerSecond: 10 },
  },
});

Retention

Finished runs are purged after 90 days by default, along with everything hanging off them: jobs, events, step attempts, batch items and their steps. Without this, a daily 100k-item batch adds on the order of 1.5M rows a day and never gives any back.

The sweep runs inside the execution worker, so it needs no wiring. Change the window, or turn it off entirely, at construction:

const catamorphic = createCatamorphic({
  database,
  storage,
  environmentProvider,
  retention: { runRetentionDays: 30 }, // or { enabled: false } to keep forever
});

Individual tenants can be given a longer or shorter window with retentionDays above.

See ADR 0027, ADR 0028, and ADR 0030.

Contributing

Catamorphic is issue-first. If you do not have repository write access, open a GitHub issue with the problem, concrete use case, constraints, and desired outcome. Public pull request creation is restricted to the core team and collaborators with write access.

AI makes producing code cheap, but it does not make a large generated change cheap to review. We review the compact source of intent first, then maintainers own the implementation with the repository's full context. Read CONTRIBUTING.md for the reasoning and issue guidance.

Local development

bun install

# Only for an explicitly provisioned host database. These are not part of the
# normal local app workflow; never point them at a database another session uses.
DATABASE_URL="<host-owned database URL>" bun run db:migrate
DATABASE_URL="<host-owned database URL>" bun run db:codegen

# Build everything
bun run build

# Regenerate the OpenAPI spec + typed api-client after route/DTO changes
cd packages/fastify-plugin && bun run generate-spec
cd ../api-client && bun run generate

Optional shared OTel, ClickHouse, and sandbox-bridge services run persistently. Start them in a separate terminal only when you need them:

bun run dev:infra

The framework packages are embed-only. In production, run either the stock server host or an application that boots Catamorphic in-process. For local development, bun run dev starts the combined desktop and stock-server manual environment. bun run dev:desktop and bun run dev:server are focused variants of the same orchestrator, which assigns each worktree its own data directories and loopback ports. To iterate on Catamorphic alongside your own host, link the packages via file: (see .agents/skills/using-catamorphic/SKILL.md and its "Local dev linking" section).

Scripts

bun install         # Install dependencies in this worktree
bun run dev         # Combined desktop and stock-server manual environment
bun run dev:desktop # Desktop-focused development environment
bun run dev:server  # Stock-server-focused development environment
bun run dev:infra   # Optional OTel, ClickHouse, and sandbox-bridge services
bun run build       # Build all packages
bun run test        # Deterministic root and Postgres-complete workspace tests
bun run test:external # Explicit opt-in for credentialed external integrations
bun run check       # 11-phase merge gate
bun run typecheck   # Typecheck all packages with tsgo
bun run lint        # Lint with Biome
bun run lint:fix    # Auto-fix lint issues
bun run db:migrate  # Apply migrations to the DB pointed at by DATABASE_URL
bun run db:codegen  # Regenerate Kysely types from the `catamorphic` schema
bun run db:reset    # Drop + recreate the catamorphic schema (dev only)
bun run db:status   # Show applied / pending migrations

Manual development profiles live under ~/.catamorphic/dev/<worktree-instance>. The runner assigns separate desktop/server profiles and ports per worktree. Legacy temporary profiles are copied on first launch, with the originals retained. Do not put manual project databases in the operating system temporary directory.

Testing

Tests use Vitest, orchestrated by Turborepo. Docker must be running: bun run test and bun run check create an isolated disposable Postgres database for every invocation. They do not need bun run dev:infra; that command is optional shared observability and sandbox-bridge infrastructure, not test infrastructure.

bun run test                                   # deterministic root and Postgres-complete workspace tests
bun run check                                  # 11 phases: lint, root/workspace types, build, migrations, root/workspace tests, and E2E
bun run test:external                          # explicit authority for credentialed integrations
bun run --filter @catamorphic/parser test      # one package through repository Node
bun run --cwd packages/parser test src/__tests__/parser.test.ts  # one file through repository Node

Each test file uses fresh temporary state and every root test invocation gets its own Postgres database and temporary caches, so parallel worktrees do not share test output. Install dependencies in each worktree and never copy a credentialed .env wholesale. Add only the individual settings a task needs. Root orchestration tests under scripts/**/*.test.ts run before the package graph. Package-local Vitest commands also select the repository-pinned Node, including when the ambient node is Node 20.

External integration tests never run merely because credentials are present. They require the explicit bun run test:external authority, which sets CATAMORPHIC_EXTERNAL_INTEGRATIONS=1, and skip when their service-specific configuration is absent:

  • Daytona (packages/daytona): needs DAYTONA_API_KEY.
  • Cloudflare Sandbox (packages/cloudflare, packages/core): needs CLOUDFLARE_SANDBOX_API_URL, CLOUDFLARE_SANDBOX_API_KEY, and CF_SANDBOX_INTEGRATION=1. Start the sandbox bridge separately if the test needs it.
  • Cloudflare Artifacts (packages/cloudflare): needs CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_API_TOKEN, and CLOUDFLARE_ARTIFACTS_NAMESPACE; it skips with a warning while feature-gated.
  • S3-compatible storage (packages/s3): needs S3_BUCKET, S3_ACCESS_KEY_ID, and S3_SECRET_ACCESS_KEY. Set S3_REGION when the provider requires it, S3_ENDPOINT for R2, MinIO, or another non-AWS endpoint, and S3_FORCE_PATH_STYLE=true when that endpoint requires path-style requests.

Unit tests run with no setup. bun run --cwd apps/desktop test:e2e runs the complete desktop suite in Docker, using real Electron on a private Xvfb display with Openbox and a deterministic fake agent. Windows use normal focus and rendering without reaching your desktop. The first run builds a cached Linux image; subsequent runs reuse dependencies. File filters and Vitest shard flags pass through. Logs, screenshots, and JUnit results go to test-results/. bun run check includes this isolated suite. Native macOS coverage runs in CI. See desktop testing for runner and sharding details.

PR checks reuse successful package and shard tasks by content hash. Changing a package invalidates its downstream consumers; local commands and main still execute tests fresh. See test caching and selective execution for cache inputs, failure handling, and execution summaries.

The local development runner uses POSIX process groups and supports macOS and Linux. Windows development orchestration is not supported.

Tech stack

  • Bun: runtime, package manager, bundler (also inside sandboxes)
  • TypeScript: tsgo for typechecking
  • Postgres: all state, schema-scoped; run queues, retries, pauses, and scheduling use the same DB
  • Electron + electron-vite: the desktop reference implementation
  • Fastify: HTTP surface with Zod + OpenAPI (mounted by the host)
  • React Flow: workflow visualization; Jotai + TanStack Query: frontend state
  • ts-morph: TypeScript AST parsing
  • Kysely: type-safe SQL
  • isomorphic-git: portable git, no native CLI dependency (system git is a soft dependency for the desktop's read surfaces)
  • OpenTelemetry: @opentelemetry/api instrumentation throughout
  • Turborepo: build orchestration

Roadmap

Direction, not shipped. Tracked in TODO.md:

  • More desktop platforms: Intel Mac validation, then signed Windows and Linux packages.
  • ACP harness: project agent definitions already accept kind: "acp"; the Agent Client Protocol client (local command and remote endpoint transports) is the planned harness behind it.
  • TS defineAgent: a typed authoring layer that compiles to the committed .work/agents/<slug>.json substrate.
  • Review-mode collaboration: PR-first sync for shared projects, invites, native PR review depth (comments, approvals, merge).
  • Remote blob storage: user-connected stores (S3/R2/Drive-style) for large binaries beside git-tracked text.

About

Work, the open-source desktop app for getting work done with AI agents (work.software), and Catamorphic, the framework underneath for embedding agents, workflows and apps in your own products

Topics

Resources

Contributing

Stars

3 stars

Watchers

0 watching

Forks

Used by

Contributors

Languages