diff --git a/docs-draft-v2/README.md b/docs-draft-v2/README.md new file mode 100644 index 0000000..47e33db --- /dev/null +++ b/docs-draft-v2/README.md @@ -0,0 +1,60 @@ +# OpsCoach + +Learn Linux and AWS operations on a real, throwaway cloud server instead of a quiz. You get a dedicated Linux host and a task list, and a grader inspects the actual state of the machine as you work. + +This is a portfolio project. The parts worth a look: a live browser-to-SSH terminal, per-session disposable lab hosts on AWS, and grading that reads real system state rather than comparing answers. + +## What it does + +Open a lab and you get a dedicated Linux host in a minute or two. You work in a terminal, in the browser or over your own SSH client, to fix broken services, harden config, triage logs, wire up systemd, and so on. A grader runs against the live machine and streams pass/fail checks to a dashboard beside the terminal. When you finish or go idle, the host is destroyed. + +Three content packs ship today: **Linux Foundations** and **AWS Foundations** drills, and the **Beaconkeeper** capstone, a 20-step systemd and operations scenario on a single box. + +## How it works + +A request reaches a shared ALB, authenticates through Cognito, and lands on the OpsCoach web app on ECS Fargate. Starting a lab launches a dedicated EC2 host; the web app bridges the browser terminal to it over SSH and runs the grader over SSH. Every host self-destructs on an idle or lifetime timer. [architecture.md](architecture.md) has the diagram and the full provision, terminal, grade, and teardown flows. + +## Why it is built this way + +Four decisions shaped the system. Each is defended in full in [architecture.md](architecture.md); the short version: + +- **Real hosts, not a fake shell.** Learning operations means touching real systemd, real packages, real logs, so each session is an actual EC2 instance rather than an emulation. +- **Disposable and single-tenant.** The lab runs as a privileged user, so each session gets its own host that is wiped on a timer. The security boundary is the cheap, isolated host, not container isolation. See [security.md](security.md). +- **Grade real state, not answers.** The grader SSHes in and inspects the machine, so a check passes only when the box is genuinely in the right state. +- **Borrow the platform, do not rebuild it.** The infrastructure plugs into an existing shared ALB, Cognito, and VPC by ID rather than standing up its own. The real resource IDs live in local config kept out of the repo. + +## Repository layout + +| Path | What it is | +| --- | --- | +| `web/` | Next.js app and custom Node server (WebSocket-to-SSH bridge), API routes, UI | +| `infra/` | AWS CDK app: Fargate service, lab hosts, teardown automation | +| `ContentPacks/` | Labs and graders, including the Beaconkeeper game | +| `scripts/` | Build, deploy, and smoke-test scripts | +| `docs/` | Architecture, security, and design notes | + +## Run it locally + +```bash +cd web +npm install +cp .env.example .env # most values are optional locally +npm run dev +``` + +With no AWS credentials and no `DATABASE_URL`, the app runs in mock mode: an in-memory store and faked provisioning, so you can click through the UI and content offline. Run the tests with `npm test`. Mock mode does not start a lab container for you, so a few flows still need Docker running locally; the gaps and the proposed local-Docker mode are in [local-dev-without-aws.md](local-dev-without-aws.md). + +## Configure and deploy + +Two things stay out of version control: the app's `web/.env` (copied from `web/.env.example`) and the CDK platform context (`infra/cdk.context.json`, copied from the example or generated by `scripts/discover-platform-context.sh`). With those in place, `scripts/deploy-platform.sh` builds the images and deploys the stacks. The walkthrough is in [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md). + +## Docs + +- **[architecture.md](architecture.md)**: the system walkthrough, from the three moving parts down to the design decisions worth defending. Start here. +- **[security.md](security.md)**: how untrusted users get root without putting anything else at risk. +- **[lab-lifecycle-design.md](lab-lifecycle-design.md)**: provisioning and the three-layer teardown, next to the code. +- **[local-dev-without-aws.md](local-dev-without-aws.md)**: what runs without AWS today, and what a full local-Docker mode would take. + +## License + +MIT. See [`../LICENSE`](../LICENSE). diff --git a/docs-draft-v2/architecture.md b/docs-draft-v2/architecture.md new file mode 100644 index 0000000..43a4955 --- /dev/null +++ b/docs-draft-v2/architecture.md @@ -0,0 +1,164 @@ +# OpsCoach architecture + +OpsCoach gives each learner a real, throwaway Linux host and grades what they actually do to it. The system is three parts: a **web app** on ECS Fargate (UI, terminal bridge, grading), a **per-session EC2 lab host** the learner operates and that dies on a timer, and a **shared platform** (ALB, Cognito, VPC) the app plugs into rather than rebuilding. Everything below is how those three talk to each other safely. + +**Scope.** Single-region training infrastructure for one app on a borrowed ALB/Cognito platform. Non-goals: this is not a hostile multi-tenant sandbox (the isolation model assumes a cooperative-but-curious learner, see [security.md](security.md)), and not a high-availability production service. + +This doc starts at the diagram and components, then defends the design decisions. The step-by-step request flows are in a collapsible block so you can skip them unless you need the wire-level detail. + +## The system + +![OpsCoach AWS architecture](architecture.svg) + +The numbered arrows above: + +1. **Request.** Browser to the ALB over HTTPS. The ALB runs `authenticate-cognito` (Cognito hosted UI, Google); a post-auth passphrase gate in the app must also pass before any page renders. Only the health check and logout skip Cognito. +2. **Route.** The ALB forwards to the OpsCoach web service on Fargate, in a private, egress-only subnet. +3. **Provision, operate, grade.** Fargate launches and drives the per-session EC2 host: `RunInstances` at the start, an SSH PTY for the browser terminal, and the grader over SSH. +4. **Ready webhook.** The host calls back to the service once it is up, resolved through Cloud Map and authenticated with a shared secret. +5. **Terminal.** The browser streams over a WebSocket to the custom Node server, which relays to the host's shell over SSH. +6. **Teardown.** A per-session EventBridge Scheduler one-shot triggers the terminator Lambda; a 5-minute sweep is the backstop. + +Grey dashed lines are supporting paths: Fargate to RDS PostgreSQL in the isolated subnet, the host pulling its lab image from ECR, and the opt-in direct SSH from a learner's laptop. + +Colour key: purple is networking (ALB, Cloud Map); orange is compute and containers (Fargate, EC2, Lambda, ECR); red is identity and secrets (Cognito, Secrets Manager); blue is the database (RDS); pink is app integration (EventBridge Scheduler). + +## Components + +| Component | AWS service | Role | +| --- | --- | --- | +| Web app + terminal bridge | ECS Fargate | Next.js app + custom Node server; WebSocket-to-SSH PTY bridge; session lifecycle, grading, dashboard | +| Edge auth | ALB + Cognito | `authenticate-cognito` at the load balancer (hosted UI to Google), then a post-auth passphrase gate | +| Lab host | EC2 (per session) | Ephemeral AL2023 / arm64 host running the lab container; learner SSH target | +| Container images | ECR | Images for the web service and each lab | +| Database | RDS PostgreSQL | Sessions, check runs, grader results (isolated subnet) | +| Service discovery | Cloud Map | In-VPC address for lab-host-to-web callbacks | +| Teardown | EventBridge Scheduler + Lambda | One-shot per-session schedule fires a terminator Lambda; 5-minute sweep backstop | +| Secrets | Secrets Manager | Database credentials and the callback HMAC secret | + +## The design decisions + +Each of these was a fork where the cheaper, obvious choice was the wrong one. The pattern is constraint, decision, trade-off. + +**A real host per session, not a shared sandbox.** +Teaching operations means real systemd, real root, and real packages, but you cannot hand untrusted users root on shared infrastructure. The cheaper alternatives, a shared multi-tenant host or browser-only containers, make container isolation the only wall between a hostile learner and everyone else, so one container escape compromises every session. Instead, every session gets its own ephemeral EC2 host, and the security model treats that host (not the container on it) as the real boundary, killed on a timer. The cost is a minute or two of provisioning latency and per-session compute, bought back as realism and a blast radius of exactly one throwaway box. Full model in [security.md](security.md). + +**A custom Node server for the browser terminal.** +A browser terminal needs a long-lived, two-way connection to a shell, and Next.js on its own does not hold one. The alternative, a managed real-time service or a separate WebSocket process, adds a moving part for what is a thin bridge. Instead, a small `server.js` wraps Next, upgrades the WebSocket, authenticates the session through an internal call, and bridges to the host with an `ssh2` PTY. The cost is running a custom server instead of stock Next, plus a 25-second keepalive so the ALB does not cut an idle terminal, in exchange for a real shell in the browser with nothing for the learner to install. + +**Grade real state, not answers.** +Multiple-choice cannot tell you whether someone can actually run a box. Instead, the grader SSHes into the live host with a least-privilege environment and checks real state (services up, files in place, config correct), returning structured results. The cost is that graders are per-pack code that has to run against a live machine, in exchange for a pass that means the box is genuinely in the right state. + +**Three independent teardown paths.** +A leaked instance costs real money, and the control plane never observes what the learner does inside their SSH session (direct-SSH learners bypass it entirely). A single cleanup mechanism is a single point of failure for the one control that bounds spend. Instead there are three: an on-host SSH-idle watcher, a one-shot EventBridge Scheduler set at provision time, and a 5-minute sweep over expiry tags. Any one is enough, and all are idempotent. The cost is more moving parts, for a hard guarantee that nothing runs forever. Full design in [lab-lifecycle-design.md](lab-lifecycle-design.md). + +**Borrow the platform; import by ID.** +A demo should not stand up its own ALB, Cognito, and VPC. Instead the CDK imports a shared platform's resources by ID from local context, and the real IDs stay out of the repo. The cost is that the app cannot bootstrap its own world from nothing, in exchange for dropping cleanly into a real shared environment. + +## The flows, step by step + +
+Auth, provision, terminal, grading, teardown (expand for the wire-level sequences) + +  + +**1 · Authentication and access gate** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant ALB as ALB + participant Cog as Cognito + participant App as Fargate + U->>ALB: HTTPS request + ALB->>Cog: authenticate-cognito (if enabled) + Cog-->>ALB: OIDC tokens + ALB->>App: forward + signed identity header + alt no gate cookie + App-->>U: redirect to /gate + U->>App: POST /api/gate (passphrase) + App-->>U: set gate cookie, continue + end + App-->>U: app page + Note over ALB,App: /api/health and /logout skip Cognito +``` + +**2 · Provision a lab session** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant App as Fargate + participant EC2 as Lab host + participant ECR as ECR + participant Sch as Scheduler + U->>App: POST /api/sessions (start lab) + App->>App: create session in RDS; mint keys + callback token + App->>EC2: RunInstances (Launch Template + user-data) + App->>Sch: CreateSchedule (terminate at T + maxLifetime) + Note over EC2: user-data: install Docker, block IMDS,
ECR login, run lab container, set authorized_keys + EC2->>ECR: pull lab image + EC2->>App: ready webhook (via Cloud Map + shared secret) + App-->>U: session ready (host, port) +``` + +**3 · Browser terminal and SSH** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant Srv as Node server + participant EC2 as Lab host + U->>Srv: WebSocket /api/sessions/:id/shell?token + Srv->>Srv: shell-auth (internal secret) resolves host + key + Srv->>EC2: ssh2 PTY on port 22 (per-session key) + EC2-->>Srv: stdout / stderr + Srv-->>U: stream to xterm.js + Note over Srv,EC2: 25s keepalive ping holds the ALB idle timer + Note over U,EC2: opt-in: SSH directly to the host's public IP +``` + +**4 · Live grading** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant App as Fargate + participant G as Grader + participant EC2 as Lab host + U->>App: POST /api/sessions/:id/grade + App->>G: spawn grader (allowlisted env, no task-role creds) + G->>EC2: SSH port 22 (grader key) runs ops status + EC2-->>G: JSON check results + G-->>App: { passed, checks[] } + App->>App: persist check run in RDS + App-->>U: live results on the dashboard +``` + +**5 · Idle and lifetime teardown** + +```mermaid +sequenceDiagram + autonumber + participant Sch as Scheduler + participant L as Terminator Lambda + participant EC2 as Lab host + participant App as Fargate + Sch->>L: fire at T + maxLifetime + L->>L: read callback secret (Secrets Manager) + L->>EC2: TerminateInstances + L->>App: shutdown callback (mark session stopped) + Note over EC2,App: backstop: 5-min EventBridge sweep
terminates any host past its ExpiresAt tag +``` + +
+ +## See also + +- **[security.md](security.md)** for the security model and its trade-offs. +- **[lab-lifecycle-design.md](lab-lifecycle-design.md)** for provisioning and the three-layer teardown in depth. +- **[`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md)** for plugging into a shared ALB/Cognito platform. diff --git a/docs-draft-v2/architecture.svg b/docs-draft-v2/architecture.svg new file mode 100644 index 0000000..4ec1f3d --- /dev/null +++ b/docs-draft-v2/architecture.svg @@ -0,0 +1,149 @@ + + OpsCoach AWS architecture + A user reaches a shared Application Load Balancer with Cognito auth and a passphrase gate, which forwards to the OpsCoach web service on ECS Fargate in a private subnet. Fargate uses an isolated-subnet RDS PostgreSQL database, launches per-session EC2 lab hosts in a public subnet (which pull a lab container from ECR and call back via Cloud Map), bridges the browser terminal to the host over SSH, runs a grader over SSH, and schedules automatic teardown via EventBridge Scheduler and a terminator Lambda. + + + + + + + + + + + + + OpsCoach — AWS architecture + In-browser & SSH Linux labs · shared ALB/Cognito platform · per-session EC2 hosts · live grading · auto-teardown + + + USER + + Browser + xterm.js terminal + + SSH client + direct opt-in + + + + + AWS Account · Region (us-east-1) + + + + Shared platform — imported by the app (not created here) + + Application Load Balancer + Cognito auth · routing + + Amazon Cognito + hosted UI · Google IdP + + + + VPC + + + + Public subnet + + EC2 — lab host (per session) + AL2023 · arm64 · Launch Template + runs lab container · IMDS blocked + + + + Private subnet (egress) + + ECS Fargate — OpsCoach web + Next.js + custom Node server + WebSocket → SSH PTY bridge + + AWS Cloud Map + in-VPC discovery + + + + Isolated subnet + + Amazon RDS — PostgreSQL + sessions · check runs · results + + + + Automation & services + + Amazon ECR + web + lab images + + Secrets Manager + DB · callback · keys + + EventBridge + Scheduler + one-shot per session + + AWS Lambda + session terminator + + + + + 1 + + + + 2 + + + + 3 + + provision · SSH PTY · grade + + + + 4 + + + ready webhook + + + + 5 + + WSS + + + + SSH opt-in + + + + + 6 + + terminate EC2 + + + + SQL + + + + pull lab image + + + diff --git a/docs-draft-v2/lab-lifecycle-design.md b/docs-draft-v2/lab-lifecycle-design.md new file mode 100644 index 0000000..102bccd --- /dev/null +++ b/docs-draft-v2/lab-lifecycle-design.md @@ -0,0 +1,216 @@ +# Lab instance lifecycle design + +**Status: implemented (web + CDK).** + +Each learner session runs on a dedicated EC2 host that must be torn down reliably, because a leaked instance costs real money and there is no human watching. Teardown uses **three independent paths**, any one of which is sufficient and all of which are safe to run twice. This doc records how hosts are provisioned and destroyed, and why a single cleanup mechanism was not enough. + +## Problem + +Starting a session launches a dedicated EC2 instance (`t4g.micro`) running Docker with the lab container. If teardown fails or never runs, instances leak and accumulate cost. + +The control plane (Next.js on Fargate) cannot reliably observe when a learner is done. It bridges the browser terminal over a WebSocket, but a learner can also SSH straight to the host's public IP and bypass the control plane entirely, and even within the bridged terminal the control plane does not see SSH connect or disconnect events. So idle has to be detected on the host, not inferred from the control plane. + +Teardown therefore needs to be: + +- **Prompt** when the learner is done (SSH idle). +- **Reliable** when callbacks fail, schedules are missed, or the learner never connects. +- **Idempotent** when multiple paths fire close together. + +## Architecture context + +```mermaid +sequenceDiagram + participant Learner + participant Web as OpsCoachWeb_Fargate + participant EC2 as LabHost_EC2 + participant Lambda as SessionTerminator + participant Scheduler as EventBridgeScheduler + + Learner->>Web: POST /api/sessions (pubkey) + Web->>EC2: RunInstances + user-data + Web->>Scheduler: CreateSchedule T+maxLifetime + EC2->>Web: POST /ready (public + private IP) + Learner->>EC2: SSH :22 + Note over EC2: idle watcher monitors :22 + Learner--xEC2: SSH disconnect + EC2->>Web: POST /shutdown reason=ssh_idle + Web->>EC2: TerminateInstances + Web->>Scheduler: DeleteSchedule + + Note over Scheduler,Lambda: If still running at T+maxLifetime + Scheduler->>Lambda: invoke terminate + Lambda->>EC2: TerminateInstances + Lambda->>Web: POST /shutdown reason=max_ttl +``` + +**Key constraints:** + +- SSH is pubkey-only on a hardened host (v1: public IP; no Tailscale). +- The grader runs from Fargate inside the VPC and SSHes to the instance's **private IP**. +- Session state lives in Postgres; EC2 lifecycle is driven by the web task role and the terminator Lambda. + +## Why three paths + +A single teardown mechanism is a single point of failure for the one control that bounds spend, so the design layers three independent paths. Any one succeeding is sufficient, and all are safe to run more than once. + +| Layer | Trigger | Actor | Typical latency | +|-------|---------|-------|-----------------| +| 1. SSH idle watcher | No established TCP sessions on host `:22` for a grace period, after at least one session was seen | EC2 user-data background script | ~2 min after disconnect | +| 2. Max TTL schedule | One-time EventBridge Scheduler at provision time | `OpsCoachSessionTerminator` Lambda | Exactly at T + max lifetime | +| 3. ExpiresAt sweep | `ExpiresAt` EC2 tag in the past | Same Lambda, every 5 min | Up to 5 min after tag expiry | + +Manual **Stop lab** (the learner button) and the authenticated `POST /api/sessions/:id/stop` use the same internal shutdown path as the webhooks. + +The layers cover each other's failure modes: + +- **Idle watcher alone is not enough.** The control plane does not terminate SSH, so without a host-side agent it would learn that a session is idle only when the learner clicks Stop or a coarse timer fires. The watcher closes the common case: the learner closes their terminal and walks away. +- **A timer alone is not enough.** Fixed timers are either too aggressive (they kill active sessions) or too loose (they leak nodes). A max lifetime is still necessary as a backstop for failed shutdown webhooks, learners who never SSH (the instance still costs money), and bugs in the idle watcher. +- **Scheduler and sweep are both kept** because they fail differently. The Scheduler fires once per session at a precise time and deletes itself, which is the primary hard cap; the `ExpiresAt` tag plus sweep catches instances where schedule creation itself failed (missing IAM, an API error during provision) or where AWS Scheduler drifted. + +## Layer 1: SSH idle watcher + +**Location:** generated shell user-data in [`../web/lib/lab-user-data.ts`](../web/lib/lab-user-data.ts) (also mirrored in [`../infra/lib/lab-user-data.sh`](../infra/lib/lab-user-data.sh) for launch-template defaults). + +**Behavior:** + +1. After bootstrap, a background subshell loops every 15 seconds. +2. Count established connections on local port 22 via `ss -tn state established '( sport = :22 )'`. +3. Track `had_session`: set to 1 once the count is greater than 0 at least once. +4. When `had_session` is 1 and the count is 0, start an idle clock. +5. If idle for `SSH_IDLE_GRACE_SECONDS` (default **120**), POST the shutdown webhook. + +**Webhook:** + +```http +POST /api/sessions/:id/shutdown +X-Internal-Secret: +Content-Type: application/json + +{ "reason": "ssh_idle" } +``` + +The 120-second grace avoids tearing down during a brief disconnect (a network blip or an `ssh` reconnect) and lets grader SSH from Fargate finish without racing the learner's disconnect in edge cases. + +**Two known limitations, both acceptable for v1:** + +- The watcher counts **all** connections on host `:22`, including grader SSH from the VPC. A learner who never opens a terminal but repeatedly runs checks from the web UI may still see the instance terminated shortly after grading quiesces. A follow-up could filter by source IP (count only non-RFC1918 addresses). +- If the learner never SSHes, `had_session` stays 0 and the watcher never fires. Max TTL (layer 2) handles that case. + +## Layer 2: One-time EventBridge schedule (max TTL) + +**Location:** [`../web/lib/session-scheduler.ts`](../web/lib/session-scheduler.ts), invoked from [`../web/lib/ec2-labs.ts`](../web/lib/ec2-labs.ts) after `RunInstances`. + +**Behavior:** + +1. On successful provision, Fargate creates schedule `opscoach-{sessionId}` (truncated to 64 chars). +2. Expression: `at(yyyy-mm-ddThh:mm:ss)` in UTC, **T + maxLifetimeMinutes** from provision time. +3. Target: the `OpsCoachSessionTerminator` Lambda with payload: + + ```json + { "action": "terminate", "instanceId": "i-…", "sessionId": "…", "reason": "max_ttl" } + ``` + +4. `ActionAfterCompletion: DELETE` removes the schedule after it fires. +5. Any explicit shutdown (`manual`, `ssh_idle`) calls `DeleteSchedule` for idempotency. + +The default max lifetime is **60 minutes** (`OPSCOACH_MAX_LIFETIME_MINUTES` / CDK context `maxLifetimeMinutes`): long enough for a typical lab session and assessment retries, short enough to bound cost if every other teardown path fails, and orthogonal to the SSH idle grace. + +EventBridge Scheduler is used rather than EventBridge Rules because it supports **one-time** schedules natively, with per-session names and auto-delete. Rules suit recurring patterns, which is why the 5-minute sweep (layer 3) uses one. + +## Layer 3: ExpiresAt tag sweep + +**Location:** [`../infra/lib/session-terminator/handler.py`](../infra/lib/session-terminator/handler.py), triggered every 5 minutes by an EventBridge Rule in [`../infra/lib/lab-host-stack.ts`](../infra/lib/lab-host-stack.ts). + +**Behavior:** + +1. At provision, `RunInstances` tags the instance with `ExpiresAt=` aligned to the **max lifetime** (the same horizon as the scheduler). +2. The Lambda scans running OpsCoach instances (`OpsCoach=true`). +3. If `ExpiresAt <= now`, it terminates the instance and calls the shutdown API with `reason=expires_at_sweep`. + +This is the cheap safety net for when schedule creation failed or an instance outlived its schedule because of API errors. + +## Unified shutdown path + +All automated and manual teardown converges on [`shutdownSessionInternal`](../web/lib/sessions.ts): + +1. Idempotent if already `stopped` or `stopping`. +2. Set status `stopping`. +3. `DeleteSchedule` (best effort). +4. `TerminateInstances` if an instance id is present (errors are logged, the session is still marked stopped). +5. Set status `stopped`, publish an SSE event. + +**Entry points:** + +| Entry | Auth | Reason | +|-------|------|--------| +| `POST /api/sessions/:id/stop` | Session token | `manual` | +| `POST /api/sessions/:id/shutdown` | `X-Internal-Secret` | `ssh_idle`, `max_ttl`, `expires_at_sweep`, `manual` | + +The terminator Lambda terminates EC2 first, then calls the shutdown API, so Postgres stays in sync. + +## Security + +- Shutdown and ready callbacks require `X-Internal-Secret` (stored in Secrets Manager, created in the lab-host stack, read by the Fargate task role). +- EC2 terminate IAM is scoped with `OpsCoach=true` resource and request tags where possible. +- The host watcher can only initiate shutdown; it cannot terminate instances directly, because the lab instance role has no terminate permission. + +## Configuration + +### Runtime (Fargate task environment) + +| Variable | Default | Purpose | +|----------|---------|---------| +| `OPSCOACH_MAX_LIFETIME_MINUTES` | `60` | Scheduler fire time and `ExpiresAt` tag | +| `OPSCOACH_SSH_IDLE_GRACE_SECONDS` | `120` | Host idle debounce before the shutdown webhook | +| `SESSION_TERMINATOR_LAMBDA_ARN` | (CDK) | Schedule target | +| `SCHEDULER_INVOKE_ROLE_ARN` | (CDK) | Scheduler execution role | +| `INTERNAL_CALLBACK_SECRET` | Secrets Manager | Authenticates host and Lambda callbacks | + +### CDK context ([`../infra/lib/web-config.ts`](../infra/lib/web-config.ts)) + +| Key | Default | Purpose | +|-----|---------|---------| +| `maxLifetimeMinutes` | `60` | Hard cap | +| `sshIdleGraceSeconds` | `120` | Passed to user-data | +| `idleTimeoutMinutes` | `10` | Legacy name; **not** used for `ExpiresAt` anymore | + +### Mock / local dev + +Without `EC2_LAUNCH_TEMPLATE_ID`, provisioning is mock-only: no scheduler, no host watcher. Sessions are in-memory unless `DATABASE_URL` is set. See [local-dev-without-aws.md](local-dev-without-aws.md). + +## CDK components + +| Resource | Stack | Role | +|----------|-------|------| +| `OpsCoachSessionTerminator` Lambda | `-OpsCoachLabHost` | Direct terminate, sweep, and shutdown-API notify | +| `OpsCoachSchedulerInvoke` IAM role | Lab host | Lets Scheduler invoke the Lambda | +| EventBridge Rule (5 min) | Lab host | Sweep trigger | +| Callback secret | Lab host | Shared with Fargate | +| Scheduler IAM on task role | `-OpsCoach` / web stack | `CreateSchedule` / `DeleteSchedule` | + +Deploy wiring: [`../infra/bin/opscoach-platform.ts`](../infra/bin/opscoach-platform.ts), [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md). + +## Alternatives considered + +| Approach | Rejected because | +|----------|------------------| +| Idle timeout from provision only (old `ExpiresAt = now + 10m`) | Kills active sessions; not true idle semantics | +| Web-only activity timeout | No visibility into SSH; a learner can be active in the terminal while the web tab is idle | +| Tailscale-only SSH | Explicit v1 product decision: hardened public SSH only | +| Lambda per session (standalone) | Scheduler plus one shared terminator Lambda is simpler and cheaper | +| EC2 instance self-terminate via IAM | Broader blast radius on a compromised lab host; prefer API/Lambda with a secret | + +## Future improvements + +- **Source-aware idle detection:** count only SSH from non-VPC (learner) addresses, so web-only grading does not arm the idle watcher incorrectly. +- **Activity extension:** refresh `ExpiresAt` and reschedule the max TTL on a grader run or explicit heartbeat (trading cost for longer labs). +- **Metrics:** CloudWatch counters per teardown reason (`ssh_idle`, `max_ttl`, `expires_at_sweep`, `manual`) to tune grace and TTL. +- **Hardened AMI:** a Packer image with Docker, fail2ban, and the watcher baked in, instead of a full user-data bootstrap. + +## Related files + +- [`../web/lib/ec2-labs.ts`](../web/lib/ec2-labs.ts): provision, tag, schedule +- [`../web/lib/session-scheduler.ts`](../web/lib/session-scheduler.ts): EventBridge Scheduler client +- [`../web/app/api/sessions/[id]/shutdown/route.ts`](../web/app/api/sessions/[id]/shutdown/route.ts): internal shutdown API +- [`../infra/lib/lab-host-stack.ts`](../infra/lib/lab-host-stack.ts): terminator Lambda and scheduler role +- [`../infra/lib/opscoach-service-stack.ts`](../infra/lib/opscoach-service-stack.ts): Fargate env and scheduler IAM diff --git a/docs-draft-v2/local-dev-without-aws.md b/docs-draft-v2/local-dev-without-aws.md new file mode 100644 index 0000000..aa8d575 --- /dev/null +++ b/docs-draft-v2/local-dev-without-aws.md @@ -0,0 +1,72 @@ +# Local development without AWS + +**Status: deferred.** Most of the web app already runs on a developer machine with no AWS; the gap is automatic per-session Docker, which is not yet built. This doc records what works today and what a full local-Docker mode (`OPSCOACH_LOCAL_DEV`) would take, so the work can be picked up when local iteration becomes a bottleneck. The deploy target stays Fargate plus per-session EC2; only local dev would use Docker on the host. + +The bar to match is the native macOS app's loop: open the app, start a lab, SSH from your terminal, see live grading, stop the lab. + +## What already works (no AWS) + +| Piece | Behavior | +|-------|----------| +| Web UI | `cd web && npm run dev`: catalog, play flow, session page, SSE grading | +| Session store | In-memory when `DATABASE_URL` is unset; optional local Postgres | +| Provisioning | **Mock EC2** when `EC2_LAUNCH_TEMPLATE_ID` is unset: the session is immediately `ready` at `127.0.0.1:22` | +| Graders | The same ContentPack shell scripts as the native app (SSH from the API process) | +| Smoke | `scripts/smoke-web-session.sh`; with `OPSCOACH_SMOKE_START_COMPOSE=1`, it starts a foundations lab on port 22 | + +The catch: mock mode does not start a lab container for you. Something has to be listening on `:22`, so today you run Docker Compose yourself, or use the smoke script's compose helper. + +## The gap: no per-session Docker + +The native app's `ContainerManager` starts a per-session `docker compose` project, maps a dynamic SSH port, injects the learner and grader keys, and tears it down on stop. The web app does not implement that yet. Five things stand between mock mode and a hands-off local loop: + +| Gap | Impact | Likely fix | +|-----|--------|------------| +| No auto `docker compose` per session | Manual lab startup | `LocalLabProvisioner` behind `provisionLabInstance()` when `OPSCOACH_LOCAL_DEV=1` | +| AWS labs call STS/CFN on create | `aws-security-basics` fails without platform stacks | Skip `prepareAwsSession()` in local mode, or use fixture `aws-session/` files | +| Beaconkeeper seeds / images | The capstone needs the right image and seed | Read `runtime.directory` and `defaultSeed` from the manifest (as the native app does) | +| `next dev` + in-memory sessions | Hot reload clears sessions, so `/grade` returns 403 | Prefer `npm start` after a build, or use Postgres locally | +| Port 22 conflicts | Only one lab on the default port | Dynamic host ports (`127.0.0.1::22` in compose) | + +## Proposed local mode + +Gate the behavior on `OPSCOACH_LOCAL_DEV=1`, with `EC2_LAUNCH_TEMPLATE_ID` left unset so provisioning does not fall into the AWS path: + +```bash +# web/.env.local (future) +OPSCOACH_LOCAL_DEV=1 +# EC2_LAUNCH_TEMPLATE_ID unset +CONTENT_ROOT=../ContentPacks +SESSIONS_ROOT=/tmp/opscoach-sessions +``` + +When `OPSCOACH_LOCAL_DEV=1`: + +1. `provisionLabInstance()` runs `docker compose` from the lab's `runtime.directory`. +2. It maps a free host port and sets `sshHost` / `graderHost` to `127.0.0.1` and that port. +3. It injects the learner public key and the per-session grader key into the container. +4. `stop` runs `docker compose down` for that session's project. +5. It skips AWS lab prep, the EventBridge Scheduler, and the EC2 terminate paths. + +Optionally, a `scripts/dev.sh` would check Docker, export the env, and start Next.js. + +## Effort estimate + +Deferred, so these are sizing guesses, not commitments: + +| Scope | Effort | Outcome | +|-------|--------|---------| +| MVP: `linux-foundations` auto-docker | ~1 to 2 days | Full UI and grader loop without AWS | +| All SSH packs plus Beaconkeeper | ~3 to 5 days | Parity with the native non-AWS labs | +| AWS lab local | Extra | Mock/fixture credentials, or an explicit "requires platform" error | + +## Why this shape + +Keep one orchestration interface, `provisionLabInstance` / `terminateLabInstance`, with two implementations behind it: production runs EC2 user-data plus Docker on the instance, local runs Docker on the developer's machine. The API routes, graders, and ContentPacks stay identical across both, so local dev exercises the real code paths rather than a parallel mock. + +## References + +- Mock EC2: `web/lib/ec2-labs.ts` (`isMockEc2Mode()`) +- Compose smoke: `scripts/smoke-web-session.sh` +- Lab teardown in production: [lab-lifecycle-design.md](lab-lifecycle-design.md) +- Production deploy: [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md) diff --git a/docs-draft-v2/security.md b/docs-draft-v2/security.md new file mode 100644 index 0000000..a1e0d0a --- /dev/null +++ b/docs-draft-v2/security.md @@ -0,0 +1,43 @@ +# Security model + +OpsCoach gives authenticated users a real Linux machine and lets them run real commands as a privileged user. That is the product, and it is the central security problem. The bet that contains it: **assume the lab container can be escaped, and design so it does not matter.** + +## The core bet + +A learner doing a systemd or filesystem lab needs real control of the box, so the lab container runs `--privileged`. Rather than treat that container as a strong wall, OpsCoach treats the **EC2 host as the real boundary** and makes the host cheap to lose: + +- **One host per session.** No shared tenancy, so a learner can only ever reach their own machine. +- **Nothing valuable lives on it.** The host's IAM role can pull its lab image and write logs, and that is all. Database credentials, the callback secret, and grader keys never touch the host. +- **It dies on a timer.** Idle for 10 minutes, or 60 minutes old, whichever comes first. Teardown is enforced from outside the host, so a wedged or hostile box still gets killed (see [lab-lifecycle-design.md](lab-lifecycle-design.md)). + +Worst case for an escaped container: full control of one throwaway machine, with no useful credentials, for at most an hour. + +## Getting in: two gates + +Reaching any page takes two gates. The shared ALB authenticates every request through Cognito (hosted UI, Google), and behind it the app requires a shared passphrase before any page renders (`web/lib/gate.ts`); only the ALB health check and logout skip Cognito. Each user is capped at a few concurrent sessions, and every session carries a short-lived bearer token checked against a hash in the database. + +## The host is built to be worthless to steal + +Once a host is running, an escaped container should find nothing worth taking and nowhere to go. + +There are no credentials to read. The host's IAM role is read-only (pull `opscoach-lab*` images, write to the lab log group, nothing more), and that role is firewalled off from the container anyway: a `DOCKER-USER` iptables rule drops traffic to `169.254.169.254`, and the host requires IMDSv2 with a one-hop limit, so even a privileged container cannot read the instance role or any secret from instance metadata. + +The network path is one-way. The host's security group accepts learner SSH from the internet but accepts the grader's SSH only from the web service's security group, so no learner can reach anyone else's box, and the database sits in an isolated subnet behind the private web service. + +## Real credentials stay off the host + +The pieces that hold real credentials never live where an escaped container could reach them. Database credentials and the callback HMAC secret live in Secrets Manager, read at runtime by the web task and the terminator Lambda, never written to a host or baked into an image. Grading is least-privilege: the grader runs as a subprocess with an allowlisted environment (`PATH`, `HOME`, `LANG`, `AWS_REGION`, and a few more) and none of the web task's AWS credentials, while AWS labs that need cloud access get their own scoped, per-session STS credentials. Each session also mints two fresh keypairs, the learner's (public key on the box, private key theirs) and the grader's (server-only), so a leaked learner key reaches one already-owned box and nothing else. + +## Deliberate trade-offs + +The sharp edges, and why they are acceptable here. + +- **The lab container is privileged.** Needed for systemd and realistic admin work. Acceptable because the host around it is single-tenant, credential-poor, and short-lived (see the core bet). +- **The terminal bridge skips SSH host-key verification** (`hostVerifier: () => true` in `web/server.js`). The host was created seconds earlier with a per-session key and is reached over the AWS private network, so there is no prior key to pin and no third party in the path. The keys are discarded with the session. +- **Learner SSH is open to the internet.** That is the feature. The grader's path is not open; it is locked to the web service's security group. + +## What this is not + +OpsCoach is training infrastructure, not a hostile multi-tenant sandbox. It does not stop a learner from using their own host's outbound network during the session, and it bounds spend with a per-user session cap rather than hard cost controls. Both are reasonable given the audience and the one-hour blast radius. + +Teardown is the load-bearing control, so it has three independent layers: an SSH-idle watcher, a one-shot timer, and a periodic sweep. The full design is in **[lab-lifecycle-design.md](lab-lifecycle-design.md)**. diff --git a/docs-draft-v3/README.md b/docs-draft-v3/README.md new file mode 100644 index 0000000..9389330 --- /dev/null +++ b/docs-draft-v3/README.md @@ -0,0 +1,62 @@ +# OpsCoach + +OpsCoach teaches Linux and AWS operations on a real, throwaway cloud server instead of a quiz. You get a live Linux host and a task list, and a grader checks the actual state of the machine as you work. + +This is a portfolio project. Three parts are worth a look: a live browser-to-SSH terminal, per-session disposable lab hosts on AWS, and grading that reads real system state instead of comparing answers. + +## Start here + +This page is the two-minute pitch. The system walkthrough is **[architecture.md](architecture.md)**, which starts at 30,000 feet (the three moving parts), drops to 10,000 feet (the diagram, components, and flows), and lands at 1,000 feet (the design decisions worth defending). + +From there, two ground-level docs sit next to the code: the **[security model](security.md)** explains how untrusted users get root without putting anything else at risk, and the **[lab lifecycle](lab-lifecycle-design.md)** explains how a host gets provisioned and reliably destroyed. **[Local dev](local-dev-without-aws.md)** covers running the app with no AWS. + +## What it does + +A learner opens a lab and gets a dedicated Linux host in a minute or two. They work in a terminal, in the browser or over their own SSH client, to fix broken services, harden config, triage logs, wire up systemd, and so on. A grader runs against the live machine and streams pass/fail checks to a dashboard beside the terminal. When the learner finishes or goes idle, the host is destroyed. + +Three content packs ship today: **Linux Foundations** and **AWS Foundations** drills, and the **Beaconkeeper** capstone, a 20-step systemd and operations scenario on a single box. + +## How it works + +A request reaches a shared ALB, authenticates through Cognito, and lands on the OpsCoach web app on ECS Fargate. Starting a lab launches a dedicated EC2 host; the web app bridges the browser terminal to it over SSH and runs the grader over SSH. Every host self-destructs on an idle or lifetime timer. [architecture.md](architecture.md) has the diagram and the full provision, terminal, grade, and teardown flows. + +## Why it is built this way + +Four decisions shaped the system, each a fork where the cheap option was the wrong one. + +**Real hosts, not a fake shell.** Learning operations means touching real systemd, real packages, real logs, so each session is an actual EC2 instance rather than an emulation. + +**Disposable and single-tenant.** Labs run as a privileged user, so every session gets its own host that is wiped on a timer. The security model leans on the cheap, isolated host instead of on container isolation. See [security.md](security.md). + +**Grade real state, not answers.** The grader SSHes in and inspects the machine, so a check passes only when the box is genuinely in the right state. + +**Borrow the platform, do not rebuild it.** The infrastructure plugs into an existing shared ALB, Cognito, and VPC by ID instead of standing up its own. The real resource IDs live in local config kept out of the repo. + +## Repository layout + +| Path | What it is | +| --- | --- | +| `web/` | Next.js app and custom Node server (WebSocket-to-SSH bridge), API routes, UI | +| `infra/` | AWS CDK app: Fargate service, lab hosts, teardown automation | +| `ContentPacks/` | Labs and graders, including the Beaconkeeper game | +| `scripts/` | Build, deploy, and smoke-test scripts | +| `docs/` | Architecture, security, and design notes | + +## Run it locally + +```bash +cd web +npm install +cp .env.example .env # most values are optional locally +npm run dev +``` + +With no AWS credentials and no `DATABASE_URL`, the app runs in mock mode: an in-memory store and faked provisioning, so you can click through the UI and content offline. Run the tests with `npm test`. Details in [local-dev-without-aws.md](local-dev-without-aws.md). + +## Configure and deploy + +Two things stay out of version control: the app's `web/.env` (copied from `web/.env.example`) and the CDK platform context (`infra/cdk.context.json`, copied from the example or generated by `scripts/discover-platform-context.sh`). With those in place, `scripts/deploy-platform.sh` builds the images and deploys the stacks. The walkthrough is in [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md). + +## License + +MIT. See [`../LICENSE`](../LICENSE). diff --git a/docs-draft-v3/architecture.md b/docs-draft-v3/architecture.md new file mode 100644 index 0000000..912ad85 --- /dev/null +++ b/docs-draft-v3/architecture.md @@ -0,0 +1,165 @@ +# OpsCoach architecture + +OpsCoach gives each learner a real, throwaway Linux host and grades what they actually do to it. This doc starts at 30,000 feet and drops to the decisions worth defending. Skim the top, stop where you care. + +**Scope:** single-region training infrastructure for one app that borrows a shared ALB/Cognito platform. Not a hostile multi-tenant sandbox, and not a high-availability production service. + +## 30,000 ft: the shape + +The whole system is three parts: + +- A **web app** (Next.js on ECS Fargate) that serves the UI, bridges the terminal, and runs grading. +- A **per-session lab host** (a dedicated EC2 instance) that the learner operates and that is destroyed on a timer. +- A **shared platform** (ALB, Cognito, VPC) that the app plugs into instead of rebuilding. + +Everything below is how those three talk to each other safely. + +## 10,000 ft: the system + +![OpsCoach AWS architecture](architecture.svg) + +The numbered arrows above: + +1. **Request.** Browser to the ALB over HTTPS. The ALB runs `authenticate-cognito` (Cognito hosted UI, Google); a post-auth passphrase gate in the app must also pass before any page renders. Only the health check and logout skip Cognito. +2. **Route.** The ALB forwards to the OpsCoach web service on Fargate, in a private, egress-only subnet. +3. **Provision, operate, grade.** Fargate launches and drives the per-session EC2 host: `RunInstances` at the start, an SSH PTY for the browser terminal, and the grader over SSH. +4. **Ready webhook.** The host calls back to the service once it is up, resolved through Cloud Map and authenticated with a shared secret. +5. **Terminal.** The browser streams over a WebSocket to the custom Node server, which relays to the host's shell over SSH. +6. **Teardown.** A per-session EventBridge Scheduler one-shot triggers the terminator Lambda; a 5-minute sweep is the backstop. + +Grey dashed lines are supporting paths: Fargate to RDS PostgreSQL in the isolated subnet, the host pulling its lab image from ECR, and the opt-in direct SSH from a learner's laptop. + +Colour key: purple is networking (ALB, Cloud Map); orange is compute and containers (Fargate, EC2, Lambda, ECR); red is identity and secrets (Cognito, Secrets Manager); blue is the database (RDS); pink is app integration (EventBridge Scheduler). + +### Components + +| Component | AWS service | Role | +| --- | --- | --- | +| Web app + terminal bridge | ECS Fargate | Next.js app + custom Node server; WebSocket-to-SSH PTY bridge; session lifecycle, grading, dashboard | +| Edge auth | ALB + Cognito | `authenticate-cognito` at the load balancer (hosted UI to Google), then a post-auth passphrase gate | +| Lab host | EC2 (per session) | Ephemeral AL2023 / arm64 host running the lab container; learner SSH target | +| Container images | ECR | Images for the web service and each lab | +| Database | RDS PostgreSQL | Sessions, check runs, grader results (isolated subnet) | +| Service discovery | Cloud Map | In-VPC address for lab-host to web callbacks | +| Teardown | EventBridge Scheduler + Lambda | One-shot per-session schedule fires a terminator Lambda; 5-minute sweep backstop | +| Secrets | Secrets Manager | Database credentials and the callback HMAC secret | + +
+The flows, step by step (auth, provision, terminal, grading, teardown) + +  + +**1 · Authentication and access gate** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant ALB as ALB + participant Cog as Cognito + participant App as Fargate + U->>ALB: HTTPS request + ALB->>Cog: authenticate-cognito (if enabled) + Cog-->>ALB: OIDC tokens + ALB->>App: forward + signed identity header + alt no gate cookie + App-->>U: redirect to /gate + U->>App: POST /api/gate (passphrase) + App-->>U: set gate cookie, continue + end + App-->>U: app page + Note over ALB,App: /api/health and /logout skip Cognito +``` + +**2 · Provision a lab session** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant App as Fargate + participant EC2 as Lab host + participant ECR as ECR + participant Sch as Scheduler + U->>App: POST /api/sessions (start lab) + App->>App: create session in RDS; mint keys + callback token + App->>EC2: RunInstances (Launch Template + user-data) + App->>Sch: CreateSchedule (terminate at T + maxLifetime) + Note over EC2: user-data: install Docker, block IMDS,
ECR login, run lab container, set authorized_keys + EC2->>ECR: pull lab image + EC2->>App: ready webhook (via Cloud Map + shared secret) + App-->>U: session ready (host, port) +``` + +**3 · Browser terminal and SSH** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant Srv as Node server + participant EC2 as Lab host + U->>Srv: WebSocket /api/sessions/:id/shell?token + Srv->>Srv: shell-auth (internal secret) resolves host + key + Srv->>EC2: ssh2 PTY on port 22 (per-session key) + EC2-->>Srv: stdout / stderr + Srv-->>U: stream to xterm.js + Note over Srv,EC2: 25s keepalive ping holds the ALB idle timer + Note over U,EC2: opt-in: SSH directly to the host's public IP +``` + +**4 · Live grading** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant App as Fargate + participant G as Grader + participant EC2 as Lab host + U->>App: POST /api/sessions/:id/grade + App->>G: spawn grader (allowlisted env, no task-role creds) + G->>EC2: SSH port 22 (grader key) runs ops status + EC2-->>G: JSON check results + G-->>App: { passed, checks[] } + App->>App: persist check run in RDS + App-->>U: live results on the dashboard +``` + +**5 · Idle and lifetime teardown** + +```mermaid +sequenceDiagram + autonumber + participant Sch as Scheduler + participant L as Terminator Lambda + participant EC2 as Lab host + participant App as Fargate + Sch->>L: fire at T + maxLifetime + L->>L: read callback secret (Secrets Manager) + L->>EC2: TerminateInstances + L->>App: shutdown callback (mark session stopped) + Note over EC2,App: backstop: 5-min EventBridge sweep
terminates any host past its ExpiresAt tag +``` + +
+ +## 1,000 ft: the decisions that shaped it + +Each of these was a fork where the obvious choice was the wrong one: constraint, decision, trade-off. + +**A real host per session, not a shared sandbox.** Teaching operations means real systemd, real root, real packages, but you cannot hand untrusted users root on shared infrastructure. The cheaper alternatives (a shared multi-tenant host, or browser-only containers) make container isolation the only wall between a hostile learner and everyone else, so one escape compromises every session. So every session gets its own ephemeral EC2 host, and the security model treats that host, not the container on it, as the real boundary, killed on a timer. The cost is a minute or two of provisioning latency and per-session spend, bought back as realism and a blast radius of exactly one throwaway box. Full model in [security.md](security.md). + +**A custom Node server for the browser terminal.** A browser terminal needs a long-lived, two-way connection to a shell, and Next.js on its own does not hold one. So a thin `server.js` wraps Next, upgrades the WebSocket, authenticates the session through an internal call, and bridges to the host with an `ssh2` PTY. The cost is a custom server instead of stock Next, plus a 25-second keepalive so the ALB does not cut an idle terminal, in exchange for a real shell in the browser with nothing to install. + +**Grade real state, not answers.** Multiple-choice cannot tell you whether someone can actually run a box. So the grader SSHes into the live host with a least-privilege environment and checks real state (services up, files in place, config correct), returning structured results. The cost is that graders are per-pack code that has to run against a live machine, in exchange for a pass that means the box is genuinely in the right state. + +**Three independent teardown paths.** A leaked instance costs real money, and the control plane never sees the learner's SSH activity. So teardown runs three ways: an on-host SSH-idle watcher, a one-shot EventBridge Scheduler set at provision time, and a 5-minute sweep over expiry tags. Any one is enough, and all are idempotent. The cost is more moving parts, for a hard guarantee that nothing runs forever. Full design in [lab-lifecycle-design.md](lab-lifecycle-design.md). + +**Borrow the platform; import by ID.** A demo should not stand up its own ALB, Cognito, and VPC. So the CDK imports a shared platform's resources by ID from local context, and the real IDs stay out of the repo. The cost is that the app cannot bootstrap its own world from nothing, in exchange for dropping cleanly into a real shared environment. + +## See also + +- **[security.md](security.md)** for the security model and its trade-offs. +- **[lab-lifecycle-design.md](lab-lifecycle-design.md)** for provisioning and the three-layer teardown in depth. +- **[../infra/PLATFORM_INTEGRATION.md](../infra/PLATFORM_INTEGRATION.md)** for plugging into a shared ALB/Cognito platform. diff --git a/docs-draft-v3/architecture.svg b/docs-draft-v3/architecture.svg new file mode 100644 index 0000000..4ec1f3d --- /dev/null +++ b/docs-draft-v3/architecture.svg @@ -0,0 +1,149 @@ + + OpsCoach AWS architecture + A user reaches a shared Application Load Balancer with Cognito auth and a passphrase gate, which forwards to the OpsCoach web service on ECS Fargate in a private subnet. Fargate uses an isolated-subnet RDS PostgreSQL database, launches per-session EC2 lab hosts in a public subnet (which pull a lab container from ECR and call back via Cloud Map), bridges the browser terminal to the host over SSH, runs a grader over SSH, and schedules automatic teardown via EventBridge Scheduler and a terminator Lambda. + + + + + + + + + + + + + OpsCoach — AWS architecture + In-browser & SSH Linux labs · shared ALB/Cognito platform · per-session EC2 hosts · live grading · auto-teardown + + + USER + + Browser + xterm.js terminal + + SSH client + direct opt-in + + + + + AWS Account · Region (us-east-1) + + + + Shared platform — imported by the app (not created here) + + Application Load Balancer + Cognito auth · routing + + Amazon Cognito + hosted UI · Google IdP + + + + VPC + + + + Public subnet + + EC2 — lab host (per session) + AL2023 · arm64 · Launch Template + runs lab container · IMDS blocked + + + + Private subnet (egress) + + ECS Fargate — OpsCoach web + Next.js + custom Node server + WebSocket → SSH PTY bridge + + AWS Cloud Map + in-VPC discovery + + + + Isolated subnet + + Amazon RDS — PostgreSQL + sessions · check runs · results + + + + Automation & services + + Amazon ECR + web + lab images + + Secrets Manager + DB · callback · keys + + EventBridge + Scheduler + one-shot per session + + AWS Lambda + session terminator + + + + + 1 + + + + 2 + + + + 3 + + provision · SSH PTY · grade + + + + 4 + + + ready webhook + + + + 5 + + WSS + + + + SSH opt-in + + + + + 6 + + terminate EC2 + + + + SQL + + + + pull lab image + + + diff --git a/docs-draft-v3/lab-lifecycle-design.md b/docs-draft-v3/lab-lifecycle-design.md new file mode 100644 index 0000000..f6d0eae --- /dev/null +++ b/docs-draft-v3/lab-lifecycle-design.md @@ -0,0 +1,175 @@ +# Lab instance lifecycle design + +**Status:** implemented (web + CDK). + +Each learner session runs on a dedicated EC2 lab host (`t4g.micro`, Docker, one lab container). A leaked host costs money, so teardown is the load-bearing control. OpsCoach uses **three independent teardown paths** instead of one: any single path is enough to kill a host, and all three are safe to run more than once. The rest of this doc is why one path was not, and how each works. + +## Problem + +The web product does not embed a terminal. Learners SSH from their own client directly to the lab host's public IP, so the control plane (Next.js on Fargate) never sees SSH connect or disconnect events and cannot infer idle from browser activity. If teardown fails or never runs, instances leak and accumulate cost. + +Teardown therefore has to be: + +- **Prompt** when the learner is done, which means detecting SSH idle on the host itself. +- **Reliable** when callbacks fail, schedules are missed, or the learner never connects. +- **Idempotent** when several paths fire close together. + +Non-goals for v1: SSH over a private overlay (no Tailscale; public IP only), and activity-based lifetime extension. + +## The shape + +```mermaid +sequenceDiagram + participant Learner + participant Web as OpsCoachWeb_Fargate + participant EC2 as LabHost_EC2 + participant Lambda as SessionTerminator + participant Scheduler as EventBridgeScheduler + + Learner->>Web: POST /api/sessions (pubkey) + Web->>EC2: RunInstances + user-data + Web->>Scheduler: CreateSchedule T+maxLifetime + EC2->>Web: POST /ready (public + private IP) + Learner->>EC2: SSH :22 + Note over EC2: idle watcher monitors :22 + Learner--xEC2: SSH disconnect + EC2->>Web: POST /shutdown reason=ssh_idle + Web->>EC2: TerminateInstances + Web->>Scheduler: DeleteSchedule + + Note over Scheduler,Lambda: If still running at T+maxLifetime + Scheduler->>Lambda: invoke terminate + Lambda->>EC2: TerminateInstances + Lambda->>Web: POST /shutdown reason=max_ttl +``` + +SSH is pubkey-only on a hardened host. The grader runs from Fargate inside the VPC and reaches the instance over its **private** IP. Session state lives in Postgres; EC2 lifecycle is driven by the web task role and the terminator Lambda. + +## The three paths + +Each path stands alone. The table is the summary; the sections below are the mechanics and the reasoning for each tunable. + +| Layer | Trigger | Actor | Typical latency | +|-------|---------|-------|-----------------| +| 1. SSH idle watcher | No established TCP sessions on host `:22` for the grace period, after at least one was seen | EC2 user-data background script | ~2 min after disconnect | +| 2. Max TTL schedule | One-time EventBridge Scheduler set at provision time | `OpsCoachSessionTerminator` Lambda | Exactly at T + max lifetime | +| 3. ExpiresAt sweep | `ExpiresAt` EC2 tag in the past | Same Lambda, every 5 min | Up to 5 min after tag expiry | + +Manual **Stop lab** (the learner button) and authenticated `POST /api/sessions/:id/stop` use the same internal shutdown path as the webhooks. + +Why three and not one: the idle watcher is prompt but blind to the case where the learner never connects; the timer is reliable but either too aggressive (kills active sessions) or too loose (leaks nodes) on its own; the sweep catches the host whose schedule was never created. Each covers the others' failure mode. + +### Layer 1: SSH idle watcher + +**Where:** generated shell user-data in [`../web/lib/lab-user-data.ts`](../web/lib/lab-user-data.ts), mirrored in [`../infra/lib/lab-user-data.sh`](../infra/lib/lab-user-data.sh) for launch-template defaults. + +A background subshell loops every 15 seconds and counts established connections on local port 22 via `ss -tn state established '( sport = :22 )'`. It tracks `had_session`, set once the count exceeds zero. When `had_session` is set and the count returns to zero, it starts an idle clock; after `SSH_IDLE_GRACE_SECONDS` (default **120**) it POSTs the shutdown webhook: + +```http +POST /api/sessions/:id/shutdown +X-Internal-Secret: +Content-Type: application/json + +{ "reason": "ssh_idle" } +``` + +The 120-second grace avoids tearing down during brief disconnects (a network blip or an `ssh` reconnect) and lets a grader SSH from Fargate finish without racing the learner's disconnect. + +Two known limits, both acceptable for v1. The watcher counts **all** connections on `:22`, including grader SSH from the VPC, so a learner who only runs checks from the web UI (never opening a terminal) may see the host torn down shortly after grading quiesces; a follow-up could filter by source IP. And if the learner never SSHes at all, `had_session` stays zero and this layer never fires, which is exactly what layer 2 is for. + +### Layer 2: one-time EventBridge schedule (max TTL) + +**Where:** [`../web/lib/session-scheduler.ts`](../web/lib/session-scheduler.ts), invoked from [`../web/lib/ec2-labs.ts`](../web/lib/ec2-labs.ts) after `RunInstances`. + +On a successful provision, Fargate creates schedule `opscoach-{sessionId}` (truncated to 64 chars) with an `at(...)` expression in UTC set to **T + maxLifetimeMinutes**. The target is the `OpsCoachSessionTerminator` Lambda: + +```json +{ "action": "terminate", "instanceId": "i-…", "sessionId": "…", "reason": "max_ttl" } +``` + +`ActionAfterCompletion: DELETE` removes the schedule once it fires, and any explicit shutdown (`manual`, `ssh_idle`) calls `DeleteSchedule` for idempotency. + +The default max lifetime is **60 minutes** (`OPSCOACH_MAX_LIFETIME_MINUTES` / CDK context `maxLifetimeMinutes`): long enough for a typical session and a few assessment retries, short enough to bound cost if every other path fails. It is orthogonal to the SSH idle grace. + +EventBridge **Scheduler** rather than EventBridge **Rules** because Scheduler supports one-time schedules natively, with per-session names and auto-delete. Rules suit recurring patterns, which is what layer 3 uses. + +### Layer 3: ExpiresAt tag sweep + +**Where:** [`../infra/lib/session-terminator/handler.py`](../infra/lib/session-terminator/handler.py), triggered every 5 minutes by an EventBridge Rule in [`../infra/lib/lab-host-stack.ts`](../infra/lib/lab-host-stack.ts). + +At provision, `RunInstances` tags the instance `ExpiresAt=` on the same horizon as the scheduler. The Lambda scans running OpsCoach instances (`OpsCoach=true`), and terminates any whose `ExpiresAt` is in the past, calling the shutdown API with `reason=expires_at_sweep`. This is the cheap safety net for the host whose schedule was never created (missing IAM, an API error during provision) or whose schedule drifted. + +## Unified shutdown path + +Every automated and manual teardown converges on [`shutdownSessionInternal`](../web/lib/sessions.ts): + +1. Return early if already `stopped` or `stopping`. +2. Set status `stopping`. +3. `DeleteSchedule` (best effort). +4. `TerminateInstances` if an instance id is present (errors logged, still mark stopped). +5. Set status `stopped`, publish an SSE event. + +There are two entry points. `POST /api/sessions/:id/stop` is authenticated by the session token and carries `reason=manual`. `POST /api/sessions/:id/shutdown` is authenticated by `X-Internal-Secret` and carries `ssh_idle`, `max_ttl`, `expires_at_sweep`, or `manual`. The terminator Lambda terminates EC2 first, then calls the shutdown API, so Postgres stays in sync. + +## Configuration + +**Runtime (Fargate task environment):** + +| Variable | Default | Purpose | +|----------|---------|---------| +| `OPSCOACH_MAX_LIFETIME_MINUTES` | `60` | Scheduler fire time + `ExpiresAt` tag | +| `OPSCOACH_SSH_IDLE_GRACE_SECONDS` | `120` | Host idle debounce before the shutdown webhook | +| `SESSION_TERMINATOR_LAMBDA_ARN` | (CDK) | Schedule target | +| `SCHEDULER_INVOKE_ROLE_ARN` | (CDK) | Scheduler execution role | +| `INTERNAL_CALLBACK_SECRET` | Secrets Manager | Authenticates host and Lambda callbacks | + +**CDK context** ([`../infra/lib/web-config.ts`](../infra/lib/web-config.ts)): + +| Key | Default | Purpose | +|-----|---------|---------| +| `maxLifetimeMinutes` | `60` | Hard cap | +| `sshIdleGraceSeconds` | `120` | Passed to user-data | +| `idleTimeoutMinutes` | `10` | Legacy name; **not** used for `ExpiresAt` anymore | + +**Mock / local dev:** without `EC2_LAUNCH_TEMPLATE_ID`, provisioning is mock-only (no scheduler, no host watcher), and sessions are in-memory unless `DATABASE_URL` is set. See [local-dev-without-aws.md](local-dev-without-aws.md). + +## Security + +Shutdown and ready callbacks require `X-Internal-Secret` (stored in Secrets Manager, created in the lab-host stack, read by the Fargate task role). EC2 terminate IAM is scoped with `OpsCoach=true` resource and request tags where possible. The host watcher can only initiate shutdown through the API; the lab instance role has no terminate permission, so it cannot kill instances directly. + +## CDK components + +| Resource | Stack | Role | +|----------|-------|------| +| `OpsCoachSessionTerminator` Lambda | `Dev-OpsCoachLabHost` | Direct terminate + sweep + shutdown API notify | +| `OpsCoachSchedulerInvoke` IAM role | Lab host | Lets Scheduler invoke the Lambda | +| EventBridge Rule (5 min) | Lab host | Sweep trigger | +| Callback secret | Lab host | Shared with Fargate | +| Scheduler IAM on task role | `Dev-OpsCoach` / web stack | `CreateSchedule` / `DeleteSchedule` | + +Deploy wiring: [`../infra/bin/opscoach-platform.ts`](../infra/bin/opscoach-platform.ts), [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md). + +## Alternatives considered + +| Approach | Rejected because | +|----------|------------------| +| Idle timeout from provision only (old `ExpiresAt = now + 10m`) | Kills active sessions; not true idle semantics | +| Web-only activity timeout | No visibility into SSH; a learner can be active in the terminal while the web tab is idle | +| Tailscale-only SSH | Explicit product decision: v1 is hardened public SSH only | +| Lambda per session (standalone) | Scheduler plus one shared terminator Lambda is simpler and cheaper | +| EC2 self-terminate via IAM | Broader blast radius on a compromised lab host; prefer API/Lambda gated by a secret | + +## Future improvements + +- **Source-aware idle detection:** count only SSH from non-VPC (learner) addresses, so web-only grading does not arm the idle watcher. +- **Activity extension:** refresh `ExpiresAt` and reschedule max TTL on a grader run or explicit heartbeat (trade cost for longer labs). +- **Metrics:** CloudWatch counters per teardown reason (`ssh_idle`, `max_ttl`, `expires_at_sweep`, `manual`) to tune the grace and TTL. +- **Hardened AMI:** a Packer image with Docker, fail2ban, and the watcher baked in, instead of a full user-data bootstrap. + +## Related files + +- [`../web/lib/ec2-labs.ts`](../web/lib/ec2-labs.ts): provision, tag, schedule. +- [`../web/lib/session-scheduler.ts`](../web/lib/session-scheduler.ts): EventBridge Scheduler client. +- [`../web/app/api/sessions/[id]/shutdown/route.ts`](../web/app/api/sessions/[id]/shutdown/route.ts): internal shutdown API. +- [`../infra/lib/lab-host-stack.ts`](../infra/lib/lab-host-stack.ts): terminator Lambda + scheduler role. +- [`../infra/lib/opscoach-service-stack.ts`](../infra/lib/opscoach-service-stack.ts): Fargate env + scheduler IAM. diff --git a/docs-draft-v3/local-dev-without-aws.md b/docs-draft-v3/local-dev-without-aws.md new file mode 100644 index 0000000..4af7c48 --- /dev/null +++ b/docs-draft-v3/local-dev-without-aws.md @@ -0,0 +1,69 @@ +# Local development without AWS + +**Status: deferred.** The web app already runs offline well enough to build most features (mock provisioning, in-memory or local Postgres, the real graders). What it does not yet do is start a lab container for you per session, the way the native macOS app does. This doc records what works today, the gap, and the local mode that would close it. The plan is to ship the platform deploy first and build `OPSCOACH_LOCAL_DEV` when local iteration becomes a bottleneck. + +## Goal + +Match the macOS app's day-to-day loop on a developer machine, with no EC2, Fargate, RDS, or AWS API calls: open the app, start a lab, SSH from your terminal, see live grading, stop the lab. Production still uses Fargate plus per-session EC2; local dev would use Docker on the host. + +## What already works (no AWS) + +| Piece | Behavior | +|-------|----------| +| Web UI | `cd web && npm run dev`: catalog, play flow, session page, SSE grading | +| Session store | In-memory when `DATABASE_URL` is unset; optional local Postgres | +| Provisioning | **Mock EC2** when `EC2_LAUNCH_TEMPLATE_ID` is unset, so a session is immediately `ready` at `127.0.0.1:22` | +| Graders | The same ContentPack shell scripts as the native app (SSH from the API process) | +| Smoke | `scripts/smoke-web-session.sh`; with `OPSCOACH_SMOKE_START_COMPOSE=1`, starts a foundations lab on port 22 | + +The catch today: mock mode does not start a lab container for you. You run Docker Compose yourself (or use the smoke script's compose helper) so something listens on `:22`. + +## What is missing for "just works" local dev + +The native app's `ContainerManager` starts a per-session `docker compose` project, maps a dynamic SSH port, injects learner and grader keys, and tears down on stop. The web app does not implement that yet. + +| Gap | Impact | Likely fix | +|-----|--------|------------| +| No auto `docker compose` per session | Manual lab startup | `LocalLabProvisioner` behind `provisionLabInstance()` when `OPSCOACH_LOCAL_DEV=1` | +| AWS labs call STS/CFN on create | `aws-security-basics` fails without platform stacks | Skip `prepareAwsSession()` in local mode, or use fixture `aws-session/` files | +| Beaconkeeper seeds / images | The capstone needs the correct image and seed | Read `runtime.directory` and `defaultSeed` from the manifest (same as native) | +| `next dev` plus in-memory sessions | Hot reload clears sessions, so `/grade` returns 403 | Prefer `npm start` after a build, or use Postgres locally | +| Port 22 conflicts | Only one lab on the default port | Dynamic host ports (`127.0.0.1::22` in compose) | + +## Proposed local mode + +```bash +# web/.env.local (future) +OPSCOACH_LOCAL_DEV=1 +# EC2_LAUNCH_TEMPLATE_ID unset +CONTENT_ROOT=../ContentPacks +SESSIONS_ROOT=/tmp/opscoach-sessions +``` + +When `OPSCOACH_LOCAL_DEV=1`, `provisionLabInstance()` would: + +1. Run `docker compose` from the lab's `runtime.directory`. +2. Map a free host port and set `sshHost` / `graderHost` to `127.0.0.1` and that port. +3. Inject the learner public key and a per-session grader key into the container. +4. On `stop`, run `docker compose down` for that session's project. +5. Skip AWS lab prep, EventBridge Scheduler, and the EC2 terminate paths. + +Optionally, a `scripts/dev.sh` would check Docker, export the env, and start Next.js. + +## Effort estimate (deferred) + +| Scope | Effort | Outcome | +|-------|--------|---------| +| MVP: `linux-foundations` auto-docker | ~1 to 2 days | Full UI plus grader loop without AWS | +| All SSH packs plus beaconkeeper | ~3 to 5 days | Parity with native non-AWS labs | +| AWS lab local | Extra | Mock/fixture credentials, or an explicit "requires platform" error | + +## Architecture note + +Keep one orchestration interface (`provisionLabInstance` / `terminateLabInstance`). Production runs EC2 user-data plus Docker on the instance; local would run Docker on the developer's machine. Same API routes, graders, and ContentPacks. That single seam is what keeps a local provisioner from forking the codebase. + +## References + +- Mock EC2: `web/lib/ec2-labs.ts` (`isMockEc2Mode()`). +- Compose smoke: `scripts/smoke-web-session.sh`. +- Production deploy: [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md). diff --git a/docs-draft-v3/security.md b/docs-draft-v3/security.md new file mode 100644 index 0000000..013e7f1 --- /dev/null +++ b/docs-draft-v3/security.md @@ -0,0 +1,41 @@ +# Security model + +OpsCoach gives authenticated users a real Linux machine and lets them run real commands as a privileged user. That is the product, and it is also the main security problem. Here is how the risk is contained. + +## The core bet + +Assume the lab container can be escaped, and design so it does not matter. + +A learner doing a systemd or filesystem lab needs real control of the box, so the lab container runs `--privileged`. Rather than treat that container as a strong wall, OpsCoach treats the **EC2 host as the real boundary** and makes the host cheap to lose: + +- **One host per session.** No shared tenancy, so a learner can only ever reach their own machine. +- **Nothing valuable lives on it.** The host's IAM role can pull its lab image and write logs, and that is all. Database credentials, the callback secret, and grader keys never touch the host. +- **It dies on a timer.** Idle for 10 minutes, or 60 minutes old, whichever comes first. Teardown is enforced from outside the host, so a wedged or hostile box still gets killed. + +Worst case for an escaped container: full control of one throwaway machine, with no useful credentials, for at most an hour. + +## Layers + +The controls, from the edge inward. + +**Getting in takes two gates.** The shared ALB authenticates every request through Cognito (hosted UI, Google), and behind it the app requires a shared passphrase before any page renders (`web/lib/gate.ts`); only the ALB health check and logout skip Cognito. Each user is capped at a few concurrent sessions, and every session carries a short-lived bearer token checked against a hash in the database. + +**A running host is built to be worthless to steal.** Its IAM role is read-only (pull `opscoach-lab*` images, write to the lab log group, nothing more), and that role is firewalled off from the container anyway: a `DOCKER-USER` iptables rule drops traffic to the instance metadata endpoint, and the host requires IMDSv2 with a one-hop limit, so even a privileged container cannot read the instance role or any secret from metadata. The network path is one-way. The host's security group takes learner SSH from the internet but takes the grader's SSH only from the web service's security group, so no learner can reach anyone else's box, and the database sits in an isolated subnet behind the private web service. + +**The pieces that hold real credentials stay off the host entirely.** Database credentials and the callback HMAC secret live in Secrets Manager, read at runtime by the web task and the terminator Lambda, never written to a host or baked into an image. Grading is least-privilege: the grader is a subprocess with an allowlisted environment (`PATH`, `HOME`, `LANG`, `AWS_REGION`, and a few more) and none of the web task's AWS credentials, while AWS labs that need cloud access get their own scoped, per-session STS credentials. Each session also mints two fresh keypairs, the learner's (public key on the box, private key theirs) and the grader's (server-only), so a leaked learner key reaches one already-owned box and nothing else. + +## Deliberate trade-offs + +The sharp edges, and why they are acceptable here. + +**The lab container is privileged.** Needed for systemd and realistic admin work. Acceptable because the host around it is single-tenant, credential-poor, and short-lived (see the core bet). + +**The terminal bridge skips SSH host-key verification** (`hostVerifier: () => true` in `web/server.js`). The host was created seconds earlier with a per-session key and is reached over the AWS private network, so there is no prior key to pin and no third party in the path. The keys are discarded with the session. + +**Learner SSH is open to the internet.** That is the feature. The grader's path is not open; it is locked to the web service's security group. + +## What this is not + +OpsCoach is training infrastructure, not a hostile multi-tenant sandbox. It does not stop a learner from using their own host's outbound network during the session, and it bounds spend with a per-user session cap rather than hard cost controls. Both are reasonable given the audience and the one-hour blast radius. + +Teardown is the load-bearing control, so it has three independent layers: an SSH-idle watcher, a one-shot timer, and a periodic sweep. The full design is in **[lab-lifecycle-design.md](lab-lifecycle-design.md)**. diff --git a/docs-draft-v4/README.md b/docs-draft-v4/README.md new file mode 100644 index 0000000..f7061d3 --- /dev/null +++ b/docs-draft-v4/README.md @@ -0,0 +1,59 @@ +# OpsCoach + +Learn Linux and cloud operations on a real, throwaway server instead of a quiz. You open a lab, get your own Linux host within a minute or two, fix what is broken from a terminal in the browser, and a grader SSHes into the live machine to check whether you actually fixed it. + +Three pieces carry it: a browser-to-SSH terminal, a dedicated lab host per session on AWS, and grading that inspects real system state instead of comparing answers. + +The catalog ships with **Linux Foundations** and **AWS Foundations** drills plus **Beaconkeeper**, a capstone that drops you onto a broken Ubuntu host with twenty faults to diagnose and repair on one box before time runs out. + +## How it works + +A request reaches a shared load balancer, authenticates through Cognito, and lands on the OpsCoach web app running on ECS Fargate. Starting a lab launches the session's EC2 host; the web app bridges the browser terminal to it over SSH, and runs the grader over SSH. Grading streams pass/fail checks to a dashboard beside the terminal, so progress is visible as the box changes. + +For the whole system, **[architecture.md](architecture.md)** is the walkthrough: the numbered diagram, the provision, terminal, grade, and teardown flows, and the design forks behind them. Where the stakes are highest, two docs go deeper: + +- **[security.md](security.md)**: how untrusted users get root without putting anything else at risk. +- **[lab-lifecycle-design.md](lab-lifecycle-design.md)**: how a host is provisioned and the three independent paths that guarantee it dies. + +To run it on your laptop, jump to [Run it locally](#run-it-locally). + +## Why it is built this way + +The security model starts from a concession: the lab container has to run privileged, so assume it will be escaped and make escaping it worthless. A learner fixing systemd or filesystem state needs genuine root, and no container flag makes that fully safe. So the boundary that matters is the host under the container, not the container itself. Each host is single-tenant, so an escape reaches the learner's own machine and no other; each carries nothing worth stealing, since its credentials only pull a lab image and write logs; and each dies on a timer. The worst an escape can do is take full control of one throwaway box, holding no useful credentials, for at most an hour. The deeper model is in **[security.md](security.md)**. + +That only holds if the host is real. A shared host or a browser-only container would make container isolation the lone wall between a hostile learner and everyone else, so one escape would compromise every session at once. A per-session host spends a visible cost instead, a minute or two of provisioning and a per-session bill, to buy the isolation the security model rests on. + +Grading made its own demand. A fixed answer key tells you a learner typed the expected command, not that the box ended up right. So the grader SSHes into the live host with a least-privilege environment and inspects the machine itself: services up, files in place, config correct. A check passes only when the box is genuinely in the right state. The cost lands in authoring: every pack ships grader code that runs against a live machine. A check can fail because the system changed under it, not only because the learner did. + +The system owns no platform of its own, on purpose. It does not create a VPC, load balancer, or Cognito pool; it imports an existing shared set by ID, the way a service joins an organization that already runs those. The real IDs stay in local config, out of the repo. That is the deployment a real environment forces, and building to it from the start is the point. + +## Repository layout + +| Path | What it is | +| --- | --- | +| `web/` | Next.js app and custom Node server (the WebSocket-to-SSH bridge), API routes, UI | +| `infra/` | AWS CDK app: Fargate service, per-session lab hosts, teardown automation | +| `ContentPacks/` | Labs and graders, including the Beaconkeeper capstone | +| `scripts/` | Build, deploy, and smoke-test scripts | +| `docs/` | Architecture, security, and design notes | + +## Run it locally + +```bash +cd web +npm install +cp .env.example .env # most values are optional locally +npm run dev +``` + +That gets you the catalog and the play flow offline. With no `EC2_LAUNCH_TEMPLATE_ID` set, the app runs in mock mode: provisioning is faked and a session reports ready immediately. With no `DATABASE_URL`, sessions live in an in-memory store. Run the tests with `npm test`. + +One gap is left on purpose: mock mode does not start a lab container, so a live terminal and real grading need a container listening locally. The smoke script can bring one up. What is wired and what is still stubbed is in **[local-dev-without-aws.md](local-dev-without-aws.md)**. + +## Configure and deploy + +Deployment needs two files that stay out of version control (both are in `.gitignore`): the app's `web/.env`, copied from `web/.env.example`, and the CDK platform context `infra/cdk.context.json`, copied from `infra/cdk.context.example.json` or generated by `scripts/discover-platform-context.sh`. With those in place, `scripts/deploy-platform.sh` builds the images and deploys the stacks. The step-by-step walkthrough is in **[../infra/PLATFORM_INTEGRATION.md](../infra/PLATFORM_INTEGRATION.md)**. + +## License + +MIT. See [../LICENSE](../LICENSE). diff --git a/docs-draft-v4/architecture.md b/docs-draft-v4/architecture.md new file mode 100644 index 0000000..0a62a11 --- /dev/null +++ b/docs-draft-v4/architecture.md @@ -0,0 +1,171 @@ +# OpsCoach architecture + +OpsCoach hands each learner a real, throwaway Linux host and grades what they actually do to it. This doc descends from 30,000 feet to the design decisions behind it. + +**Scope.** Single-region training infrastructure for one app that borrows a shared ALB, Cognito, and VPC platform instead of standing up its own. Explicitly not a hostile multi-tenant sandbox, and not a high-availability production service. + +## 30,000 ft · the shape + +Three parts, and everything else is how they talk to each other safely: + +- A **web app** (Next.js on ECS Fargate) that serves the UI, bridges the browser terminal to a shell, and runs grading. +- A **per-session lab host** (a dedicated EC2 instance) that the learner operates and that is destroyed on a timer. +- A **shared platform** (ALB, Cognito, VPC) that the app plugs into by ID rather than rebuilding. + +The engineering lives in the seams: handing an untrusted user root on a real machine without putting anything else at risk, holding a live shell open through a load balancer, and guaranteeing that no host outlives its session. + +## 10,000 ft · the system + +![OpsCoach AWS architecture](architecture.svg) + +The numbered arrows trace one session from first request to teardown. + +1. **Request.** Browser to the ALB over HTTPS. The ALB runs `authenticate-cognito` (Cognito hosted UI, Google), and a post-auth passphrase gate inside the app must also pass before any page renders. Only the health check and logout skip Cognito. +2. **Route.** The ALB forwards to the OpsCoach web service on Fargate, in a private, egress-only subnet. +3. **Provision and grade.** Fargate launches the host with `RunInstances`, then later runs the grader against it over SSH. Both ride one control link (a single arrow in the diagram). The learner's own terminal is a separate path (step 5). +4. **Ready webhook.** The host calls back to the service once it is up, resolved through Cloud Map and authenticated with a shared secret. +5. **Terminal.** The learner operates the box here: the browser streams over a WebSocket to a custom Node server, which relays to the host's shell over SSH. +6. **Teardown.** Three independent, idempotent paths guarantee the host dies (the numbered arrow is the terminate). Any one suffices. + +Grey dashed lines are supporting paths: Fargate to RDS PostgreSQL in the isolated subnet, the host pulling its lab image from ECR, and the opt-in direct SSH from a learner's laptop. + +### Components + +| Component | AWS service | Role | +| --- | --- | --- | +| Web app + terminal bridge | ECS Fargate | Serves UI; WebSocket-to-SSH PTY bridge; session lifecycle + grading | +| Edge auth | ALB + Cognito | `authenticate-cognito` at the load balancer (hosted UI to Google), then a shared-passphrase gate in the app | +| Lab host | EC2 (per session) | Ephemeral AL2023 / arm64 host running the lab container; learner's SSH target | +| Container images | ECR | Images for the web service and each lab | +| Database | RDS PostgreSQL | Sessions, check runs, grader results (isolated subnet) | +| Service discovery | Cloud Map | In-VPC address for lab-host-to-web callbacks | +| Teardown | On-host watcher + EventBridge Scheduler + Lambda | SSH-idle shutdown, one-shot max-lifetime schedule, expiry-tag sweep | +| Secrets | Secrets Manager | Database credentials and the callback HMAC secret | + +
+The flows, step by step (auth, provision, terminal, grading, teardown) + +  + +**Authentication and access gate** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant ALB as ALB + participant Cog as Cognito + participant App as Fargate + U->>ALB: HTTPS request + ALB->>Cog: authenticate-cognito (if enabled) + Cog-->>ALB: OIDC tokens + ALB->>App: forward + signed identity header + alt no gate cookie + App-->>U: redirect to /gate + U->>App: POST /api/gate (passphrase) + App-->>U: set gate cookie, continue + end + App-->>U: app page + Note over ALB,App: /api/health and /logout skip Cognito +``` + +**Provision a lab session** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant App as Fargate + participant EC2 as Lab host + participant ECR as ECR + participant Sch as Scheduler + U->>App: POST /api/sessions (start lab) + App->>App: create session in RDS; mint keys + callback token + App->>EC2: RunInstances (Launch Template + user-data) + App->>Sch: CreateSchedule (terminate at T + maxLifetime) + Note over EC2: user-data: install Docker, block IMDS,
ECR login, run lab container, set authorized_keys + EC2->>ECR: pull lab image + EC2->>App: ready webhook (via Cloud Map + shared secret) + App-->>U: session ready (host, port) +``` + +**Browser terminal and SSH** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant Srv as Node server + participant EC2 as Lab host + U->>Srv: WebSocket /api/sessions/:id/shell?token + Srv->>Srv: shell-auth (internal secret) resolves host + key + Srv->>EC2: ssh2 PTY on port 22 (per-session key) + EC2-->>Srv: stdout / stderr + Srv-->>U: stream to xterm.js + Note over Srv,EC2: 25s keepalive ping holds the ALB idle timer + Note over U,EC2: opt-in: SSH directly to the host's public IP +``` + +**Live grading** + +```mermaid +sequenceDiagram + autonumber + actor U as Browser + participant App as Fargate + participant G as Grader + participant EC2 as Lab host + U->>App: POST /api/sessions/:id/grade + App->>G: spawn grader (allowlisted env, no task-role creds) + G->>EC2: SSH port 22 (grader key) runs ops status + EC2-->>G: JSON check results + G-->>App: { passed, checks[] } + App->>App: persist check run in RDS + App-->>U: live results on the dashboard +``` + +**Idle and lifetime teardown** + +```mermaid +sequenceDiagram + autonumber + participant EC2 as Lab host + participant App as Fargate + participant Sch as Scheduler + participant L as Terminator Lambda + Note over EC2: path 1 (prompt): on-host watcher polls :22 + EC2->>App: no SSH for grace period: shutdown (reason=ssh_idle) + App->>EC2: TerminateInstances; delete schedule + Note over Sch,L: path 2 (hard cap): if still running at T + maxLifetime + Sch->>L: fire one-shot + L->>EC2: TerminateInstances + L->>App: shutdown callback (mark session stopped) + Note over EC2,App: path 3 (backstop): 5-min EventBridge sweep
terminates any host past its ExpiresAt tag +``` + +
+ +## 1,000 ft · the decisions that shaped it + +Each of these is a place the obvious choice would have been wrong. Constraint, decision, trade-off. + +**A real host per session, not a shared sandbox.** +Constraint: teaching operations means real systemd, real root, real packages, and you cannot hand untrusted users root on shared infrastructure. Both cheaper alternatives (a shared multi-tenant host, or browser-only containers) make container isolation the only wall between a hostile learner and everyone else, so a single escape compromises every session. Decision: every session gets its own ephemeral EC2 host, killed on a timer. The security model treats that host, not the container on it, as the real boundary. Trade-off: a minute or two of provisioning latency and per-session cost, in exchange for a blast radius of exactly one throwaway box, and a learner operating a real machine rather than a simulation. Full model in [security.md](security.md). + +**A custom Node server for the browser terminal.** +Constraint: a browser terminal needs a long-lived, two-way connection to a shell, and Next.js on its own does not hold one. Decision: wrap Next in a thin `server.js`. It leaves every normal request to the stock Next handler and intercepts only the WebSocket upgrade, where it authenticates the session through an internal call and bridges to the host with an `ssh2` PTY. Trade-off: a custom server instead of stock Next, plus a 25-second keepalive ping so the ALB does not reap an idle terminal, in exchange for a real shell in the browser with nothing for the learner to install. + +**Grade real state, not answers.** +Constraint: a fixed answer key tells you the learner typed the expected command, not that the box ended up in the right state. Decision: the grader SSHes into the live host with a least-privilege environment and checks real state (services up, files in place, config correct), returning structured results that stream to the dashboard as each check completes. Trade-off: graders are per-pack code that has to run against a live machine, in exchange for a pass that means the box is genuinely in the right state, not that the learner ran the expected commands. + +**Three independent teardown paths.** +Constraint: a leaked instance costs real money, and the control plane never sees the learner's SSH activity, so it cannot infer idle from browser traffic. Decision: an on-host SSH-idle watcher, a one-shot EventBridge Scheduler set at provision time, and a 5-minute sweep over expiry tags. Any one path is sufficient, and all are idempotent. Trade-off: more moving parts, for a hard guarantee that nothing runs forever. Full design in [lab-lifecycle-design.md](lab-lifecycle-design.md). + +**Borrow the platform; import by ID.** +Constraint: a demo should not stand up its own ALB, Cognito, and VPC. Decision: the CDK imports a shared platform's resources by ID from local context, and the real IDs stay out of the repo (placeholders ship in their place). Trade-off: the app cannot bootstrap its own world from nothing, in exchange for dropping cleanly into an existing shared environment. + +## See also + +- **[security.md](security.md)** for the security model and its trade-offs. +- **[lab-lifecycle-design.md](lab-lifecycle-design.md)** for provisioning and the three-layer teardown in depth. +- **[../infra/PLATFORM_INTEGRATION.md](../infra/PLATFORM_INTEGRATION.md)** for plugging into a shared ALB/Cognito platform. diff --git a/docs-draft-v4/architecture.svg b/docs-draft-v4/architecture.svg new file mode 100644 index 0000000..4ec1f3d --- /dev/null +++ b/docs-draft-v4/architecture.svg @@ -0,0 +1,149 @@ + + OpsCoach AWS architecture + A user reaches a shared Application Load Balancer with Cognito auth and a passphrase gate, which forwards to the OpsCoach web service on ECS Fargate in a private subnet. Fargate uses an isolated-subnet RDS PostgreSQL database, launches per-session EC2 lab hosts in a public subnet (which pull a lab container from ECR and call back via Cloud Map), bridges the browser terminal to the host over SSH, runs a grader over SSH, and schedules automatic teardown via EventBridge Scheduler and a terminator Lambda. + + + + + + + + + + + + + OpsCoach — AWS architecture + In-browser & SSH Linux labs · shared ALB/Cognito platform · per-session EC2 hosts · live grading · auto-teardown + + + USER + + Browser + xterm.js terminal + + SSH client + direct opt-in + + + + + AWS Account · Region (us-east-1) + + + + Shared platform — imported by the app (not created here) + + Application Load Balancer + Cognito auth · routing + + Amazon Cognito + hosted UI · Google IdP + + + + VPC + + + + Public subnet + + EC2 — lab host (per session) + AL2023 · arm64 · Launch Template + runs lab container · IMDS blocked + + + + Private subnet (egress) + + ECS Fargate — OpsCoach web + Next.js + custom Node server + WebSocket → SSH PTY bridge + + AWS Cloud Map + in-VPC discovery + + + + Isolated subnet + + Amazon RDS — PostgreSQL + sessions · check runs · results + + + + Automation & services + + Amazon ECR + web + lab images + + Secrets Manager + DB · callback · keys + + EventBridge + Scheduler + one-shot per session + + AWS Lambda + session terminator + + + + + 1 + + + + 2 + + + + 3 + + provision · SSH PTY · grade + + + + 4 + + + ready webhook + + + + 5 + + WSS + + + + SSH opt-in + + + + + 6 + + terminate EC2 + + + + SQL + + + + pull lab image + + + diff --git a/docs-draft-v4/lab-lifecycle-design.md b/docs-draft-v4/lab-lifecycle-design.md new file mode 100644 index 0000000..94b1405 --- /dev/null +++ b/docs-draft-v4/lab-lifecycle-design.md @@ -0,0 +1,164 @@ +# Lab host lifecycle and teardown + +Status: implemented (web app + CDK). + +Every learner session gets its own EC2 lab host, and the load-bearing problem is killing it again. This doc explains how OpsCoach provisions and tears down per-session hosts, and why teardown runs as three independent paths instead of one. For the wider security argument that teardown protects, see [security.md](security.md); for where this sits in the system, see [architecture.md](architecture.md). + +## Problem + +A session can launch a dedicated `t4g.micro` host (Amazon Linux 2023, arm64) running the lab container under Docker. If teardown fails or never runs, hosts leak and bill by the hour. Teardown is therefore the control that has to work. + +The trap is assuming the control plane knows when a learner is done. It does not. The web app embeds a browser terminal: xterm.js streams over a WebSocket to a custom Node server (`web/server.js`), which bridges to the host with an `ssh2` PTY on port 22. That bridge holds a connection open with a 25-second keepalive so the load balancer does not cut an idle terminal, so its liveness tracks the socket, not the human. + +The learner can also opt into SSH straight from their laptop to the host's public IP, and those sessions never touch the control plane at all. Either way, the Next.js control plane on Fargate cannot reliably infer idle from a browser tab, and it is blind to direct SSH. Whatever decides "this host is idle" has to live on the host. + +So teardown has to be: + +- **Prompt** once the learner is actually gone (no live SSH on the box). +- **Reliable** when a callback is dropped, a schedule misfires, or the learner never connects at all. +- **Idempotent**, because more than one path can fire on the same host within seconds. + +**Non-goals.** This is not a cost-optimization or autoscaling design: there is no bin-packing, no spot fallback, no warm pool. It does not extend a session on activity (a started host dies at its cap regardless of use; see Future improvements). And it is not the security model. It is the mechanism that guarantees the security model's one-hour blast radius actually expires. + +## Defense in depth: three teardown paths + +Teardown runs as three independent paths. Any one of them succeeding is enough, and every one is safe to run more than once. The point is that no single mechanism is trusted: the host can wedge, a callback can drop, schedule creation can fail at provision time. Layering trades extra moving parts for a hard guarantee that nothing runs forever. + +| Path | Fires when | Runs on | Typical latency | +|------|------------|---------|-----------------| +| 1. SSH idle watcher | No established TCP on host `:22` for the grace period, after at least one session was seen | Background script in host user-data | ~2 min after the last disconnect | +| 2. Max-TTL schedule | A one-shot EventBridge Scheduler entry created at provision | Terminator Lambda | At T + max lifetime, exactly | +| 3. ExpiresAt sweep | An `ExpiresAt` instance tag is in the past | Same Lambda, every 5 min | Up to 5 min after the tag expires | + +The three "Why not X alone?" sections below justify the layering, one path at a time. First, the learner's **Stop lab** button and the authenticated `POST /api/sessions/:id/stop` are not a fourth path: they call the same internal shutdown routine the automated paths converge on (see Unified shutdown). + +### Why not the idle watcher alone? + +Because it runs on the host, and the host can misreport or go silent. The watcher fails to fire in exactly the cases that cost the most: the learner provisions a host and never connects (so the watcher never sees a session to start its idle clock), a bug stalls the loop, or the host wedges badly enough that nothing on it runs. The watcher is the prompt path, not the guaranteed one, so it needs a backstop that runs off-host. + +### Why not a timer alone? + +A timer with no idle signal is either too aggressive or too loose. Tune it short and it kills learners mid-lab; tune it long and idle hosts bill for the slack. The earlier design made exactly this mistake: it set `ExpiresAt = now + 10 min` at provision and called it idle teardown, which just terminated active sessions ten minutes in. A fixed cap is the right tool for *bounding* cost, not for detecting idle. It belongs as the backstop, with a real idle signal in front of it. + +### Why both a schedule and a sweep? + +They cover different failures. The schedule is precise but can fail to exist: if `CreateSchedule` fails at provision (missing IAM, a transient API error), there is no timer for that host, and a host with no timer is exactly the host that leaks. The sweep needs nothing but a tag that `RunInstances` already wrote, so it catches hosts the schedule missed, plus any that outlived their schedule through API errors. One path is precise but can fail to exist, the other is blunt but certain, and a single Lambda runs both. + +## Path 1: SSH idle watcher + +Generated as shell user-data in `web/lib/lab-user-data.ts` (mirrored in `infra/lib/lab-user-data.sh` for launch-template defaults). A background loop runs on the host: + +1. Every 15 seconds, count established connections on local port 22 with `ss -tn state established '( sport = :22 )'`. +2. Once that count goes above zero, latch a `had_session` flag. The watcher will not act until it has seen at least one real session. +3. When `had_session` is set and the count falls back to zero, start an idle clock. +4. After `SSH_IDLE_GRACE_SECONDS` (default **120**) of continuous idle, POST the shutdown webhook: + +```http +POST /api/sessions/:id/shutdown +X-Internal-Secret: +Content-Type: application/json + +{ "reason": "ssh_idle" } +``` + +The 120-second grace debounces two things: a brief network blip or `ssh` reconnect that should not end the session, and the grader's SSH from Fargate (the platform grades a lab by logging in to check real state) that should be allowed to finish without racing a learner disconnect. + +The watcher only *asks* for teardown. It cannot terminate anything: the lab instance role grants log writes and ECR image pulls and nothing else (no `ec2:TerminateInstances`), so a compromised host cannot turn the watcher into a weapon against other instances. Termination always runs off-host, gated by the callback secret. + +Two limits are deliberate for v1: + +- **The watcher counts every connection on `:22`, including grader SSH from inside the VPC.** A learner who never opens a terminal but runs checks repeatedly from the web UI can see their host torn down soon after grading goes quiet. Acceptable now; the fix is source-aware counting (below). +- **If the learner never SSHes, `had_session` stays unset and the watcher never fires.** That host is precisely what the max-TTL path exists to catch. + +## Path 2: Max-TTL schedule + +`web/lib/session-scheduler.ts`, called from `web/lib/ec2-labs.ts` right after `RunInstances`. On a successful provision, Fargate creates a one-shot EventBridge Scheduler entry named `opscoach-{sessionId}` (truncated to 64 characters): + +- Expression `at()`, set to **T + maxLifetimeMinutes** from provision. +- Target: the terminator Lambda, with `{ "action": "terminate", "instanceId": "...", "sessionId": "...", "reason": "max_ttl" }`. +- `ActionAfterCompletion: DELETE`, so the entry cleans itself up after it fires. +- Any earlier shutdown (`manual`, `ssh_idle`) calls `DeleteSchedule`, so a host that dies early does not leave a dangling timer. + +**Default max lifetime: 60 minutes** (`OPSCOACH_MAX_LIFETIME_MINUTES`, CDK context `maxLifetimeMinutes`). Sixty minutes is long enough for a full lab plus assessment retries and short enough to bound the bill if every other path fails. It is intentionally orthogonal to the idle grace: one bounds the worst case, the other handles the normal case. + +**Why EventBridge Scheduler, not EventBridge Rules?** Scheduler does one-shot schedules natively, with per-session names and auto-delete after firing. Rules are built for recurring patterns, which is why the 5-minute sweep (path 3) uses a Rule and the per-session cap does not. + +## Path 3: ExpiresAt sweep + +`infra/lib/session-terminator/handler.py`, triggered every 5 minutes by an EventBridge Rule defined in `infra/lib/lab-host-stack.ts`. At provision, `RunInstances` tags the host `ExpiresAt=` on the same horizon as the schedule. On each run, the Lambda lists running OpsCoach hosts (tag `OpsCoach=true`) and, for any whose `ExpiresAt` is in the past, terminates the instance and calls the shutdown API with `reason=expires_at_sweep`. + +The sweep depends on nothing but a tag the launch already wrote, which is the whole point: it is the safety net for the rare host whose schedule was never created or somehow outlived itself. + +## Unified shutdown + +Every path, automated or manual, converges on `shutdownSessionInternal` in `web/lib/sessions.ts`. Keeping one routine is what makes the idempotency real: concurrent triggers collapse onto the same guarded transition instead of racing. + +1. Return immediately if the session is already `stopped` or `stopping`. +2. Set status `stopping`. +3. `DeleteSchedule` (best effort). +4. `TerminateInstances` if an instance id is present. Errors are logged, and the session is still marked stopped, so an API hiccup never strands a session in limbo. +5. Set status `stopped` and publish the event the dashboard streams over SSE. + +There are two ways in, both reaching the same routine: + +| Entry point | Auth | Reasons | +|-------------|------|---------| +| `POST /api/sessions/:id/stop` | Session token | `manual` | +| `POST /api/sessions/:id/shutdown` | `X-Internal-Secret` | `ssh_idle`, `max_ttl`, `expires_at_sweep`, `manual` | + +The terminator Lambda terminates EC2 first, then calls the shutdown API, so the EC2 state of the world and the Postgres record converge rather than drift. The shutdown route rejects an unsigned call with 401. + +## Alternatives considered + +| Approach | Rejected because | +|----------|------------------| +| Provision-time timer as "idle" (`ExpiresAt = now + 10 min`) | Kills active sessions; a fixed timer is not idle detection. This is the bug the current design replaced. | +| Web-only activity timeout | The control plane is blind to SSH. A learner working in the terminal looks idle if the browser tab is quiet, and direct SSH is invisible entirely. | +| Tailscale-only SSH (no public IP) | A product decision for v1: hardened public SSH only. The watcher and timers are written so this can change without reworking teardown. | +| A dedicated Lambda per session | One shared terminator Lambda plus a per-session schedule is simpler to operate and cheaper than N functions. | +| Host self-terminates via its own IAM | Widens the blast radius of a compromised host. Termination stays off-host, behind the callback secret. | + +## Configuration + +The two knobs that change behavior are the max lifetime and the idle grace. Both are set from the Fargate task environment and plumbed from CDK context; the rest of the environment wires the actors together. + +| Variable | Default | Purpose | +|----------|---------|---------| +| `OPSCOACH_MAX_LIFETIME_MINUTES` | `60` | Schedule fire time and `ExpiresAt` horizon | +| `OPSCOACH_SSH_IDLE_GRACE_SECONDS` | `120` | Host idle debounce before the shutdown webhook | +| `SESSION_TERMINATOR_LAMBDA_ARN` | (from CDK) | Schedule target | +| `SCHEDULER_INVOKE_ROLE_ARN` | (from CDK) | Scheduler execution role | +| `INTERNAL_CALLBACK_SECRET` | (Secrets Manager) | Authenticates host and Lambda callbacks | + +One CDK field is a deliberate trap to avoid: `idleTimeoutMinutes` (default `10`) is a legacy name from the old timer-as-idle design and **no longer drives `ExpiresAt`**. The horizon comes from `maxLifetimeMinutes`. The name is kept only to avoid a churning rename; do not wire teardown to it. + +Without `EC2_LAUNCH_TEMPLATE_ID`, provisioning is mock-only: no real host, no scheduler, no watcher, and sessions live in memory unless `DATABASE_URL` is set. That is the local-development path, covered in [local-dev-without-aws.md](local-dev-without-aws.md). + +## CDK components + +| Resource | Stack | Role | +|----------|-------|------| +| Terminator Lambda | Lab host stack | Direct terminate, the 5-minute sweep, and the shutdown-API notify | +| Scheduler-invoke IAM role | Lab host stack | Lets EventBridge Scheduler invoke the Lambda | +| EventBridge Rule (5 min) | Lab host stack | Sweep trigger | +| Callback secret | Lab host stack | Shared HMAC secret, read by the Fargate task | +| Scheduler IAM on the task role | Web stack | `CreateSchedule` / `DeleteSchedule` | + +For how these stacks plug into a borrowed ALB/Cognito/VPC platform, see [../infra/PLATFORM_INTEGRATION.md](../infra/PLATFORM_INTEGRATION.md). + +## Future improvements + +- **Source-aware idle detection.** Count only SSH from non-VPC (learner) addresses, so web-only grading stops arming the watcher and the first known limitation goes away. +- **Activity extension.** Refresh `ExpiresAt` and reschedule the max-TTL entry on a grader run or an explicit heartbeat, trading some cost for longer working sessions. +- **Teardown metrics.** CloudWatch counters per reason (`ssh_idle`, `max_ttl`, `expires_at_sweep`, `manual`) to tune the grace and the cap against real usage instead of guesses. +- **Pre-baked AMI.** A Packer image with Docker, the watcher, and host hardening already installed, replacing the full user-data bootstrap and shaving provision time. + +## Related files + +- `web/lib/ec2-labs.ts`: provision, tag, kick off the schedule +- `web/lib/session-scheduler.ts`: EventBridge Scheduler client +- `web/lib/lab-user-data.ts`: host bootstrap and the idle watcher +- `web/app/api/sessions/[id]/shutdown/route.ts`: internal shutdown API +- `web/lib/sessions.ts`: `shutdownSessionInternal`, the unified path +- `infra/lib/lab-host-stack.ts`: terminator Lambda, sweep Rule, scheduler role +- `infra/lib/session-terminator/handler.py`: terminate and sweep logic diff --git a/docs-draft-v4/local-dev-without-aws.md b/docs-draft-v4/local-dev-without-aws.md new file mode 100644 index 0000000..3d6aac1 --- /dev/null +++ b/docs-draft-v4/local-dev-without-aws.md @@ -0,0 +1,47 @@ +# Local development without AWS + +Run the OpsCoach web app on a laptop with no AWS account: clone the repo, `npm run dev`, and the full UI plus grader loop come up in minutes. This doc is the design for closing the last manual gap in that loop. It is a deferred proposal, not a how-to: most of what is below does not exist yet, and the table marks what works today from what is still planned. + +## Scope + +The goal is the local equivalent of the production lifecycle (provision, operate, grade, tear down) for the SSH-based content packs, close enough that a contributor never needs AWS to build and test a feature. Non-goals: reproducing the AWS control plane locally (Cognito, EventBridge Scheduler, the terminator Lambda), and the AWS-credential labs, whose offline story is deferred under Alternatives. + +## What works today, and the one gap + +Most of the loop already runs without AWS: + +| Piece | Local behavior | +|-------|----------------| +| Web UI | `cd web && npm run dev` serves the catalog, play flow, session page, and SSE grading | +| Session store | In-memory `Map` when `DATABASE_URL` is unset; point it at a local Postgres to persist | +| Provisioning | Mock EC2 when `EC2_LAUNCH_TEMPLATE_ID` is unset (`isMockEc2Mode()` in `../web/lib/ec2-labs.ts`): a session goes `ready` at `127.0.0.1:22` | +| Graders | The same ContentPack scripts as production, run as a subprocess that SSHes from the API process | +| Smoke test | `../scripts/smoke-web-session.sh`; with `OPSCOACH_SMOKE_START_COMPOSE=1` it brings up a foundations lab on `127.0.0.1:22` | + +The gap is a single missing piece. Mock mode returns `127.0.0.1:22` but starts nothing listening there. In production a host's user-data brings the lab container up; locally, nothing does, so today you supply the container yourself, by hand or through the smoke script's helper. Closing that seam is the proposal. + +(One smaller dev-loop snag lives here too: the in-memory store is wiped on every `next dev` hot reload, after which `/grade` fails with "Invalid session or token." Run `npm start` on a built bundle, or point at a local Postgres, to keep sessions across reloads.) + +## One interface, two backends + +**Constraint:** the product has to provision a lab two completely different ways, a fresh EC2 host in production and a local container in development, without forking the session lifecycle, the grader, or the API routes. + +**Decision:** keep one orchestration interface, `provisionLabInstance()` and `terminateLabInstance()`, and vary only what sits behind it. Production runs Docker on an EC2 host it boots through user-data; local mode runs Docker on the laptop. Everything upstream (the API routes, the SSH grader, the ContentPacks) is identical, because the difference is confined to that one pair of functions. It is the same boundary that already lets mock EC2 stand in for a real launch; local mode extends it from "return a fake host" to "start a real local one." This is also the bulk of what the native macOS app's `ContainerManager` did (a per-session `docker compose` project, a mapped SSH port, injected keys, teardown on stop), now rebuilt behind the shared interface. + +One flag drives it. `web/.env.local` sets `OPSCOACH_LOCAL_DEV=1`, leaves `EC2_LAUNCH_TEMPLATE_ID` unset, and points `CONTENT_ROOT` and `SESSIONS_ROOT` at local paths. With the flag on, `provisionLabInstance()` runs `docker compose` from the lab's `runtime.directory`, maps a free host port, wires `sshHost` and `graderHost` and injects the learner and per-session grader keys, and tears down with `docker compose down` on stop. It skips the AWS-only paths: STS and CloudFormation prep, the Scheduler one-shot, the EC2 terminate calls. An optional `scripts/dev.sh` would check Docker, export the env, and start Next.js, so the loop is one command. + +**Trade-off:** the interface has to stay narrow enough that both backends honor it, which keeps EC2-specific and Docker-specific details out of the callers. That narrowness is what makes the whole system testable without AWS. The work is worth doing once local iteration, not deploy, is what slows contributors down. + +The one piece this design cannot fully scope in advance is Beaconkeeper: its image and seed (`runtime.directory` and `defaultSeed` from the manifest) need a real run to pin down. + +## Alternatives considered + +A full local AWS emulation (LocalStack or similar) was rejected: it rebuilds the control plane this design treats as a non-goal, and trades a little manual setup for a large, drifting dependency that still would not run real labs. Driving everything through the smoke script's compose helper was rejected as the primary path because it is a test harness, not a dev loop: it starts one fixed lab on a fixed port rather than tracking the lab a learner picked. It stays useful as a fast end-to-end check. + +For the AWS-credential labs, the choice is deferred: either mint fixture credentials so they run offline, or have local mode return an explicit "this lab requires the platform" error. The error path is cheaper; fixtures are more complete. Whoever picks this up decides based on whether offline AWS labs justify the fixture upkeep. + +## References + +- Mock EC2 and the orchestration interface: `../web/lib/ec2-labs.ts` (`isMockEc2Mode()`, `provisionLabInstance()`, `terminateLabInstance()`) +- AWS lab prep that local mode skips: `../web/lib/aws-lab-manager.ts` (`prepareAwsSession()`) +- Production deploy and the shared environment: `../infra/PLATFORM_INTEGRATION.md` diff --git a/docs-draft-v4/security.md b/docs-draft-v4/security.md new file mode 100644 index 0000000..f493960 --- /dev/null +++ b/docs-draft-v4/security.md @@ -0,0 +1,59 @@ +# Security model + +OpsCoach hands authenticated users a real Linux machine and lets them run real commands as a privileged user. That is the product, and it is also the security problem. The whole design is about making that machine cheap to lose. + +## The core bet + +Assume the lab container can be escaped, and build so that escaping it does not matter. + +A learner doing a systemd or filesystem lab needs genuine control of the box, so the lab container runs `--privileged`. Treating that container as a strong wall would be a losing game. Instead OpsCoach treats the EC2 host as the real boundary and makes the host cheap to lose: + +- One host per session: no shared tenancy, so a learner can reach only their own machine. +- Nothing valuable lives on it: the host's IAM role can pull the lab image and write logs, and that is the whole list. Database credentials, the callback secret, and grader keys never touch the host. +- It dies on a short clock: teardown is driven from outside the host, so a wedged or hostile box still gets killed. + +Worst case for an attacker who escapes the container: full control of one throwaway machine, holding no useful credentials, for at most an hour. + +The rest of this document is the controls that make each of those claims true, from the edge inward. + +## Two gates, and why they are not the security model + +Reaching any page takes two independent checks. The shared Application Load Balancer authenticates every request through Cognito (hosted UI, Google as the identity provider). Behind it, the app requires a shared passphrase before it renders anything (`web/lib/gate.ts`); the accepted codes and cookie token come from environment variables, so no secret lives in source. Only the load balancer health check and the logout route skip Cognito. + +These gates keep the public out, but the security model does not rest on them. A shared passphrase is exactly the kind of secret that leaks, so the design never treats it as a wall. If both gates fell, an attacker would reach a session and a privileged host, which is the case the rest of this document is built to survive. Past the gates, every session also carries a short-lived bearer token checked against a hash in the database, so a guessed or stale URL gets a caller nowhere. Each user is capped at three concurrent sessions by default. + +## The host is built to be worthless to steal + +Once a lab host is running, the design assumes it may be fully compromised and removes anything worth taking. + +Its IAM role is read-only and narrow: pull `your-org/opscoach-lab*` images from the registry, write to the lab log group, and nothing else. That role is then kept out of the container's reach. + +The control that does the work is IMDSv2 with a one-hop response limit. A metadata request from inside the container crosses the Docker bridge NAT on its way out, and that extra hop spends the single allowance, so the request expires before it reaches the metadata service. A request from the host is still at hop one and succeeds. That is why the host can read its own credentials while the container cannot. A second layer backs it up: a `DOCKER-USER` iptables rule drops container traffic to the metadata address (`169.254.169.254`) outright. A host-root escape can delete that iptables rule, but the one-hop limit still stands, and behind both the role is nearly empty and the host is one isolated, short-lived machine. There is little to take and nowhere to pivot. + +The network path is one-way by design. The host's security group accepts learner SSH from the internet, because connecting from your own client is the feature. It accepts the grader's SSH only from the web service's security group, a single source rather than a VPC-wide rule on port 22, so no lab host can reach another. The database sits in an isolated subnet behind the private web service, unreachable from any lab host. + +## The valuable secrets live somewhere else + +The pieces that hold real credentials stay off the host entirely, so compromising a host never yields them. + +Database credentials and the callback HMAC secret live in Secrets Manager, read at runtime by the web task and the terminator Lambda. They are never written to a host or baked into an image. + +Grading runs least-privilege. The grader is a committed content-pack program that the web task runs as a child process with an allowlisted environment: `PATH`, `HOME`, `LANG`, `AWS_REGION`, `AWS_CONFIG_FILE`, and a few locale variables. The AWS credential variables and the ECS role-fetch URI are excluded, so a buggy or hostile grader cannot read the database, the callback secret, or assume the task's role. Labs that exercise real cloud APIs do not borrow the task role either; they authenticate with their own scoped, per-session workspace credentials. + +Each session also mints two fresh SSH keypairs: the learner's (public key on the box, private key theirs) and the grader's (server-only). A leaked learner key therefore reaches one box the learner already controls, and nothing else, and both keys are discarded when the session ends. + +## Teardown is the load-bearing control + +Every control above bounds the blast radius; teardown bounds it in time. A host comes down when it goes idle or when it hits an hour old, whichever lands first, and the deadline is enforced from outside the host so a wedged or hostile box cannot dodge it. Three independent paths back that deadline. Any one is enough, and all are safe to run twice: + +- an SSH-idle watcher on the host that fires a couple of minutes after the learner disconnects, +- a one-time scheduled timer at the sixty-minute maximum lifetime, and +- a periodic sweep that catches anything the first two missed. + +The full design, including why a single timer or a single idle check was rejected, is in [lab-lifecycle-design.md](lab-lifecycle-design.md). For where these controls sit in the larger system, see [architecture.md](architecture.md). + +## Honest limits and trade-offs + +A few sharp edges are deliberate. The lab container runs privileged, which it must for systemd and realistic administration; that is acceptable only because the host around it is single-tenant, credential-poor, and short-lived. The terminal bridge skips SSH host-key verification (`hostVerifier: () => true` in `web/server.js`): the host was created seconds earlier with a throwaway per-session key, reached over the AWS private network, and discarded with the session, so there is nothing to pin and pinning would add ceremony, not safety. + +Two things are out of scope on purpose. OpsCoach does not restrict a learner's outbound network from their own host during the session; egress lockdown waits for the planned move to a dedicated lab account, where it can be tuned to what bootstrap actually needs rather than guessed at now. And it bounds spend with the per-user session cap rather than hard cost controls. Both are reasonable given the audience and the one-hour ceiling on any single host.