Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
60 changes: 60 additions & 0 deletions docs-draft-v2/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# OpsCoach

Learn Linux and AWS operations on a real, throwaway cloud server instead of a quiz. You get a dedicated Linux host and a task list, and a grader inspects the actual state of the machine as you work.

This is a portfolio project. The parts worth a look: a live browser-to-SSH terminal, per-session disposable lab hosts on AWS, and grading that reads real system state rather than comparing answers.

## What it does

Open a lab and you get a dedicated Linux host in a minute or two. You work in a terminal, in the browser or over your own SSH client, to fix broken services, harden config, triage logs, wire up systemd, and so on. A grader runs against the live machine and streams pass/fail checks to a dashboard beside the terminal. When you finish or go idle, the host is destroyed.

Three content packs ship today: **Linux Foundations** and **AWS Foundations** drills, and the **Beaconkeeper** capstone, a 20-step systemd and operations scenario on a single box.

## How it works

A request reaches a shared ALB, authenticates through Cognito, and lands on the OpsCoach web app on ECS Fargate. Starting a lab launches a dedicated EC2 host; the web app bridges the browser terminal to it over SSH and runs the grader over SSH. Every host self-destructs on an idle or lifetime timer. [architecture.md](architecture.md) has the diagram and the full provision, terminal, grade, and teardown flows.

## Why it is built this way

Four decisions shaped the system. Each is defended in full in [architecture.md](architecture.md); the short version:

- **Real hosts, not a fake shell.** Learning operations means touching real systemd, real packages, real logs, so each session is an actual EC2 instance rather than an emulation.
- **Disposable and single-tenant.** The lab runs as a privileged user, so each session gets its own host that is wiped on a timer. The security boundary is the cheap, isolated host, not container isolation. See [security.md](security.md).
- **Grade real state, not answers.** The grader SSHes in and inspects the machine, so a check passes only when the box is genuinely in the right state.
- **Borrow the platform, do not rebuild it.** The infrastructure plugs into an existing shared ALB, Cognito, and VPC by ID rather than standing up its own. The real resource IDs live in local config kept out of the repo.

## Repository layout

| Path | What it is |
| --- | --- |
| `web/` | Next.js app and custom Node server (WebSocket-to-SSH bridge), API routes, UI |
| `infra/` | AWS CDK app: Fargate service, lab hosts, teardown automation |
| `ContentPacks/` | Labs and graders, including the Beaconkeeper game |
| `scripts/` | Build, deploy, and smoke-test scripts |
| `docs/` | Architecture, security, and design notes |

## Run it locally

```bash
cd web
npm install
cp .env.example .env # most values are optional locally
npm run dev
```

With no AWS credentials and no `DATABASE_URL`, the app runs in mock mode: an in-memory store and faked provisioning, so you can click through the UI and content offline. Run the tests with `npm test`. Mock mode does not start a lab container for you, so a few flows still need Docker running locally; the gaps and the proposed local-Docker mode are in [local-dev-without-aws.md](local-dev-without-aws.md).

## Configure and deploy

Two things stay out of version control: the app's `web/.env` (copied from `web/.env.example`) and the CDK platform context (`infra/cdk.context.json`, copied from the example or generated by `scripts/discover-platform-context.sh`). With those in place, `scripts/deploy-platform.sh` builds the images and deploys the stacks. The walkthrough is in [`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md).

## Docs

- **[architecture.md](architecture.md)**: the system walkthrough, from the three moving parts down to the design decisions worth defending. Start here.
- **[security.md](security.md)**: how untrusted users get root without putting anything else at risk.
- **[lab-lifecycle-design.md](lab-lifecycle-design.md)**: provisioning and the three-layer teardown, next to the code.
- **[local-dev-without-aws.md](local-dev-without-aws.md)**: what runs without AWS today, and what a full local-Docker mode would take.

## License

MIT. See [`../LICENSE`](../LICENSE).
164 changes: 164 additions & 0 deletions docs-draft-v2/architecture.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,164 @@
# OpsCoach architecture

OpsCoach gives each learner a real, throwaway Linux host and grades what they actually do to it. The system is three parts: a **web app** on ECS Fargate (UI, terminal bridge, grading), a **per-session EC2 lab host** the learner operates and that dies on a timer, and a **shared platform** (ALB, Cognito, VPC) the app plugs into rather than rebuilding. Everything below is how those three talk to each other safely.

**Scope.** Single-region training infrastructure for one app on a borrowed ALB/Cognito platform. Non-goals: this is not a hostile multi-tenant sandbox (the isolation model assumes a cooperative-but-curious learner, see [security.md](security.md)), and not a high-availability production service.

This doc starts at the diagram and components, then defends the design decisions. The step-by-step request flows are in a collapsible block so you can skip them unless you need the wire-level detail.

## The system

![OpsCoach AWS architecture](architecture.svg)

The numbered arrows above:

1. **Request.** Browser to the ALB over HTTPS. The ALB runs `authenticate-cognito` (Cognito hosted UI, Google); a post-auth passphrase gate in the app must also pass before any page renders. Only the health check and logout skip Cognito.
2. **Route.** The ALB forwards to the OpsCoach web service on Fargate, in a private, egress-only subnet.
3. **Provision, operate, grade.** Fargate launches and drives the per-session EC2 host: `RunInstances` at the start, an SSH PTY for the browser terminal, and the grader over SSH.
4. **Ready webhook.** The host calls back to the service once it is up, resolved through Cloud Map and authenticated with a shared secret.
5. **Terminal.** The browser streams over a WebSocket to the custom Node server, which relays to the host's shell over SSH.
6. **Teardown.** A per-session EventBridge Scheduler one-shot triggers the terminator Lambda; a 5-minute sweep is the backstop.

Grey dashed lines are supporting paths: Fargate to RDS PostgreSQL in the isolated subnet, the host pulling its lab image from ECR, and the opt-in direct SSH from a learner's laptop.

Colour key: purple is networking (ALB, Cloud Map); orange is compute and containers (Fargate, EC2, Lambda, ECR); red is identity and secrets (Cognito, Secrets Manager); blue is the database (RDS); pink is app integration (EventBridge Scheduler).

## Components

| Component | AWS service | Role |
| --- | --- | --- |
| Web app + terminal bridge | ECS Fargate | Next.js app + custom Node server; WebSocket-to-SSH PTY bridge; session lifecycle, grading, dashboard |
| Edge auth | ALB + Cognito | `authenticate-cognito` at the load balancer (hosted UI to Google), then a post-auth passphrase gate |
| Lab host | EC2 (per session) | Ephemeral AL2023 / arm64 host running the lab container; learner SSH target |
| Container images | ECR | Images for the web service and each lab |
| Database | RDS PostgreSQL | Sessions, check runs, grader results (isolated subnet) |
| Service discovery | Cloud Map | In-VPC address for lab-host-to-web callbacks |
| Teardown | EventBridge Scheduler + Lambda | One-shot per-session schedule fires a terminator Lambda; 5-minute sweep backstop |
| Secrets | Secrets Manager | Database credentials and the callback HMAC secret |

## The design decisions

Each of these was a fork where the cheaper, obvious choice was the wrong one. The pattern is constraint, decision, trade-off.

**A real host per session, not a shared sandbox.**
Teaching operations means real systemd, real root, and real packages, but you cannot hand untrusted users root on shared infrastructure. The cheaper alternatives, a shared multi-tenant host or browser-only containers, make container isolation the only wall between a hostile learner and everyone else, so one container escape compromises every session. Instead, every session gets its own ephemeral EC2 host, and the security model treats that host (not the container on it) as the real boundary, killed on a timer. The cost is a minute or two of provisioning latency and per-session compute, bought back as realism and a blast radius of exactly one throwaway box. Full model in [security.md](security.md).

**A custom Node server for the browser terminal.**
A browser terminal needs a long-lived, two-way connection to a shell, and Next.js on its own does not hold one. The alternative, a managed real-time service or a separate WebSocket process, adds a moving part for what is a thin bridge. Instead, a small `server.js` wraps Next, upgrades the WebSocket, authenticates the session through an internal call, and bridges to the host with an `ssh2` PTY. The cost is running a custom server instead of stock Next, plus a 25-second keepalive so the ALB does not cut an idle terminal, in exchange for a real shell in the browser with nothing for the learner to install.

**Grade real state, not answers.**
Multiple-choice cannot tell you whether someone can actually run a box. Instead, the grader SSHes into the live host with a least-privilege environment and checks real state (services up, files in place, config correct), returning structured results. The cost is that graders are per-pack code that has to run against a live machine, in exchange for a pass that means the box is genuinely in the right state.

**Three independent teardown paths.**
A leaked instance costs real money, and the control plane never observes what the learner does inside their SSH session (direct-SSH learners bypass it entirely). A single cleanup mechanism is a single point of failure for the one control that bounds spend. Instead there are three: an on-host SSH-idle watcher, a one-shot EventBridge Scheduler set at provision time, and a 5-minute sweep over expiry tags. Any one is enough, and all are idempotent. The cost is more moving parts, for a hard guarantee that nothing runs forever. Full design in [lab-lifecycle-design.md](lab-lifecycle-design.md).

**Borrow the platform; import by ID.**
A demo should not stand up its own ALB, Cognito, and VPC. Instead the CDK imports a shared platform's resources by ID from local context, and the real IDs stay out of the repo. The cost is that the app cannot bootstrap its own world from nothing, in exchange for dropping cleanly into a real shared environment.

## The flows, step by step

<details>
<summary><b>Auth, provision, terminal, grading, teardown</b> (expand for the wire-level sequences)</summary>

&nbsp;

**1 · Authentication and access gate**

```mermaid
sequenceDiagram
autonumber
actor U as Browser
participant ALB as ALB
participant Cog as Cognito
participant App as Fargate
U->>ALB: HTTPS request
ALB->>Cog: authenticate-cognito (if enabled)
Cog-->>ALB: OIDC tokens
ALB->>App: forward + signed identity header
alt no gate cookie
App-->>U: redirect to /gate
U->>App: POST /api/gate (passphrase)
App-->>U: set gate cookie, continue
end
App-->>U: app page
Note over ALB,App: /api/health and /logout skip Cognito
```

**2 · Provision a lab session**

```mermaid
sequenceDiagram
autonumber
actor U as Browser
participant App as Fargate
participant EC2 as Lab host
participant ECR as ECR
participant Sch as Scheduler
U->>App: POST /api/sessions (start lab)
App->>App: create session in RDS; mint keys + callback token
App->>EC2: RunInstances (Launch Template + user-data)
App->>Sch: CreateSchedule (terminate at T + maxLifetime)
Note over EC2: user-data: install Docker, block IMDS,<br/>ECR login, run lab container, set authorized_keys
EC2->>ECR: pull lab image
EC2->>App: ready webhook (via Cloud Map + shared secret)
App-->>U: session ready (host, port)
```

**3 · Browser terminal and SSH**

```mermaid
sequenceDiagram
autonumber
actor U as Browser
participant Srv as Node server
participant EC2 as Lab host
U->>Srv: WebSocket /api/sessions/:id/shell?token
Srv->>Srv: shell-auth (internal secret) resolves host + key
Srv->>EC2: ssh2 PTY on port 22 (per-session key)
EC2-->>Srv: stdout / stderr
Srv-->>U: stream to xterm.js
Note over Srv,EC2: 25s keepalive ping holds the ALB idle timer
Note over U,EC2: opt-in: SSH directly to the host's public IP
```

**4 · Live grading**

```mermaid
sequenceDiagram
autonumber
actor U as Browser
participant App as Fargate
participant G as Grader
participant EC2 as Lab host
U->>App: POST /api/sessions/:id/grade
App->>G: spawn grader (allowlisted env, no task-role creds)
G->>EC2: SSH port 22 (grader key) runs ops status
EC2-->>G: JSON check results
G-->>App: { passed, checks[] }
App->>App: persist check run in RDS
App-->>U: live results on the dashboard
```

**5 · Idle and lifetime teardown**

```mermaid
sequenceDiagram
autonumber
participant Sch as Scheduler
participant L as Terminator Lambda
participant EC2 as Lab host
participant App as Fargate
Sch->>L: fire at T + maxLifetime
L->>L: read callback secret (Secrets Manager)
L->>EC2: TerminateInstances
L->>App: shutdown callback (mark session stopped)
Note over EC2,App: backstop: 5-min EventBridge sweep<br/>terminates any host past its ExpiresAt tag
```

</details>

## See also

- **[security.md](security.md)** for the security model and its trade-offs.
- **[lab-lifecycle-design.md](lab-lifecycle-design.md)** for provisioning and the three-layer teardown in depth.
- **[`../infra/PLATFORM_INTEGRATION.md`](../infra/PLATFORM_INTEGRATION.md)** for plugging into a shared ALB/Cognito platform.
Loading