Coordination server for RL training runs. Three verl nodes train independently and report to this box; MLflow is the one place to watch all of it.
Code and secrets live in this repo. All data lives in /overflow/recursive-knowledge.
Start here if someone sent you a bundle (e.g. all.zip). You do not need
this repo checked out, and you do not need access to the server.
The zip contains three files:
| File | What it is |
|---|---|
<name>.env |
Tracking URI, Cloudflare Access token, rsync target |
<name>_id_ed25519 |
SSH key for checkpoint transfer |
<name>_id_ed25519.pub |
Its public half |
It is a credential. Treat it like a password: no shared filesystems, no git, no Slack.
mkdir -p ~/.rk && unzip all.zip -d ~/.rk
chmod 600 ~/.rk/all.env ~/.rk/all_id_ed25519The chmod is not optional — ssh refuses to use a private key that other
users can read, and the failure message points at the key rather than at
permissions.
uv pip install 'git+ssh://git@github.com/recursive-knowledge/mlflow-vis.git#subdirectory=client'This is what makes authentication automatic: it registers an MLflow
request-header provider that attaches the Cloudflare Access headers to every
REST call, so training code never has to know the server is gated. It also
installs the rk-ckpt-sync command.
Install it into the same interpreter that runs your training. If you use a
venv, activate it first — uv pip install targeting one environment while
python resolves to another is the single most common way this goes wrong.
Confirm it registered before going further:
python -c "from mlflow.tracking.request_header.registry import \
_request_header_provider_registry as r; \
print([type(p).__name__ for p in r])"CloudflareAccessHeaderProvider must appear in that list. If it does not, the
package is not installed in this interpreter, and every MLflow call will
come back as an HTML login page — MLflow reports that as
response body was not in a valid JSON format, naming Cloudflare rather than
the missing package.
If you do not have access to the repo, ask for the client/ directory and
uv pip install ./client instead.
source ~/.rk/all.envPut that in your shell profile and in any batch/job script — a scheduler job does not inherit your interactive shell.
Telemetry, over the tunnel:
python -c "import mlflow; print(mlflow.search_experiments())"Checkpoints, over SSH — a separate path with separate credentials, so it can fail independently:
ssh -i "$RK_CKPT_KEY" -p "$RK_CKPT_PORT" "$RK_CKPT_HOST" true && echo "rsync path OK"A list of experiments and a silent OK mean you are done. Otherwise:
| What you see | What it means |
|---|---|
response body was not in a valid JSON format with Sign in · Cloudflare Access HTML |
The client package is not registered in this interpreter — go back to step 2. This is not a token problem; MLflow sent the request without the Access headers |
403 |
The token is wrong, or it was never added to the application's Service Auth policy |
Permission denied (publickey) |
The server has not authorized your key — ask whoever sent the bundle to run make node-authorize |
Host key verification failed |
Accept the host key once: ssh -i "$RK_CKPT_KEY" -p "$RK_CKPT_PORT" "$RK_CKPT_HOST" |
The two paths use different credentials and fail independently — telemetry working tells you nothing about whether rsync will.
trainer:
project_name: rl-posttraining
experiment_name: all # unique per node
logger: ['console', 'mlflow'] # console = local fallback if tracking blips
save_freq: 50
default_local_dir: /scratch/ckpts/allMetrics go over the tunnel; checkpoint bytes never do. Cloudflare caps
request bodies at 100 MB. rk-ckpt-sync rsyncs the directory over SSH and
logs only a pointer to MLflow.
rk-ckpt-sync --local-dir /scratch/ckpts/all/global_step_50 --step 50 --run-id <run-id>| Flag | Default | Meaning |
|---|---|---|
--local-dir |
required | Checkpoint directory on this node |
--step |
required | global_step this checkpoint corresponds to |
--run-id |
$MLFLOW_RUN_ID, else the active run |
Which MLflow run to attach the pointer to |
--experiment |
$MLFLOW_EXPERIMENT_NAME, else default |
Groups the remote directory |
--kind |
sharded |
sharded = resumable, tied to topology; hf = exported weights |
--dry-run |
off | Print the size and destination, transfer nothing |
Always start with --dry-run. It shows exactly where the bytes would land
without moving any.
Two defaults that bite from a shell hook. After training exits there is no
active MLflow run and MLFLOW_RUN_ID is usually unset, so pass --run-id
explicitly — copy it from the run's URL in the UI. Likewise
MLFLOW_EXPERIMENT_NAME is not exported by the verl config, so without
--experiment your checkpoints land under default/ while the run itself
lives elsewhere. Set both:
export MLFLOW_EXPERIMENT_NAME=all
rk-ckpt-sync --local-dir /scratch/ckpts/all/global_step_50 --step 50 --run-id abc123…Bytes land at $RK_CKPT_ROOT/<experiment>/<run_id>/step_<N> on the server,
staged as step_<N>.incoming and renamed only on success — the server's
retention sweep skips *.incoming, so an interrupted transfer is never
mistaken for a complete checkpoint nor reaped mid-flight. Re-running the same
step is safe; it resumes with --partial and replaces on completion.
It then tags the MLflow run so the checkpoint is discoverable from the UI:
| Tag / metric | Value |
|---|---|
ckpt.step_<N>.uri |
user@host:/path/to/step_<N> |
ckpt.step_<N>.kind |
sharded or hf |
ckpt.step_<N>.size_gb |
Transferred size |
ckpt.step_<N>.config_hash |
Fingerprint of config.json / config.yaml / params.json |
ckpt.latest_step |
Most recent step synced |
ckpt_size_gb (metric) |
Size, plotted against step |
The config hash matters for resuming: a sharded checkpoint needs the exact config it was written with, and a mismatch is the difference between a resumable artifact and a directory of unusable tensors.
Checkpoints are pruned on the server by retention policy (
CKPT_KEEP_LASTsteps per run). To keep one permanently, ask for a.keepfile to be placed in its directory — the sweep never touches those.
Two kinds of traffic, two completely different paths:
| Telemetry | Checkpoints | |
|---|---|---|
| What | metrics, params, tags, small artifacts | sharded model weights |
| Size | kilobytes, continuous | gigabytes, every save_freq |
| Path | Cloudflare tunnel → MLflow | rsync over SSH → /overflow/.../checkpoints |
| Auth | Access service token | per-node SSH key |
Cloudflare proxies cap request bodies at 100 MB. Checkpoints are far larger, so they never touch the tunnel — nodes rsync the bytes and log only a pointer (URI, step, config hash) to MLflow. Keep that split intact.
Through the tunnel (once configured) — the normal way:
https://mlflow.<your-domain>
Cloudflare Access asks for SSO first. That hostname is the only thing this stack exposes to the internet.
Over SSH — always works, no Cloudflare needed:
ssh -L 5000:localhost:5000 ml-login
# then open http://localhost:5000
Supabase Studio (admin console: SQL editor, table browser) is deliberately not published. MLflow is the visualization surface; Studio is for poking at the database.
make studio
ssh -L 8000:localhost:8000 ml-login # then http://localhost:8000
make install # uv env, secrets, vendored Supabase assets, images
make up # postgres + storage + mlflow
make bucket # create the artifact bucket (first run only)
make smoke # prove the whole path worksmake install is idempotent and reproducible — it never overwrites an
existing secret, and dependencies come from the committed uv.lock.
Set up the tunnel and Access application first — full walkthrough in
cloudflare/README.md. Then:
# in env/server.env:
# CLOUDFLARE_TUNNEL_TOKEN=<install token from the dashboard>
# MLFLOW_HOSTNAME=mlflow.<your-domain>
# MLFLOW_ALLOWED_HOSTS=...,mlflow.<your-domain> <- required, see below
make up-tunnel
make health
MLFLOW_ALLOWED_HOSTSmust contain your tunnel hostname. MLflow 3 has a DNS-rebinding guard that returns403 Invalid Host headerfor anything not listed — while/healthkeeps returning 200, so the stack looks fine and every node fails. Host matching includes the port, so list bothhostandhost:5000for local access.
make node-provision NODE=julius # generates env/nodes/julius.env + an SSH key
make node-token NODE=julius # Access service token, written into that file
make node-authorize NODE=julius # grants rsync access (asks for confirmation)
make test-tunnel NODE=julius # verifies telemetry end to endZip env/nodes/julius.env together with its keyfile and send it over a
private channel. What the recipient does with it is
Setting up a compute node at the top of this
file — point them there rather than explaining it again.
make node-authorize is the step that is easy to forget: without it telemetry
works and rsync fails with Permission denied (publickey).
Teammates who only want to look at the dashboard need none of this — send them the URL. Access lets any address matching your policy sign in by email.
Two tiers, and the distinction is the point.
Full control of the box. Generated by make install; chmod 600, gitignored.
| Variable | What it is |
|---|---|
POSTGRES_PASSWORD |
Supabase Postgres superuser + service roles |
MLFLOW_DB_PASSWORD |
MLflow's own least-privilege database role |
JWT_SECRET |
Signs ANON_KEY / SERVICE_ROLE_KEY; changing it regenerates both |
ANON_KEY |
Supabase anonymous JWT (derived) |
SERVICE_ROLE_KEY |
Supabase full-access JWT (derived) — bypasses row-level security |
S3_PROTOCOL_ACCESS_KEY_ID / _SECRET |
MLflow's credentials for the artifact bucket |
PG_META_CRYPTO_KEY |
Studio's credential encryption key |
DASHBOARD_PASSWORD |
Studio login |
CLOUDFLARE_TUNNEL_TOKEN |
Authorizes an outbound tunnel for your account |
CLOUDFLARE_API_TOKEN |
Optional; lets make node-token issue Access service tokens |
Non-secret knobs in the same file: RK_DATA_ROOT, MLFLOW_HOSTNAME,
MLFLOW_ALLOWED_HOSTS, CKPT_SSH_*, and the retention settings
(CKPT_KEEP_LAST, BACKUP_KEEP_DAILY, BACKUP_KEEP_WEEKLY, DISK_WARN_GB).
What a teammate or node actually needs. Contains no database password and no tunnel token; grants exactly telemetry + rsync, revocable per node.
| Variable | What it is |
|---|---|
MLFLOW_TRACKING_URI |
https://mlflow.<your-domain> |
CF_ACCESS_CLIENT_ID / _SECRET |
That node's Cloudflare Access service token |
RK_CKPT_HOST / _PORT / _ROOT |
Where to rsync checkpoints |
<name>_id_ed25519 |
That node's SSH key (ships alongside) |
Revoke one node without touching anything else:
sed -i '/rk-mlflow-julius/d' ~/.ssh/authorized_keys # rsync access
# then delete its service token in the Cloudflare dashboard # telemetry/overflow/recursive-knowledge/
├── supabase/db/ Postgres — runs, metrics, checkpoint pointers
├── supabase/storage/ artifact bucket bytes
├── checkpoints/ <experiment>/<run_id>/step_<N>/ (rsync target)
└── backups/{daily,weekly}/ + backups/logs/ (cron output)
Postgres holds every run, metric, and checkpoint pointer — the checkpoint bytes
on disk are unusable without it. Dumps are pg_dumpall (roles included, since
storage-api authenticates as supabase_storage_admin).
make backup-cron-enable # daily 03:17, weekly Sun 04:47
make backup-cron-status
make backup-cron-disable # existing dumps are kept
make backup TIER=weekly # run one now
make backup-list
make restore F=/overflow/recursive-knowledge/backups/weekly/cluster-*.sql.gzRetention keeps 7 daily and 8 weekly, with a hard floor of 2 per tier — a bad dump can never leave you with zero recoverable copies.
/overflow is a shared pool that runs near capacity, so retention is
load-bearing rather than hygiene.
make usage # what the data root is consuming
make prune # dry run: which checkpoints would go
make prune APPLY=1 # actually delete
make health # warns below DISK_WARN_GBPruning keeps the last CKPT_KEEP_LAST steps per run and never touches an
in-flight *.incoming transfer, anything under 30 minutes old, or a step
directory containing a .keep file.
make help- Supabase is pinned to the commit in
supabase/PINNED_REF;make installvendors its init SQL and Kong config from exactly that ref. To upgrade, bump the SHA, re-run install, diffsupabase/upstream/, and update the image tags incompose.yamlto match that commit'sdocker-compose.yml. - The tracking server runs Python 3.13, not 3.14. MLflow 3.14.0's uvicorn
entry point imports
importlib.abc.Traversable, which 3.14 removed — and uvicorn is mandatory because--allowed-hostsis rejected under gunicorn. The local ops venv stays on 3.14; only the image is pinned back. STORAGE_INTERNAL_URLmust matchMLFLOW_S3_ENDPOINT_URL's host. storage-api rebuilds the S3 canonical request from it, so a mismatch fails every artifact write withSignatureDoesNotMatch.- Kong is not on the MLflow path. MLflow talks to storage-api directly, so the admin gateway being down can't stop a training run from logging.