Repository navigation
fix(restore): stream restore inputs, scope dump files per attempt, document restore 409s - #1276
Conversation
The MariaDB, PostgreSQL and MongoDB legacy restore helpers built the whole backup into an in-memory tar before uploading it to the container, and the MongoDB sidecar paths collected the S3 object into memory first. A restore therefore needed RAM proportional to the database. Stage every input on the host in constant memory instead: S3 objects stream into an attempt-owned temp file (never overwriting an existing one), gzip is decompressed file-to-file, and the file is streamed into the container with the shared container_upload helper. The Redis staging helpers move into a shared restore_staging module with typed errors that name the bucket, key and file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: David Viejo <dviejo@kfs.es>
The control-plane backup named its sidecar container and its host files after the backup, so a retry after a failed or crashed attempt collided with the container or files the earlier attempt left behind. The MariaDB dump wrote to a per-backup host file opened with a truncating create, so a retry or a concurrent attempt could overwrite another attempt's dump. Move the control-plane backup onto dump_capture: each attempt gets its own sidecar name, its own directory inside the sidecar and its own host directory, and the dump is streamed out through the Docker archive API instead of a bind mount. The MariaDB dump streams into a host directory owned by the attempt. Both now report typed failures that say where the dump failed, and an attempt only ever deletes what it created. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: David Viejo <dviejo@kfs.es>
The documented 409 only mentioned the unconfirmed cross-service restore. The endpoint also returns 409 when another restore is already active on the target service (restore-already-active, carrying active_restore_run_id) and when the backup is being deleted. Describe all three and regenerate the CLI and web clients. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: David Viejo <dviejo@kfs.es>
|
@greptile-apps review |
📓 Changelog previewThis is what your commits will add to the generated ## [Unreleased]
### Documentation
- **api:** Document every 409 the start-restore endpoint returns
### Fixed
- **restore:** Stream legacy restore inputs instead of buffering them
- **backup:** Scope control-plane and MariaDB dump files to each attempt
- **restore:** Remove partial MongoDB archives and bound uploads by stalls
- **restore:** Give the daemon time to answer after an upload's last byte |
|
…alls A MongoDB sidecar restore removed its staging directory only after the sidecar ran, so a download that failed part-way returned early and left the partial archive in the host temp dir. Hold the directory in a TempDir guard so it is removed on every exit path. Container uploads were bounded by a fixed one-hour deadline, which fails a large restore over a slow link even while it is making progress. Replace it with a stall timeout in the shared upload helper: the upload is abandoned only when the Docker daemon accepts no data for five minutes, and the error reports how many bytes were sent. Callers no longer pass a deadline. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: David Viejo <dviejo@kfs.es>
|
@greptile-apps review |
The stall timer kept running after the last chunk was sent, so a daemon still writing out a large file for more than five minutes made a complete upload fail as stalled. Track when the body is exhausted and switch to a separate confirmation window from then on: the stall timeout plus the time to write the whole file at a conservative 8 MiB/s. Running out of it is reported as a distinct Unconfirmed error that says every byte was received. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: David Viejo <dviejo@kfs.es>
|
@greptile-apps review |
Description
Follow-ups to #1273.
1. Restores no longer load the whole backup into memory.
mongorestoresidecar paths also collected the S3 object into aVecfirst. A restore needed RAM roughly the size of the database, and every in-place PostgreSQL restore of apg_dumpbackup takes the legacy path.create_new, so it never overwrites an existing one.container_uploadhelper.restore_stagingmodule with a typedStagingErrorthat names the bucket, key and file..to_str().unwrap()from the MongoDB legacy restore.2. Dump files and containers belong to one attempt.
dump_capture, like the per-service engines since fix(restore): recover interrupted restores and explain restore and backup failures #1273. Each attempt gets its own sidecar name, its own directory inside the sidecar and its own host directory under<data_dir>/backups/tmp. The dump is streamed out through the Docker archive API instead of a bind mount, so Docker in a VM or on another host works.Exportwith thepg_dumpstderr,EmptyDump,DumpUnreadable,Container,Upload).DumpCaptureError::Execkeeps the retry policy and cancellation of the underlying exec failure.3. The start-restore 409 is fully documented.
restore-already-active(withactive_restore_run_id) and the "backup is being deleted" conflict.spec:update+generate:apifor the CLI,openapi-tsfor the web. Each changed by exactly one line.Scalability
Control plane only, no hot-path code.
dump_captureengines already make, in exchange for no host bind mount.Type of change
Checklist
cargo test --lib)cargo check --libpasses with no warningsgit commit -s) per the DCOVerification
All runs used
--features temps-providers/docker-tests,temps-backup/docker-testsagainst a real Docker daemon:cargo clippy -p temps-backup -p temps-providers --lib --tests -- -D warnings: clean.cargo test --lib -p temps-backup: 338 passed.pg_dump -Fcrestores, streamed from staged files into apostgres:17-alpinecontainer.mongodumparchive throughrestore_from_legacy.cargo test --lib -p temps-providers: 830 passed. Three tests that this branch doesn't touch fail only on a macOS + Colima host:test_init_persists_actual_port_after_conflict_retry: a hostTcpListenerdoes not occupy the port inside Colima's VM, so no retry is forced.cluster_integration_testsmonitor-health tests: they time out through the VM.initand the cluster code, neither of which this PR changes; CI runs them on Linux.bun run spec:check: canonical.python3 scripts/source_attribution.py check: passes.Follow-ups noticed, not changed here
DUMP_SHELL(from reading the code, not reproduced): if the database-list query fails, for example because of a wrong root password, the shell prints a plain-text-- No user databases to dumpand exits 0. That output would be accepted as a non-gzip "dump".restore_backup_filedoes not checkpsql's exit code.docker execprocess inside the service container.Related issues
Follow-up to #1273 (#1236, #1237, #1243).
🤖 Generated with Claude Code