diff --git a/docs/superpowers/plans/2026-07-21-database-and-file-backups.md b/docs/superpowers/plans/2026-07-21-database-and-file-backups.md new file mode 100644 index 0000000..e9d4d79 --- /dev/null +++ b/docs/superpowers/plans/2026-07-21-database-and-file-backups.md @@ -0,0 +1,451 @@ +# Database and File Backups Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add a nightly, encrypted, deduplicated disaster-recovery backup for PostgreSQL, uploaded STL files, and Telegram session volumes, stored on a Synology NAS. + +**Architecture:** A Linux-host systemd timer invokes a host orchestration script. The script verifies the mounted Synology NFS share, stops the app/worker/bot services for consistency, and runs a one-shot Docker Compose backup service. The backup service creates a PostgreSQL custom-format dump and stores it together with the three persistent file volumes in a Restic repository on the NAS. A guarded restore command reconstructs the database and volumes, and documentation describes setup and testing. + +**Tech Stack:** Docker Compose, PostgreSQL 16 `pg_dump`/`pg_restore`, Restic repository encryption and retention, Synology NFS, Linux systemd service/timer, Bash. + +## Global Constraints + +- PostgreSQL data must be backed up as a logical custom-format dump; the raw `postgres_data` volume is not the primary backup. +- The `manual_uploads`, `tdlib_state`, and `tdlib_bot_state` volumes are included in every successful snapshot. +- The `tmp_zips` scratch volume is excluded. +- The Docker host must stop `app`, `worker`, and `bot` while volumes are captured; PostgreSQL remains running for `pg_dump`. +- The backup repository is encrypted and stored on a Synology NFS share restricted to the Docker host. +- Retention is 30 daily snapshots; pruning is allowed only after a verified successful backup. +- A failed run must restart services and preserve the last known-good snapshot. +- No in-app backup UI is part of this implementation. +- The repository has no automated test framework; verification uses Bash syntax checks, Docker Compose validation, logs, Restic checks, and a disposable restore rehearsal. + +--- + +## File and Responsibility Map + +Create or modify only these focused units: + +- Create `backup/Dockerfile`: build the one-shot image containing PostgreSQL client tools, Restic, Bash, and the backup entrypoint. +- Create `scripts/backup/container-entrypoint.sh`: run the backup inside Compose, including dump creation, manifest creation, Restic snapshot, verification, and retention. +- Create `scripts/backup/run-backup.sh`: host-level lock, NFS mount validation, service stop/start, and invocation of the one-shot Compose service. +- Create `scripts/backup/restore.sh`: guarded restore orchestration for a selected Restic snapshot. +- Create `deploy/systemd/dragons-stash-backup.service`: systemd unit invoking the host backup script. +- Create `deploy/systemd/dragons-stash-backup.timer`: nightly schedule. +- Create `scripts/backup/README.md`: Synology setup, host mount, secrets, first backup, restore, and operational troubleshooting. +- Modify `docker-compose.yml`: add the profile-gated one-shot `backup` service and its read-only volume mounts. +- Modify `.env.example`: document backup mount, staging, repository, and secret-file configuration without committing secrets. +- Modify `README.md`: add the production backup setup and restore entry points, linking to the detailed backup guide. + +--- + +### Task 1: Add the backup service and configuration contract + +**Files:** +- Modify: `docker-compose.yml` +- Modify: `.env.example` +- Create: `backup/Dockerfile` + +**Interfaces:** +- Consumes: existing `db`, `manual_uploads`, `tdlib_state`, and `tdlib_bot_state` Compose resources. +- Produces: a profile-gated Compose service named `backup` that mounts the three file volumes read-only, connects to the `backend` network, and exposes `/backup` and `/staging` to the container entrypoint. + +- [ ] **Step 1: Add explicit backup environment variables to `.env.example`** + +Add this block without real credentials: + +```dotenv +# Disaster recovery backups +BACKUP_MOUNT_PATH="/mnt/dragonsstash-backups" +BACKUP_STAGING_PATH="/var/lib/dragons-stash-backup/staging" +BACKUP_REPOSITORY="/backup/restic" +BACKUP_RESTIC_PASSWORD_FILE="/etc/dragons-stash/restic-password" +BACKUP_RETENTION_DAYS=30 +BACKUP_APP_VERSION="unknown" +``` + +- [ ] **Step 2: Add the profile-gated `backup` service to `docker-compose.yml`** + +Add a service with these properties: + +```yaml + backup: + profiles: ["backup"] + build: + context: . + dockerfile: backup/Dockerfile + environment: + DATABASE_URL: postgresql://${POSTGRES_USER:-dragons}:${POSTGRES_PASSWORD:-stash}@db:5432/${POSTGRES_DB:-dragonsstash} + RESTIC_REPOSITORY: ${BACKUP_REPOSITORY:-/backup/restic} + RESTIC_PASSWORD_FILE: /run/secrets/restic-password + BACKUP_RETENTION_DAYS: ${BACKUP_RETENTION_DAYS:-30} + BACKUP_APP_VERSION: ${BACKUP_APP_VERSION:-unknown} + user: "0:0" + volumes: + - manual_uploads:/data/uploads:ro + - tdlib_state:/data/tdlib-worker:ro + - tdlib_bot_state:/data/tdlib-bot:ro + - ${BACKUP_MOUNT_PATH:?Set BACKUP_MOUNT_PATH to the mounted Synology share}:/backup:rw + - ${BACKUP_STAGING_PATH:?Set BACKUP_STAGING_PATH to a local staging directory}:/staging:rw + - ${BACKUP_RESTIC_PASSWORD_FILE:?Set BACKUP_RESTIC_PASSWORD_FILE to a root-readable secret file}:/run/secrets/restic-password:ro + depends_on: + db: + condition: service_healthy + networks: + - backend +``` + +Ensure the new service does not have `restart: always` and is not started by the normal production `docker compose up -d` command unless the `backup` profile is explicitly requested. + +- [ ] **Step 3: Create the backup image definition** + +Create `backup/Dockerfile`: + +```dockerfile +FROM postgres:16-alpine + +RUN apk add --no-cache bash restic coreutils + +COPY scripts/backup/container-entrypoint.sh /usr/local/bin/dragons-stash-backup +RUN chmod 0755 /usr/local/bin/dragons-stash-backup + +ENTRYPOINT ["/usr/local/bin/dragons-stash-backup"] +``` + +- [ ] **Step 4: Validate the Compose contract** + +Run on a Linux host with the required variables available: + +```bash +docker compose --profile backup config --quiet +``` + +Expected: exit code `0` and no Compose validation errors. If the required NAS/secret paths are absent, the command must fail with the explicit variable-name error rather than silently using a host path. + +- [ ] **Step 5: Commit the service boundary** + +```bash +git add backup/Dockerfile docker-compose.yml .env.example +git commit -m "feat: add backup compose service" +``` + +### Task 2: Implement the one-shot backup container + +**Files:** +- Create: `scripts/backup/container-entrypoint.sh` + +**Interfaces:** +- Consumes: `DATABASE_URL`, `RESTIC_REPOSITORY`, `RESTIC_PASSWORD_FILE`, `BACKUP_RETENTION_DAYS`, `/data/uploads`, `/data/tdlib-worker`, `/data/tdlib-bot`, `/backup`, and `/staging`. +- Produces: exit `0` only after a verified Restic snapshot and successful retention pruning; non-zero on any failed dump, snapshot, verification, or prune step. + +- [ ] **Step 1: Define strict shell behavior and required inputs** + +The script must begin with: + +```bash +#!/usr/bin/env bash +set -Eeuo pipefail +``` + +Validate that `DATABASE_URL`, `RESTIC_REPOSITORY`, `RESTIC_PASSWORD_FILE`, and `BACKUP_RETENTION_DAYS` are set, that the password file is readable, and that `/backup` and `/staging` are mounted directories. + +- [ ] **Step 2: Create a per-run staging directory and cleanup trap** + +Use a directory below `/staging` named with UTC timestamp and process ID. Register an `EXIT` trap that removes only that directory. Never remove `/staging` itself or any directory under `/backup`. + +- [ ] **Step 3: Create the PostgreSQL dump** + +Run `pg_dump` using the connection URL and custom format: + +```bash +pg_dump --format=custom --file="$RUN_DIR/database.dump" "$DATABASE_URL" +``` + +After the command succeeds, require the dump to be a non-empty regular file. Generate a SHA-256 checksum for the dump in the manifest directory. + +- [ ] **Step 4: Create the manifest** + +Write a JSON manifest containing the UTC backup timestamp, repository path, retention value, dump filename, dump checksum, and the three volume paths captured. Obtain the application image/version from an explicit `BACKUP_APP_VERSION` environment value when supplied; otherwise record `unknown` rather than guessing from mutable container state. + +- [ ] **Step 5: Create one Restic snapshot** + +Run one `restic backup` command against the staged database dump, manifest, and the three mounted persistent volumes. Use stable source labels so the snapshot can be recognized during restore. Do not include `/data/uploads` more than once and do not pass `/data/tdlib` or `/tmp/zips` paths that are not mounted. + +- [ ] **Step 6: Verify and apply retention** + +After `restic backup` succeeds: + +```bash +restic snapshots --latest 1 +restic check +restic forget --keep-daily "$BACKUP_RETENTION_DAYS" --prune +``` + +If any command fails, exit non-zero and do not run `forget --prune`. The host wrapper will restart the stopped services. When invoked with an unrecognized first argument, the entrypoint must pass the remaining arguments to the `restic` binary so operators can inspect the repository through the Compose image without installing Restic on the host: + +```bash +case "${1:-backup}" in + backup) run_backup ;; + restore) run_restore "$@" ;; + *) exec restic "$@" ;; +esac +``` + +- [ ] **Step 7: Build and run a container-only smoke test** + +Run: + +```bash +docker build -f backup/Dockerfile -t dragons-stash-backup:smoke . +bash -n scripts/backup/container-entrypoint.sh +``` + +Expected: image build succeeds and Bash reports no syntax errors. The full snapshot test waits until the host wrapper and a real PostgreSQL/volume environment exist. + +- [ ] **Step 8: Commit the backup container** + +```bash +git add backup/Dockerfile scripts/backup/container-entrypoint.sh +git commit -m "feat: implement encrypted database and volume snapshots" +``` + +### Task 3: Add the host orchestration script and nightly systemd timer + +**Files:** +- Create: `scripts/backup/run-backup.sh` +- Create: `deploy/systemd/dragons-stash-backup.service` +- Create: `deploy/systemd/dragons-stash-backup.timer` + +**Interfaces:** +- Consumes: `.env`/deployment environment, the mounted `BACKUP_MOUNT_PATH`, Docker Compose project, and the `backup` service from Task 1. +- Produces: one host command that safely stops and restarts services and returns the backup container's exit status; systemd runs it nightly. + +- [ ] **Step 1: Implement lock and mount validation** + +The host script must use `flock` on `/run/lock/dragons-stash-backup.lock`, reject a concurrent run, and validate the NAS mount with both `mountpoint --q "$BACKUP_MOUNT_PATH"` and a writable probe file that is immediately removed. A local directory at the same path must not pass validation. The script reads `BACKUP_MOUNT_PATH`, `BACKUP_STAGING_PATH`, `BACKUP_RESTIC_PASSWORD_FILE`, and `BACKUP_RETENTION_DAYS` from the systemd environment file. + +- [ ] **Step 2: Capture service state and define guaranteed restart** + +Before stopping services, record which of `app`, `worker`, and `bot` are running with `docker compose ps --status running -q SERVICE`. Stop only the services that were running. Register an `EXIT` trap that starts exactly those services and preserves the backup command's original exit code. + +- [ ] **Step 3: Invoke the profile-gated backup service** + +After the services stop, run: + +```bash +docker compose --profile backup run --rm backup backup +``` + +Pass through the container exit code. The wrapper must not call the Restic retention command itself; that responsibility stays inside the backup container. + +- [ ] **Step 4: Add the systemd service** + +Create a unit with `Type=oneshot`, `User=root`, `EnvironmentFile=-/etc/dragons-stash/backup.env`, `WorkingDirectory` set to the production Compose directory, `ExecStart` pointing to the absolute `run-backup.sh` path, and `TimeoutStartSec=infinity`. Configure `After=network-online.target docker.service` and `Requires=docker.service`. Do not put the Restic password or database password in the unit file. + +- [ ] **Step 5: Add the nightly timer** + +Create a timer using `OnCalendar=*-*-* 03:00:00`, `Persistent=true`, and `RandomizedDelaySec=15m`. Set `Unit=dragons-stash-backup.service` and `WantedBy=timers.target`. + +- [ ] **Step 6: Validate shell and systemd files** + +Run: + +```bash +bash -n scripts/backup/run-backup.sh +systemd-analyze verify deploy/systemd/dragons-stash-backup.service deploy/systemd/dragons-stash-backup.timer +``` + +Expected: both commands exit `0`. Run `systemctl list-timers dragons-stash-backup.timer` after installation and confirm the next run is scheduled. + +- [ ] **Step 7: Commit scheduling and orchestration** + +```bash +git add scripts/backup/run-backup.sh deploy/systemd/dragons-stash-backup.service deploy/systemd/dragons-stash-backup.timer +git commit -m "feat: schedule nightly off-host backups" +``` + +### Task 4: Implement guarded restore tooling + +**Files:** +- Create: `scripts/backup/restore.sh` + +**Interfaces:** +- Consumes: a Restic snapshot ID, the same repository/password configuration, the backup Compose service, and the live Compose project. +- Produces: restored PostgreSQL data and persistent volumes only after explicit confirmation for live replacement; a non-destructive staging restore by default. + +- [ ] **Step 1: Define restore modes and destructive guard** + +Support these commands: + +```bash +./scripts/backup/restore.sh list +./scripts/backup/restore.sh verify SNAPSHOT_ID +./scripts/backup/restore.sh restore-to-staging SNAPSHOT_ID STAGING_DIR +./scripts/backup/restore.sh restore-live SNAPSHOT_ID --confirm-replace-live-data +``` + +Reject `restore-live` unless the exact confirmation flag is present. `list`, `verify`, and `restore-to-staging` must not stop services or modify live volumes. + +- [ ] **Step 2: Implement snapshot verification and staging restore** + +Use `restic snapshots`, `restic check`, and `restic restore SNAPSHOT_ID --target STAGING_DIR`. Verify that the restored staging tree contains a non-empty custom-format dump, a manifest, `manual_uploads`, `tdlib-worker`, and `tdlib-bot` before reporting success. + +- [ ] **Step 3: Implement live restore sequencing** + +For `restore-live`: + +1. Confirm the Compose project and target repository. +2. Stop `app`, `worker`, and `bot`. +3. Create a safety PostgreSQL dump of the current database into local staging. +4. Restore the selected snapshot to a separate staging directory. +5. Replace the three Docker volumes only after the restored tree passes validation. +6. Recreate the configured database from the restored custom-format dump using `pg_restore --no-owner`. +7. Start services and run the health endpoint plus file-reference verification. + +If any step fails, leave the services stopped, print the exact staging path and failure, and do not delete the safety dump. + +- [ ] **Step 4: Add file-reference verification** + +Run a small SQL query against `manual_upload_files` to enumerate `filePath` values and check each path inside the restored `/data/uploads` tree. Report missing paths and return non-zero if any database reference is broken. + +- [ ] **Step 5: Validate the restore command without touching live data** + +Run: + +```bash +bash -n scripts/backup/restore.sh +./scripts/backup/restore.sh list +``` + +Expected: syntax passes and `list` prints available snapshot IDs without stopping any service or modifying a volume. + +- [ ] **Step 6: Commit restore tooling** + +```bash +git add scripts/backup/restore.sh +git commit -m "feat: add guarded backup restore workflow" +``` + +### Task 5: Document Synology setup, operations, and recovery + +**Files:** +- Create: `scripts/backup/README.md` +- Modify: `README.md` + +**Interfaces:** +- Consumes: the exact environment variables, systemd units, and restore commands from Tasks 1–4. +- Produces: operator-facing instructions that do not require reading implementation files. + +- [ ] **Step 1: Document Synology configuration** + +Document creating the `dragonsstash-backups` shared folder, enabling NFS, restricting the export to the Docker host's fixed IP, and mounting it at `/mnt/dragonsstash-backups`. Include commands for checking the mount: + +```bash +mountpoint /mnt/dragonsstash-backups +touch /mnt/dragonsstash-backups/.write-test +rm /mnt/dragonsstash-backups/.write-test +``` + +Do not document exposing NFS to the Internet. + +- [ ] **Step 2: Document secret and staging setup** + +Document creating the root-readable Restic password file at `/etc/dragons-stash/restic-password`, creating the local staging directory, setting ownership/permissions, and adding the backup variables to the production environment without committing secrets. + +- [ ] **Step 3: Document installation and first-run commands** + +Include: + +```bash +sudo install -m 0644 deploy/systemd/dragons-stash-backup.service /etc/systemd/system/ +sudo install -m 0644 deploy/systemd/dragons-stash-backup.timer /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now dragons-stash-backup.timer +sudo systemctl start dragons-stash-backup.service +sudo journalctl -u dragons-stash-backup.service -n 100 --no-pager +``` + +Explain that the first run may be long because it uploads all existing STL and session data; later Restic snapshots deduplicate unchanged data. + +- [ ] **Step 4: Document monitoring, retention, and restore** + +Document how to inspect timer status, service failures, Restic snapshots, repository checks, and the four restore modes. Explicitly state that `restore-live` is destructive and requires the confirmation flag. + +- [ ] **Step 5: Add a concise production-backup section to the root README** + +Add a link from the deployment/operations section to `scripts/backup/README.md`, state that Docker volumes are not backups, and identify PostgreSQL, STL uploads, and Telegram session volumes as the protected data set. + +- [ ] **Step 6: Commit documentation** + +```bash +git add scripts/backup/README.md README.md +git commit -m "docs: document Synology backup and recovery" +``` + +### Task 6: Verify backup, failure recovery, retention, and restore + +**Files:** +- Modify: `scripts/backup/README.md` only if verification commands need correction. + +**Interfaces:** +- Consumes: the complete backup stack from Tasks 1–5. +- Produces: evidence that the acceptance criteria are met, including a disposable restore rehearsal and a failure-path result. + +- [ ] **Step 1: Validate configuration and scripts** + +Run: + +```bash +docker compose --profile backup config --quiet +bash -n scripts/backup/container-entrypoint.sh scripts/backup/run-backup.sh scripts/backup/restore.sh +systemd-analyze verify deploy/systemd/dragons-stash-backup.service deploy/systemd/dragons-stash-backup.timer +``` + +Expected: all commands exit `0`. + +- [ ] **Step 2: Seed a recognizable test record and STL file** + +Using the existing app/database workflow, create one test upload whose database metadata and file can be identified after restore. Record the expected filename, upload ID, and file checksum before backup. + +- [ ] **Step 3: Run a real backup and inspect the snapshot** + +Run the systemd service manually, then inspect: + +```bash +sudo systemctl start dragons-stash-backup.service +sudo journalctl -u dragons-stash-backup.service --since "10 minutes ago" --no-pager +docker compose --profile backup run --rm backup snapshots +docker compose --profile backup run --rm backup check +``` + +Expected: the service succeeds, the snapshot exists, the repository check succeeds, all services are running again, and the test file checksum is represented in the backed-up volume. + +- [ ] **Step 4: Test the failure path with the NAS unavailable** + +Temporarily unmount the Synology share in a controlled maintenance session, run the systemd service, and confirm it fails before creating a new snapshot. Remount the share and confirm the previously successful snapshot remains listed. Verify that services are running after the failed attempt. + +- [ ] **Step 5: Rehearse a disposable restore** + +Restore the selected snapshot to a disposable Compose project or isolated Docker volumes. Import the database dump, restore the three file trees, start the disposable app/worker/bot services, call `/api/health`, and run the file-reference verification. Confirm the test upload metadata and STL checksum match the pre-backup record. + +- [ ] **Step 6: Verify retention behavior** + +Use a disposable repository or controlled test timestamps to create more than 30 daily snapshots, run the retention command after a successful backup, and confirm that the latest 30 daily snapshots remain. Confirm a failed backup does not invoke pruning. + +- [ ] **Step 7: Record verification evidence** + +Add the actual commands, dates, snapshot ID, restore result, and any environment-specific caveats to the operational notes. Do not commit passwords, session contents, database dumps, or NAS addresses that are intended to remain private. + +- [ ] **Step 8: Commit any documentation corrections** + +```bash +git add scripts/backup/README.md +git commit -m "test: document verified backup and restore procedure" +``` + +## Plan Self-Review + +- **Spec coverage:** PostgreSQL dump, STL volume, both Telegram volumes, Synology NFS, Restic encryption/deduplication, 30-day retention, maintenance window, service restart on failure, guarded restore, file-reference validation, monthly repository/restore checks, and NAS-loss caveat are covered by Tasks 1–6. +- **Placeholder scan:** No `TBD`, `TODO`, or unspecified implementation task remains. Environment-dependent values are explicit configuration variables or operator-supplied paths. +- **Type/interface consistency:** The Compose service name is consistently `backup`; the container command modes are `backup` and `restore`; the host wrapper owns service lifecycle; the restore script owns destructive confirmation; Restic owns snapshots and pruning. +- **Scope check:** The plan contains one operational subsystem with separate backup, restore, scheduling, and documentation units that can each be reviewed and tested independently.