Backups & restore
What is worth backing up
| Where | What | Back it up? |
|---|---|---|
| The database | organizations, studios, users, permissions, encrypted secrets (SSO, repository tokens, data-source credentials), run history, favorites, audit log | Yes |
| Data volume — uploaded files | data-source files uploaded through the portal, at studio or organization level | Yes — no other copy exists |
| Data volume — built output | materialized project roots, rendered report output | No — rebuilds from your repository |
| Object storage — built output | the same, when TRELLUM_STORAGE_BACKEND=s3 |
No — same reason, and your bucket has its own durability |
| Data volume — git checkouts | working copies of each studio's report repository | No — re-cloned on demand |
| Containers | nothing | Disposable |
They run by default
There is nothing to enable. The backup service starts with the stack, because
the bundled database has no replication and no point-in-time recovery, so these
dumps are the entire recovery story — and a backup you have to remember to turn
on is one that is off on the instances that needed it.
Nightly, into a separate backups volume — not the data volume it protects,
which would not survive the failure it exists for:
| File | What |
|---|---|
db-<stamp>.dump |
pg_dump --format=custom, verified with pg_restore --list before it counts |
studios-<stamp>.tar.gz |
Studio state, excluding git checkouts and query caches, both reproducible |
orgs-<stamp>.tar.gz |
Organization-level uploaded data-source files |
status.json |
What the last run did — read by the backups health check |
A dump that fails verification is deleted rather than kept, so a bad one can
never displace a good one during retention. Retention (BACKUP_KEEP_DAYS,
default 14) only runs after a verified success, so a run of failures cannot age
out your last good backup.
If you moved the database to one you operate yourself, backups follow it — the
service uses DATABASE_URL rather than assuming the bundled container.
Getting them off the host
BACKUP_REMOTE_TARGET=backups@nas.internal:/trellum
Set it in .env. Backups that live only on the host do not survive losing the
host, which is the failure they exist for. Unset, the service says so on every
run and the health check reports kept on this host only.
Danger
Store SECRET_ENCRYPTION_KEY somewhere separate from the backups. A backup
plus the key in one place is a single point of compromise — and without the
key, a restored database cannot decrypt a single stored credential.
Knowing they still work
docker compose exec web python manage.py doctor
The backups check fails when the newest verified backup is older than
BACKUP_MAX_AGE_HOURS (default 36). That is the alert that catches a backup
service which died six weeks ago — otherwise discovered by needing it.
It also fails if the bundled database is in use and no verified backup exists at all. The bundled database is a supported way to run; running it with nothing to restore from is not.
Prove a backup actually restores
docker compose run --rm --entrypoint /bin/sh backup /restore_drill.sh
pg_restore --list proves a file is a well-formed archive. Only a restore
proves it contains your portal. The drill restores the newest dump into a
scratch database, counts the essential tables, drops it, and records the result
— which doctor then reports as restore-tested N days ago. If a drill fails,
the backups check goes red: backups that restore empty look like protection
right up until the day they are needed.
Run it after setting backups up, and on a schedule if you want to keep knowing. An untested backup is a hope.
Restore
docker compose stop web worker
docker compose exec -T db pg_restore -U trellum -d trellum_portal --clean --if-exists \
< /path/to/db-<stamp>.dump
docker compose run --rm web sh -c "cd /data && tar -xzf /backups/studios-<stamp>.tar.gz"
docker compose run --rm web sh -c "cd /data && tar -xzf /backups/orgs-<stamp>.tar.gz"
docker compose up -d
docker compose exec web python manage.py doctor
Afterwards, rebuild report output from each studio dashboard (Run All) or simply wait for the schedules to come round.
Note
To undo a bad upgrade rather than recover from a disaster, use
scripts/rollback.sh, which does the above and re-pins the previous image.
See Upgrades.