Skip to content

Control plane backup and restore ​

The control plane keeps everything it knows (apps, domains, tokens, users, encrypted secrets, deploy history) in one SQLite database. A control plane backup is a consistent snapshot of that file, taken while the server keeps running.

This is separate from database and volume backups, which protect the data of the apps you deploy.

What is in a backup, and what is not ​

In a backupNot in a backup
Every app, domain, environment, project and user recordThe master key
API token hashes and session dataApp volumes and managed database contents
Secrets, still encryptedThe telemetry (metrics and logs) database
Deploy history and audit logTLS private keys held outside the database

The master key is never in a backup

Secrets in the database are encrypted with the master key. A restore without the same master key leaves those secrets unreadable. Keep a separate copy of the master key (a password manager or offline vault), and store it apart from the backups. Whoever holds both can read your secrets.

Backups do contain token hashes and encrypted secrets, so every backup route requires a root-scoped token and files are written with mode 0600.

Where backups live ​

Snapshots are written to <data dir>/control-plane-backups/ and named levelrail-YYYYMMDDTHHMMSSZ.db (UTC). Each one is copied with VACUUM INTO, checked with SQLite's integrity_check, and hashed with SHA-256.

Snapshots sit on the same disk as the database. To survive losing the machine, set up encrypted off-box backups, key escrow and restore drills, or download snapshots and store them elsewhere.

Automatic snapshots ​

VariableDefaultMeaning
APP_CONTROL_PLANE_BACKUP_INTERVAL24hHow often a scheduled snapshot is taken. Go duration syntax. 0 disables scheduled snapshots.
APP_CONTROL_PLANE_BACKUP_RETAIN7How many scheduled snapshots to keep. Older scheduled ones are deleted.

Manual snapshots (created from the CLI or API) and pre-upgrade snapshots are never deleted automatically.

Doctor check ​

levelrail-cli doctor and the dashboard Status page include a control_plane_backup check. It warns when the newest snapshot is older than 3 days (the fix is levelrail-cli control-plane-backups create), is ok when a recent one exists, and reports unknown when scheduled snapshots are disabled with APP_CONTROL_PLANE_BACKUP_INTERVAL=0. It never fails. See troubleshooting.

Stale backup alert ​

The doctor check only helps when someone looks. To be notified instead, create an alert rule of kind control_plane_backup_stale:

bash
levelrail-cli apps alerts create <app> --name "Control plane backup stale" --kind control_plane_backup_stale --channel-id CHANNEL

It fires once the newest snapshot is older than 3 days and sends a resolved notice after the next snapshot lands. Set --for-duration (for example 48h) to change the maximum age. The rule is platform-wide (the app only decides where it is listed), stays quiet when APP_CONTROL_PLANE_BACKUP_INTERVAL=0, and also stays quiet before the first snapshot exists. The dashboard's alert rule dialog and the alerting quick setup prompt offer it too. See observability.

Before an upgrade ​

When the server starts on an existing database and this release carries schema migrations that have not been applied yet, it takes a snapshot first. A brand new database is skipped. If that snapshot fails (for example the disk is full), the failure is logged and the migration still proceeds.

To roll back to that snapshot after a bad upgrade, stop the service, install the previous binary and run levelrail restore-snapshot --list, then levelrail restore-snapshot --dry-run latest to verify it, then levelrail restore-snapshot latest (it asks for confirmation; --yes skips the prompt). It uses the same checks and the same .before-restore-<timestamp> safety copy as restore-db. See Installing.

Downgrade guard ​

If the database has a schema version newer than the binary understands, the server refuses to start and says so. Run a newer release, or restore a snapshot taken by this version.

Taking and managing backups ​

levelrail-cli control-plane-backups create
levelrail-cli control-plane-backups list
levelrail-cli control-plane-backups download <name> --out backup.db
levelrail-cli control-plane-backups verify <name>
levelrail-cli control-plane-backups delete <name>

Without --out, download writes the raw bytes to stdout. The same operations are available in the dashboard under Settings, and over the API (see the API reference):

MethodPath
POST/api/v1/system/backups
GET/api/v1/system/backups
GET/api/v1/system/backups/{name}/download
POST/api/v1/system/backups/{name}/verify
DELETE/api/v1/system/backups/{name}

Verifying a backup ​

A backup you have never checked is a hope, not a backup. Verification proves a snapshot is still intact without restoring anything:

  1. checksum: the file's SHA-256 is recomputed and compared with the checksum recorded when the snapshot was taken. Snapshots from before checksums were recorded pass this check with a note.
  2. integrity: SQLite opens the file read-only and runs integrity_check.
  3. schema_version: the snapshot's schema version must not be newer than this binary supports, otherwise a restore would be refused.
levelrail-cli control-plane-backups verify levelrail-20260101T000000Z.db

The command exits 0 when every check passes and 1 when any fails, so it fits in a cron job or a monitoring script. Over the API the response is {"name", "ok", "checks": [{"name", "ok", "detail"}], "verified_at"}; a failed check is still a 200 with "ok": false.

The last result is kept in a small file next to the snapshot (no database change) and shows up as verified_at and verified_ok on each entry of the list response. The dashboard shows a Verify button and a last-verified badge per backup. Deleting a snapshot removes its verification record too.

Restoring ​

Restoring is an offline operation on the server, because the database cannot be swapped under a running control plane.

  1. Stop the control plane.

  2. Make sure the same master key is available to the restored server.

  3. Run the restore against the backup file:

    levelrail restore-db /path/to/levelrail-20260101T000000Z.db

    The command uses APP_DATA_DIR to find the live database.

  4. Start the control plane.

restore-db first checks the file: integrity_check must pass and its schema version must not be newer than the binary. It then moves the current database aside as levelrail.db.before-restore-<timestamp> and puts the backup in place. Nothing is deleted, so a mistaken restore can be undone by moving that copy back.

Anything that changed after the snapshot was taken (new apps, tokens, deploys) is gone after a restore. Running containers are not touched, and the reconciler converges them toward the restored desired state on startup.

Released under the Apache 2.0 License.