Site ↗
Documentation sections
Operations · 0.8.1

Recovering an interrupted deployment

After a deploy error, inspect state, resume or close the operation; handle failed first startup and reconfigure.

On this page

After a deploy error or lost connection, first determine what completed. Starting a new operation does not replace recovery of the old one.

Work from the project root: Source on the VPS as the deploy user, Archive/Registry on the controlling computer, and prod-local on your computer.

1. Find the operation ID and stage

bash ./deploy/control status --env local --json
bash ./deploy/control reconcile --env local --json
bash ./deploy/control logs --env local --service backend

For a server, replace local with the environment name. reservation means an operation holds the environment; current is the last successful release. A container can run even with current: null, for example after a failed first startup.

2. Choose resume or closure

Resume is appropriate when the cause is resolved and you want the same release:

bash ./deploy/control resume --env local \
  --operation OPERATION_ID --confirm local

Resume supports provision, deploy, backup, down, certificates, rollback, and restore. Reconcile only shows state and does not release the lock. Completed stages are skipped. An uncertain migration result is checked against the database; some side effects need manual investigation. Resume cannot substitute another image and has no --release parameter.

Abandon closes the selected operation without success:

bash ./deploy/control abandon --env local \
  --operation OPERATION_ID --confirm local

Data and history remain. It does not undo migrations, restore previous code, or reopen access automatically. A closed operation cannot be resumed. Do not delete reservation or edit its JSON manually.

Provision interrupted during HTTPS certificate issuance

Open the certbot.log path in the error message. Fix the reported cause: DNS, HTTP challenge access, or challenge file permissions. Then resume the same provision using its ID. Current tools preserve a log of each attempt and reuse prepared secrets. This applies to current generated tools: git pull does not replace an old release's saved payload; such an operation may need separate investigation.

After provision resumes successfully, you still need to deploy the application:

bash ./deploy/control deploy --env production --release release-001.json --confirm production

Use the release JSON used when provision began. A new build is unnecessary for retrying certificate checks.

If older tools report secrets.json: cannot overwrite existing file, do not delete secrets: update the project's generator tools. This is not a reason to recreate the database.

First installation created tables but did not record a release

missing release record with an existing application schema is not first install means the database is already populated but no successful first release exists. A new deploy cannot treat it as empty. There is currently no universal way to switch such an installation to another release.

If data matters, preserve the database and state, then investigate the schema's origin and migrations. Do not create current.json or fake migration records. For disposable prod-local with unneeded data, use the separate procedure below.

Recreate disposable prod-local

This option applies only to a local environment you agree to restart with an empty database. The new installation has no previous records, accounts, or images; we move the entire old directory into an archive rather than deleting it. Do not use this for production, dev-local, or valuable data instead of agreed recovery.

All commands below run from the generated project root. The environment name is deliberately fixed: local, preset prod-local. Adapt commands before using another name.

First close the known active operation

Check status --env local --json. Only if this environment has an active operation, substitute its exact ID:

bash ./deploy/control abandon --env local \
  --operation OPERATION_ID_FROM_STATUS --confirm local

Skip this if no operation is active. If abandon fails, stop and investigate; do not delete reservation manually.

Stop only local containers and preserve the directory

The following snippet explicitly runs in Bash; it does not use docker system prune, volume removal, or data rm:

bash <<'BASH'
set -euo pipefail
project_root=$(pwd -P)
env_file=$(mktemp)
trap 'rm -f "$env_file"' EXIT
bash deploy/generated/runtime/environment.sh "$project_root" local > "$env_file"
jq -e '.environment=="local" and .preset=="prod-local"' "$env_file" >/dev/null
state_root=$(jq -er .root "$env_file")
test "$state_root" = "$project_root/deploy/.state/local"
jq -e --arg root "$state_root" '
  .media.storage=="local" and .media.root==($root+"/data/media") and
  .backup.kind=="local" and .backup.root==($root+"/backups")
' "$env_file" >/dev/null
test -d "$state_root"
test ! -L "$state_root"
project_name=$(jq -er .project "$env_file")
compose_project="$project_name-local"
archive="$project_root/deploy/.state/local-archive-$(date -u +%Y%m%dT%H%M%SZ)-$$"
test ! -e "$archive"
container_ids=$(docker ps -aq --filter "label=com.docker.compose.project=$compose_project")
while IFS= read -r container_id; do
  test -n "$container_id" || continue
  docker stop "$container_id"
  docker rm "$container_id"
done <<< "$container_ids"
mv "$state_root" "$archive"
printf 'Старое окружение сохранено: %s\n' "$archive"
BASH

Commands select containers by Compose project, stopping only this prod-local. Dev containers with another Compose project name keep running.

The preserved folder contains stopped PostgreSQL files, images, keys, and history. This is a copy of the environment directory, rather than a portable SQL dump. Restoring the previous instance requires its settings and Docker images. Do not delete the archive immediately after creating the new installation.

Create a clean environment again

First ensure the required fix is present in source and generated project files. Then:

bash ./deploy/control provision --env local
bash ./deploy/control check --env local
bash ./deploy/control up --env local
bash ./deploy/control status --env local --json

provision creates images, keys, and certificates again. If you trusted the local CA in your OS, remove the old CA and import the new one: deploy/.state/local/pki/ui-ca.crt. Then create the first superadministrator following prod-local instructions.

This resets a disposable environment while preserving old files in an archive; it does not resume the previous operation. To transfer needed records, use a separate data procedure after checking the new installation.

Timeout or lost connection

After SSH loss, use status/reconcile: the server may have continued. If build failed before an operation was created, fix the network cause and repeat the build with a new output. If reservation exists, investigate that operation first.

A stage waits up to 1800 seconds by default; DEPLOY_STAGE_TIMEOUT_SECONDS accepts 1–86400. More time does not fix inaccessible DNS/registry. CI timeout must cover multiple stages.

Interrupted secret rotation

Normal resume does not handle rotate-secrets. Preserve journals and prepared files: some secrets may already have been applied. Do not start another rotation or remove old keys until you know which stage completed. The saved rotation state needs investigation.

Interrupted reconfigure

Preserve the same new JSON and repeat reconfigure --env production --confirm production --operation ID; keep --on-host when switching from Source. Normal resume is not used here.

Abandon is possible before configuration publication; first restore the local JSON to the previous saved configuration when checking the environment. After publication begins, repeat reconfigure with the same new JSON and ID rather than abandon.