Skip to content

Consistent, documented deploy process for staging and production - #1046

Merged
maxkadel merged 6 commits into
mainfrom
deployment
Jul 14, 2026
Merged

maxkadel merged 6 commits into
mainfrom
deployment

Conversation

@maxkadel

Copy link
Copy Markdown
Member

Summary

Replaces the undocumented, hand-edited-in-production deploy process with a single reproducible path for both environments: bin/deploy <environment> [git_ref] [image_tag].

  • Layered env files: .env.common (committed, non-secret, shared) + .env.<environment> (gitignored secrets), sourced from the shared ADVENTIST 1Password vault via provision/bin/fetch-secrets and pushed to servers via bin/push-env.
  • docker-compose.staging.yml rewritten to mirror production's topology (healthchecks, restart: unless-stopped, host bind mounts) so a staging deploy is actually predictive of production.
  • bin/deploy: fetches tags (--force, so a moved tag can't abort the deploy), checks out the ref, re-execs itself after any self-update so a mid-run script change can't run stale buffered logic, loads env via plain source (not the dotenv CLI — the server's installed python3-dotenv-cli doesn't merge multiple -e files), brings up the stack, and polls the /up healthcheck.
  • ops/DEPLOY.md: full runbook, including one-time 1Password setup, the fresh-environment bootstrap gotchas found on staging's first real run (Solr's security.json and precompiled Sprockets assets get shadowed by bind mounts; fcrepo needs its own Postgres role the base Postgres image never creates), and a verification checklist.
  • provision/site.yml: matches production's actual directory ownership (adventist-data group, 2775), and sets safe.directory for git operations on shared-owned paths.

Validated end-to-end against a real staging deploy (fresh environment, tenant creation, file ingest all working).

Test plan

  • bin/deploy staging completes and passes the /up healthcheck
  • Migrations run via initialize_app before web/worker start
  • Tenant creation and file ingest succeed against the real S3 buckets
  • Container restart survives (restart: unless-stopped)
  • First production run (planned for a maintenance window, using the same script)

🤖 Generated with Claude Code

maxkadel and others added 5 commits July 8, 2026 15:10
Staging ran a different, dev-oriented docker-compose.yml (named volumes, no
healthchecks/restart policy) than production's docker-compose.production.yml
(host bind mounts, healthchecks, restart: unless-stopped), so a successful
staging deploy proved nothing about production. This makes staging structurally
match production and adds the tooling to deploy either one consistently:

- .env.common: ~65 shared, non-secret config keys extracted from
  .env.production, committed so staging/production can't silently drift on
  them again. .env.staging/.env.production shrink to just secrets and
  per-environment values (gitignored, sourced from the ADVENTIST 1Password
  vault via provision/bin/fetch-secrets).
- docker-compose.staging.yml: rewritten to mirror docker-compose.production.yml
  (healthchecks, restart policy, root user, dependency graph) instead of the
  dev-oriented file, with host paths adapted to staging's single-volume
  layout. docker-compose.production.yml's AWS/Fedora S3 bucket names are
  pulled out of hardcoded compose values into env vars, since staging
  inheriting those hardcoded values would have meant it silently shared
  production's live S3 buckets.
- bin/deploy: canonical deploy path for both environments - checks out a git
  ref, pulls images, brings up the stack, and waits on the /up healthcheck.
  Stashes (doesn't discard) any local drift before pulling.
- bin/push-env: pushes current secrets from an operator's laptop (the server
  has no 1Password access) onto the target server via personal SSH access,
  backing up whatever's already there first.
- provision/site.yml: /store/keep/adventist_knapsack now gets
  root:adventist-data 2775 (matching production's actual, previously
  untracked permissions) so personal accounts can deploy without a shared
  credential; adds the git safe.directory exception needed for that.
- Remove april from ssh_users (no longer with the company).

See ops/DEPLOY.md for the full runbook.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…eploy

The server's actual installed tool (python3-dotenv-cli 2.2.0, via
provision/site.yml) doesn't support -o/-f at all, and its -e flag doesn't
merge multiple files - passing two silently drops the first file's keys
entirely (verified directly against a real deploy attempt on staging, which
also surfaced a genuine Solr credential mismatch - see .env.staging).

Replaced with plain `set -a; source .env.common; source .env.<environment>;
set +a` before invoking docker-compose directly. This has no dependency on
whatever dotenv variant happens to be installed, and correctly gives
.env.<environment> precedence over .env.common on any overlapping key
(verified: override wins, common-only keys survive, per-env-only keys
survive).

Also documents that bin/deploy must be run as a personal user, not root:
running it as root means git checkout/pull rewrite .env.common (a tracked
file) with root's default permissions, silently un-doing the group-writable
state bin/push-env needs for the next person's push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 229a621)
A moved/re-pushed tag (as happened repeatedly while fixing this branch today)
leaves a stale local tag ref on any server that already fetched it once -
the next plain \`git fetch --tags\` refuses to update it ("would clobber
existing tag") and aborts the whole fetch, taking the deploy down with it.
Deploying from a trusted, controlled remote, so forcing tag updates here is
safe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 7e34590)
Self-modifying script problem, confirmed directly: a git pull that updated
bin/deploy still ran the OLD dotenv invocation for the rest of that same
process, even though the file on disk was already correct (Fast-forward had
completed). The running bash process had already read past that point using
the pre-update content. Re-exec after the checkout/pull/submodule-update step
so everything downstream always runs from the actual current file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 36f297a)
Documents the fixes now in bin/deploy (dotenv independence, tag --force
fetch, re-exec after self-update) plus the fresh-environment bootstrap
gotchas found on staging's first real deploy: seeding Solr's security.json
and precompiled assets past bind-mount shadowing, creating fcrepo's own
Postgres role, and the corrected samvera-original-files-staging bucket note.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

Test Results

      4 files  ±0        4 suites  ±0   41s ⏱️ -1s
362 935 tests ±0  362 894 ✅ ±0  41 💤 ±0  0 ❌ ±0 
    939 runs  ±0      898 ✅ ±0  41 💤 ±0  0 ❌ ±0 

Results for commit eabd209. ± Comparison against base commit ef58553.

♻️ This comment has been updated with latest results.

Comment thread provision/inventory.yml Outdated
@ShanaLMoore

Copy link
Copy Markdown
Contributor

ah I overlooked that that was a comment 🙃

@maxkadel
maxkadel merged commit af991f6 into main Jul 14, 2026
11 checks passed
@maxkadel
maxkadel deleted the deployment branch July 14, 2026 16:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants