Deploying with Docker and GHCR: Ship Images, Not Code

DevOps•

It's Friday, 6:40 PM. You SSH into production, run git pull, then pip install -r requirements.txt (or npm install, or composer install, pick your poison). Somewhere deep in your dependency tree, a package you've never heard of published a new release forty minutes ago, and it doesn't compile on your server. The install that passed CI this morning dies halfway through. The code is new, half the dependencies are old, the app won't start, and your weekend just got cancelled.

Sound familiar? The problem isn't pip. It isn't npm, Composer, or Maven either. The problem is that you are building your application on the production server. Every deploy is a fresh roll of the dice, and the thing you tested is never exactly the thing you shipped.

A while back I wrote about deploying Django on a Linux server with a Git hook, Gunicorn, and Supervisor. I still like that setup for one app on one box. But add a Next.js frontend, a background worker, a teammate's Go service, and a staging server that should behave exactly like production, and "git pull on the server" stops scaling. I recently moved a project with a Django API, a Next.js frontend, and a background worker onto the pipeline in this article, and this is everything I'd tell a teammate doing the same.

The idea fits in one sentence: build your app once into a Docker image, store it in GitHub Container Registry (GHCR), and let your servers do exactly one thing: pull and run. It doesn't care whether your app is Django, FastAPI, Node.js, PHP, Go, or Java. Only the Dockerfile changes. The pipeline, the server, and the rollback story stay the same.

So grab a coffee (or tea, I don't judge), and let's ship some images.

What you'll learn:

  • Why "build once, deploy everywhere" beats building on the server, and what actually changes.
  • The container contract every app must follow, whatever the language.
  • Production-ready Dockerfiles for Django, FastAPI, Node.js, PHP, Go, and Java.
  • A GitHub Actions pipeline that pushes to GHCR, and a tagging strategy that makes every deploy traceable.
  • How to deploy without opening SSH to the internet, roll back in seconds, and dodge the pitfalls that get everyone once.

1. The Core Idea: Ship Artifacts, Not Source Code

Before shipping containers existed, loading a cargo ship was chaos. Sacks of coffee, barrels of oil, crates of machine parts, each loaded by hand with its own rules. Then someone standardized the steel box, and the crane stopped caring what was inside. It just moved boxes.

A Docker image does the same thing for your deployments. Your server stops caring whether it runs Python, Node.js, PHP, Go, or Java. It moves boxes.

               git push
                   │
                   ▼
┌────────────────────────────────────┐
│ GitHub Actions                     │
│ test + build ONCE                  │
└──────────────────┬─────────────────┘
                   │ docker push
                   ▼
┌────────────────────────────────────┐
│ GHCR                               │
│ ghcr.io/your-org/myapp:sha-3f2a1c9 │
└──────────────────┬─────────────────┘
                   │ docker pull
      ┌────────────┼────────────┐
      ▼            ▼            ▼
   staging      prod-1       prod-2

Same image everywhere. Only the .env differs.

Here's what that changes in practice:

Build on the server (git pull)Ship an image (Docker + GHCR)
What gets deployedYour code plus whatever the package manager resolves todayThe exact bytes CI built and tested
What the server needsPython, Node, PHP, a JDK, compilers, build tools…Docker. That's it.
Where a broken build failsOn production, mid-deployIn CI, where production never notices
Two servers, same result?HopefullyGuaranteed: same image, same digest
Rollbackgit revert, rebuild, prayPoint at the previous tag. Seconds.
Adding a Go or Java serviceA new toolchain on every serverSame pipeline, new Dockerfile

If you remember one rule from this article, make it this one:

Build once, then promote the same artifact everywhere. Environments differ in configuration, never in code.


2. Why GHCR?

A container registry is the warehouse between your CI and your servers. Docker Hub, AWS ECR, Google Artifact Registry, and Azure Container Registry all work, and everything in this article applies to them (only the login changes). But if your code already lives on GitHub, GHCR (ghcr.io) is the path of least resistance:

  • No extra secrets for CI. GitHub Actions pushes with the built-in GITHUB_TOKEN, which expires when the job finishes. No long-lived password sitting in your repository settings.
  • Permissions follow the repo. An image pushed from a workflow is linked to its repository and inherits that repository's access. Can see the code? Can pull the image. Can't? Can't.
  • Private by default. New packages start private. Going public is opt-in.
  • Cost. Public images are free. At the time of writing, GitHub's billing docs also say container storage and bandwidth are currently free for private images, with at least a month's notice before that changes.
RegistryPick it when
GHCRYour code and CI already live on GitHub
Docker HubYou publish public images for the whole world (free-tier pulls are rate-limited)
ECR / Artifact Registry / ACRYou're all-in on one cloud and want IAM-native auth next to your compute

3. The Container Contract: What Every Stack Must Agree To

This is the part most tutorials skip, and it's the real reason one pipeline can deploy six different stacks. Before you write a single line of Dockerfile, your app has to honor a contract. The platform only works if every box has the same handles.

RuleWhat it meansThe classic violation
Config from the environmentDatabase URLs, secrets, and flags come from env vars at runtimeA .env or settings file baked into the image
Listen on 0.0.0.0Bind to all interfaces inside the containerServer bound to 127.0.0.1, unreachable from outside its own container
Log to stdout/stderrPrint logs and let Docker collect themWriting to /app/logs/app.log, gone when the container is replaced
Stay statelessUploads and sessions live in a volume, object storage, or RedisUser uploads saved inside the container filesystem
Shut down gracefullyCatch SIGTERM, finish in-flight work, exitShell-form CMD swallows the signal, and Docker kills the app 10 seconds later
Expose a health endpoint/healthz returns 200 only when the app can serve trafficNo healthcheck: the container is "running" but broken
Run as non-rootThe process runs as an unprivileged userEverything runs as root

Two of these deserve a closer look.

Graceful shutdown. On every deploy, Docker sends your process SIGTERM, waits 10 seconds, then sends SIGKILL. If your CMD uses shell form (CMD gunicorn app:app), your app runs as a child of /bin/sh, and the signal may never reach it. Always use exec form (CMD ["gunicorn", "app:app"]), and if you use an entrypoint script, end it with exec "$@" so your app replaces the shell.

The build-time config trap. "Config from the environment" has a sneaky exception: frontend bundles. Next.js inlines every NEXT_PUBLIC_* variable into the JavaScript bundle during next build, and Vite does the same with VITE_*. The Next.js docs say it plainly: build one Docker image, deploy it to several environments, and those values are frozen at whatever they were during the build. Congratulations, your production site is calling the staging API.

The fix is architectural, not a flag:

  • Call your API through a same-origin relative path (/api/...) and let the reverse proxy route it. The browser never needs to know the API's hostname.
  • For values that truly differ per environment, read them on the server at request time, or serve a tiny config endpoint that the client fetches on startup.

4. The Dockerfile: One Pattern, Six Stacks

Every good production Dockerfile follows the same five rules, whatever the language:

  1. Multi-stage builds. Compilers, dev dependencies, and build tools live in a builder stage. The final runtime stage gets only what it needs to run.
  2. Dependency manifests first. Copy requirements.txt, package-lock.json, composer.lock, go.sum, or pom.xml and install dependencies before copying your source. Docker caches that layer, so changing one line of code doesn't reinstall the entire internet.
  3. Pin the base image. python:3.14-slim, not python:latest. Latest is a moving target.
  4. Run as a non-root user, by numeric UID (USER 10001). Kubernetes' runAsNonRoot check can't verify a user name, so numeric keeps your images portable.
  5. Exec-form CMD, starting a real production server.

And every repo gets a .dockerignore, so your build context doesn't ship secrets, your Git history, or 800 MB of node_modules to the builder:

# .dockerignore
.git
.env
.env.*
node_modules
vendor
.venv
__pycache__
target
dist
build
.next
*.log

Now the part that actually changes per stack.

Python: Django and FastAPI

One Dockerfile covers both. Only the last line differs.

# ---------- Stage 1: build wheels (compilers live here, never in production) ----------
FROM python:3.14-slim AS builder
RUN apt-get update && apt-get install -y --no-install-recommends build-essential \
    && rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY requirements.txt .
RUN pip wheel --no-cache-dir --wheel-dir /wheels -r requirements.txt

# ---------- Stage 2: slim runtime ----------
FROM python:3.14-slim
ENV PYTHONDONTWRITEBYTECODE=1 \
    PYTHONUNBUFFERED=1
RUN useradd --create-home --uid 10001 app
WORKDIR /app
COPY --from=builder /wheels /wheels
RUN pip install --no-cache-dir /wheels/* && rm -rf /wheels
COPY --chown=app:app . .
USER 10001
EXPOSE 8000

# Django
CMD ["gunicorn", "config.wsgi:application", "--bind", "0.0.0.0:8000", "--workers", "3"]
# FastAPI (use this instead)
# CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]
  • gunicorn or uvicorn must be in your requirements.txt. manage.py runserver and uvicorn --reload belong on your laptop, never in an image.
  • PYTHONUNBUFFERED=1 makes Python write logs immediately instead of buffering them, which is exactly what "log to stdout" needs.
  • Django tip: run collectstatic during the build (with a throwaway SECRET_KEY if your settings demand one) and serve static files with WhiteNoise, so the image is fully self-contained.

Node.js: Express, NestJS, and Friends

# ---------- Stage 1: install everything, build, then drop dev dependencies ----------
FROM node:24-alpine AS builder
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev

# ---------- Stage 2: runtime with production dependencies only ----------
FROM node:24-alpine
ENV NODE_ENV=production
WORKDIR /app
COPY --from=builder --chown=node:node /app/package.json ./
COPY --from=builder --chown=node:node /app/node_modules ./node_modules
COPY --from=builder --chown=node:node /app/dist ./dist
# 1000 is the image's built-in "node" user
USER 1000
EXPOSE 3000
CMD ["node", "dist/main.js"]
  • Run node directly, not npm start. npm sits between Docker and your app and can swallow SIGTERM. Node running as PID 1 also won't exit on SIGTERM unless your code handles it, so add a process.on('SIGTERM', ...) handler or set init: true on the service in Compose.
  • Next.js: set output: 'standalone' in your Next.js config, copy .next/standalone, .next/static, and public into the runtime stage, and start it with CMD ["node", "server.js"]. And remember the build-time config trap from section 3.

PHP: Laravel and Symfony

# ---------- Stage 1: Composer dependencies (the composer image ships git + unzip) ----------
FROM composer:2 AS vendor
WORKDIR /app
COPY composer.json composer.lock ./
RUN composer install --no-dev --prefer-dist --no-scripts --no-autoloader --ignore-platform-reqs
COPY . .
RUN composer dump-autoload --no-dev --optimize --no-scripts

# ---------- Stage 2: Apache + mod_php runtime ----------
FROM php:8.4-apache
RUN docker-php-ext-install pdo_mysql opcache \
    && a2enmod rewrite \
    && sed -i 's#/var/www/html#/var/www/html/public#' /etc/apache2/sites-available/000-default.conf
COPY --from=vendor --chown=www-data:www-data /app /var/www/html
EXPOSE 80
  • --ignore-platform-reqs is only safe because the runtime stage installs the PHP extensions. Add every extension your composer.json needs to that docker-php-ext-install line.
  • php-fpm speaks FastCGI, not HTTP. If you use the -fpm images, you need nginx (or another FastCGI-aware server) in front; a plain HTTP proxy_pass to port 9000 won't work. The Apache image above sidesteps that, and FrankenPHP is a modern single-binary alternative.
  • Laravel users: don't run php artisan config:cache in the Dockerfile. It snapshots your environment variables at build time, which is the build-time config trap all over again. Run it when the container starts.

Go

# ---------- Stage 1: compile a static binary ----------
FROM golang:1.27 AS builder
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -trimpath -ldflags="-s -w" -o /out/server ./cmd/server

# ---------- Stage 2: distroless (no shell, no package manager, just your binary) ----------
FROM gcr.io/distroless/static-debian13:nonroot
COPY --from=builder /out/server /server
EXPOSE 8080
ENTRYPOINT ["/server"]
  • CGO_ENABLED=0 matters. A CGO-linked binary on a static image dies with the wonderfully misleading exec /server: no such file or directory. The file exists; the C library it needs doesn't.
  • Distroless has no shell, so a curl-based healthcheck can't work. Give your binary a -healthcheck flag, or check health from the outside.

Java: Spring Boot

# ---------- Stage 1: build the jar with the full JDK ----------
FROM eclipse-temurin:25-jdk AS builder
WORKDIR /app
COPY .mvn/ .mvn/
COPY mvnw pom.xml ./
RUN ./mvnw -q dependency:go-offline
COPY src ./src
RUN ./mvnw -q package -DskipTests

# ---------- Stage 2: the JRE is all we need at runtime ----------
FROM eclipse-temurin:25-jre
RUN useradd --system --uid 10001 app
WORKDIR /app
COPY --from=builder /app/target/*.jar app.jar
USER 10001
EXPOSE 8080
ENTRYPOINT ["java", "-XX:MaxRAMPercentage=75", "-jar", "app.jar"]
  • The JVM is container-aware, but by default it only uses 25% of the container's memory for the heap. -XX:MaxRAMPercentage=75 lets it use most of the limit while leaving room for everything else.
  • -DskipTests is fine here because your tests already ran in CI before this image was built. Using Gradle? Same shape, with ./gradlew bootJar.

The Only Part That Changes

StackBuild stageRuntime imageStart commandThe gotcha that bites
Djangopip wheelpython:3.14-slimgunicornrunserver is not a production server
FastAPIpip wheelpython:3.14-slimuvicorn--reload stays on your laptop
Node.jsnpm ci + build + prunenode:24-alpinenode dist/main.jsnpm start can swallow SIGTERM
PHPcomposer install --no-devphp:8.4-apacheApache (built in)php-fpm speaks FastCGI, not HTTP
GoCGO_ENABLED=0 go builddistroless/static/serverA CGO binary says "no such file or directory"
Java./mvnw packageeclipse-temurin:25-jrejava -jarThe heap defaults to 25% of container memory

Everything else in this article is identical for all six.


5. The Pipeline: Build Once, Push to GHCR

Now the fun part. This workflow builds an image on every push to main (and on every v* version tag), pushes it to GHCR, and deploys it. It's the same file for every stack. The only stack-specific input is your Dockerfile.

# .github/workflows/deploy.yml
name: Build and Deploy

on:
  push:
    branches: [main]
    tags: ['v*']

jobs:
  build:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      packages: write # lets the built-in GITHUB_TOKEN push to ghcr.io
    steps:
      - uses: actions/checkout@v7

      - uses: docker/setup-buildx-action@v4

      - name: Log in to GHCR
        uses: docker/login-action@v4
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - name: Generate tags and labels
        id: meta
        uses: docker/metadata-action@v6
        with:
          images: ghcr.io/${{ github.repository }} # lowercased for you
          flavor: latest=false # we never deploy :latest, so don't even publish it
          tags: |
            type=sha
            type=ref,event=branch
            type=semver,pattern={{version}}

      - name: Build and push
        uses: docker/build-push-action@v7
        with:
          context: .
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          cache-from: type=gha
          cache-to: type=gha,mode=max

  deploy:
    needs: build
    if: github.ref == 'refs/heads/main'
    runs-on: [self-hosted, production]
    environment: production
    concurrency:
      group: production-deploy
      cancel-in-progress: false
    steps:
      - name: Deploy the exact image we just built
        run: ~/app/deploy.sh "sha-${GITHUB_SHA::7}"

A few things worth noting:

  • permissions: packages: write is what lets the built-in GITHUB_TOKEN push. Grant it per job (least privilege) instead of flipping your whole repository's default token permissions to read/write.
  • metadata-action writes the tags, lowercases the image name (Docker rejects uppercase, and github.repository happily contains it), and adds OCI labels like org.opencontainers.image.revision that tie the image back to its exact commit.
  • The type=gha cache keeps Docker layers in GitHub's Actions cache. When only your code changed, the dependency layers are reused and just the last few layers rebuild.
  • The deploy job runs on a self-hosted runner that lives on your server. More on that choice in section 7.
  • concurrency on the deploy job means two quick merges never deploy on top of each other.
  • environment: production scopes production-only secrets to this job and, depending on your GitHub plan, lets you require a human to approve it before it runs.
  • Add a test job and make build depend on it with needs: test. A red test should never become an image.

Tagging Strategy: Why :latest Is a Lie

A tag is a sticky note on a box. Anyone can peel it off and stick it on a different box, and that's exactly what happens to :latest and :main on every push. A digest (sha256:...) is the serial number stamped into the box itself. It never changes.

TagExampleDoes it move?Use it for
Commit SHAsha-3f2a1c9NeverDeploying. Traceable to one exact commit
BranchmainEvery pushHumans asking "what's on main right now?"
SemVer1.4.2Never (by convention)Releases, changelogs, customers
latestlatestWhenever anyone pushes anythingNothing in production. Ever.

Deploy :latest and your server runs "whatever was pushed most recently, by anyone, from any branch." Two servers pulling it five minutes apart can run different code, and when something breaks, latest can't tell you what changed. sha-3f2a1c9 can: just run git show 3f2a1c9.

That's why the workflow deploys the SHA tag, never a branch tag. (Want to be bulletproof? Reference the image by digest, ghcr.io/your-org/myapp@sha256:..., and nothing can ever move it.)


6. The Server: Pull, Don't Build

Here's the beautiful part: your server no longer needs Python, Node.js, PHP, Go, or a JDK. It needs Docker and four small files (plus a log that writes itself).

~/app/
├── compose.yaml   # which images run, and how they connect
├── .env           # secrets + the live TAG (chmod 600, never committed)
├── nginx.conf     # reverse proxy: the only thing facing the internet
├── deploy.sh      # pull → migrate → swap → verify → prune
└── deploys.log    # written by deploy.sh: what went live, and when

Keep these files in your repo too (a deploy/ folder works well), so changes get reviewed like code. The server just holds a copy.

One-Time Setup

Install Docker Engine and the Compose plugin from Docker's official apt repository. Their convenience script is meant for dev machines, not production. If this is a fresh VPS, harden SSH first. Then:

# A dedicated deploy user that can run Docker
sudo adduser deploy
sudo usermod -aG docker deploy

# As that user: log in to GHCR ONCE, with a token that can only READ packages
read -s GHCR_TOKEN   # paste the token; it won't echo or land in your shell history
echo "$GHCR_TOKEN" | docker login ghcr.io -u your-github-username --password-stdin

Two senior notes here:

  • The docker group is effectively root. Anyone in it can mount the host's filesystem into a container. That's fine for a dedicated deploy user on a dedicated server. Just don't hand it out casually.
  • GHCR only accepts classic personal access tokens, so create one with only the read:packages scope, ideally on a dedicated bot account that can see nothing but this app's packages. If the server is ever compromised, that token can pull images and not much else.

compose.yaml

Shown for a Python app. For another stack, only the image name, the worker command, and the healthcheck command change.

# ~/app/compose.yaml
services:
  app:
    image: ghcr.io/your-org/myapp:${TAG:?TAG is not set}
    restart: unless-stopped
    env_file: .env
    healthcheck:
      # Use a tool that exists in YOUR image: python -c, node -e, wget (Alpine)...
      test: ["CMD", "python", "-c", "import urllib.request as r; r.urlopen('http://127.0.0.1:8000/healthz', timeout=2)"]
      interval: 10s
      timeout: 3s
      retries: 5
      start_period: 30s
    depends_on:
      db:
        condition: service_healthy
    networks: [edge, backend]

  worker:
    image: ghcr.io/your-org/myapp:${TAG:?TAG is not set} # same image...
    command: ["celery", "-A", "config", "worker", "--loglevel=info"] # ...different process
    restart: unless-stopped
    env_file: .env
    depends_on:
      db:
        condition: service_healthy
    networks: [backend]

  db:
    image: postgres:18-alpine
    restart: unless-stopped
    environment:
      POSTGRES_DB: ${POSTGRES_DB}
      POSTGRES_USER: ${POSTGRES_USER}
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
    volumes:
      - pgdata:/var/lib/postgresql # Postgres 18+ path (17 and older: /var/lib/postgresql/data)
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U ${POSTGRES_USER} -d ${POSTGRES_DB}"]
      interval: 5s
      timeout: 3s
      retries: 10
    networks: [backend]

  proxy:
    image: nginx:1.30-alpine
    restart: unless-stopped
    ports:
      - "80:80"
    volumes:
      - ./nginx.conf:/etc/nginx/conf.d/default.conf:ro
    depends_on: [app]
    networks: [edge]

networks:
  edge: # proxy <-> app
  backend: # app and worker <-> database, never exposed

volumes:
  pgdata:

A few things to understand here:

  • ${TAG:?TAG is not set} makes Compose refuse to start without an explicit version. No accidental latest, ever.
  • Same image, two jobs. app and worker run the identical image with different commands (Celery here; Django Q, BullMQ, or a Go queue consumer work the same way). One build, one tag, always in sync.
  • Only proxy publishes a port. The database has no ports: entry, so nothing outside the server can reach it. And because proxy and db share no network, the proxy can't even see it.
  • depends_on with service_healthy starts the app only after Postgres actually accepts connections, not just after its container exists.
  • Healthchecks use what's already in the image. python:*-slim has no curl, so this one uses Python itself.

nginx.conf

# ~/app/nginx.conf
server {
    listen 80;
    server_name example.com;
    client_max_body_size 50m;

    # Resolve "app" through Docker's DNS at request time (cached 10s),
    # so a redeployed container with a new IP doesn't give you 502s
    resolver 127.0.0.11 valid=10s;
    set $app_backend http://app:8000;

    location / {
        proxy_pass $app_backend;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

The resolver + variable trick matters more than it looks. Without it, nginx resolves app to an IP address once, at startup. After a redeploy, the new container can get a different IP, and nginx keeps knocking on the old door. For HTTPS, add a listen 443 ssl server block with your certificates (and publish port 443), or swap nginx for Caddy or Traefik, which fetch Let's Encrypt certificates on their own.

.env

# ~/app/.env  (chmod 600, never committed)

# App config: read at runtime, never baked into the image
SECRET_KEY=change-me
DATABASE_URL=postgres://app:change-me@db:5432/app
ALLOWED_HOSTS=example.com

# Postgres container
POSTGRES_DB=app
POSTGRES_USER=app
POSTGRES_PASSWORD=change-me

# The live image tag. deploy.sh rewrites this line, so keep it last.
TAG=sha-9b1e4d2

Notice @db:5432. Inside a container, localhost means that container, so services find each other by their Compose service names.

deploy.sh

This is the whole deployment: about twenty lines, and the only thing CI ever calls.

#!/usr/bin/env bash
# ~/app/deploy.sh <tag>: deploy (or roll back to) any image tag
set -euo pipefail
cd "$(dirname "$0")"

NEW_TAG="${1:?usage: ./deploy.sh <image-tag>}"
OLD_TAG="$(grep '^TAG=' .env | cut -d= -f2 || true)"
export TAG="$NEW_TAG" # the shell environment wins over .env for the commands below

# 1. Pull first. If the registry or the tag is broken, nothing running is touched.
docker compose pull app worker

# 2. Migrate with the NEW image, before it takes traffic. Other stacks:
#    FastAPI: alembic upgrade head    Laravel: php artisan migrate --force
#    Node: npx prisma migrate deploy  Spring Boot: Flyway runs on startup, skip this
docker compose run --rm app python manage.py migrate --noinput

# 3. Bring the stack up to date and wait for healthchecks. Unhealthy? Put the old version back.
if ! docker compose up --wait --wait-timeout 120; then
  echo "!! ${NEW_TAG} failed its healthcheck, rolling back to ${OLD_TAG}"
  TAG="$OLD_TAG" docker compose up -d
  exit 1
fi

# 4. Record what's live, keep a history, and prune images older than a week
sed -i '/^TAG=/d' .env && echo "TAG=${NEW_TAG}" >> .env
echo "$(date -u +%FT%TZ) ${NEW_TAG}" >> deploys.log
docker image prune -af --filter "until=168h"
echo "==> ${NEW_TAG} is live"

A few things to understand here:

  • Pull before touching anything. If GHCR is down or the tag doesn't exist, the script dies at step 1 and production keeps running the old version. Only app and worker get pulled, so an app deploy never surprises you by upgrading Postgres or nginx.
  • One command for the first deploy and every one after. docker compose up converges the whole stack: the first run starts everything (database and proxy included), and later runs recreate only the containers whose image or config changed. Your database and proxy keep running through app deploys.
  • Migrations run once, as a one-off container, with the new image, before the swap. Not in every container's entrypoint: three replicas racing to migrate the same database is not how you want to spend your afternoon.
  • Healthchecks gate the deploy. --wait blocks until the new containers are healthy. If they aren't within 120 seconds, the script puts the previous tag back and exits non-zero, so the CI job goes red and you know.
  • The .env records what's live, so docker compose ps and docker compose logs always refer to the running version, and deploys.log gives you an audit trail for free.
  • Prune with a time window. Images from the last week stay on disk, so rolling back to a recent version doesn't even need a download.

About downtime: docker compose up replaces a container in place: stop the old one, start the new one. Expect a few seconds of errors during the swap. For internal tools and plenty of products, that's fine. If it isn't, you want two replicas behind the proxy and rolling replacement, which is exactly what tools like Kamal, Docker Swarm, and Kubernetes give you. Your images, tags, and pipeline carry over unchanged.


7. Connecting CI to the Server: Push vs. Pull

The pipeline built and pushed an image. How does the server find out? There are three honest options:

ApproachHow it worksBest forWatch out for
SSH from CIThe workflow SSHes in and runs deploy.shA public VPS; the simplest setup there isSSH must be reachable from GitHub's runners, and a private key lives in your repo secrets
Self-hosted runnerAn agent on the server polls GitHub over outbound HTTPS and runs the deploy job locallyServers behind NAT, VPNs, or corporate firewallsNever on public repos, and the runner user can run Docker, so treat it like root
GitOps / pull agentSomething in the cluster watches Git or the registry and reconcilesKubernetes, fleets of serversMore moving parts than one VM needs

The workflow above uses the self-hosted runner, and for private infrastructure I really like it: the server reaches out to GitHub, and GitHub never reaches in. A deploy needs zero inbound ports, not even SSH.

To set one up, open your repository's Settings → Actions → Runners → New self-hosted runner, choose Linux, and run the download and config.sh commands it shows you (they include the current runner version and a one-time registration token). Add a label so the workflow can target this machine, then install it as a service that runs as your deploy user:

./config.sh --url https://github.com/your-org/myapp --token <REGISTRATION_TOKEN> --labels production
sudo ./svc.sh install deploy
sudo ./svc.sh start

⚠️ Public repositories: GitHub's own security guidance says self-hosted runners should almost never be used with public repos, because anyone can open a pull request that runs code on your machine. For public projects, deploy over SSH instead.

If you'd rather deploy over SSH, only the deploy job changes:

  deploy:
    needs: build
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    environment: production
    concurrency:
      group: production-deploy
      cancel-in-progress: false
    steps:
      - name: Deploy over SSH
        env:
          SSH_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
          KNOWN_HOSTS: ${{ secrets.DEPLOY_KNOWN_HOSTS }}
          HOST: ${{ secrets.DEPLOY_HOST }}
        run: |
          mkdir -p ~/.ssh
          echo "$SSH_KEY" > ~/.ssh/id_ed25519 && chmod 600 ~/.ssh/id_ed25519
          echo "$KNOWN_HOSTS" > ~/.ssh/known_hosts
          # The remote command starts in the deploy user's home directory
          ssh deploy@"$HOST" "app/deploy.sh sha-${GITHUB_SHA::7}"

Run ssh-keyscan your-server once, verify the fingerprint, and store the output in DEPLOY_KNOWN_HOSTS. Don't run ssh-keyscan inside the workflow: that trusts whichever machine answers.


8. Rollbacks: The Superpower You Just Unlocked

Remember that Friday night from the intro? Here's the new version of the story:

# What went live recently?
tail -n 3 ~/app/deploys.log
# 2026-09-23T16:02:11Z sha-5c8d0e7
# 2026-09-24T09:12:03Z sha-9b1e4d2
# 2026-09-24T14:40:51Z sha-3f2a1c9   <- the broken one

# Put the last good version back
~/app/deploy.sh sha-9b1e4d2

That's it. No rebuild, no git revert under pressure, no reinstalling dependencies. The previous image is usually still on disk, so the rollback takes seconds. Prefer clicking? Open the last good run in your repository's Actions tab and re-run just its deploy job. A re-run uses the original run's commit, so it redeploys exactly that image (GitHub allows re-runs for up to 30 days).

The Catch: Images Roll Back, Databases Don't

Here's what separates a senior deploy strategy from a junior one. Rolling back the image is instant. Rolling back the database schema is not. If sha-3f2a1c9 dropped a column, the old image will crash looking for it, and your one-command rollback just became a restore-from-backup evening.

The fix is a discipline called expand and contract: every migration must work with both the new release and the previous one.

ChangeJunior move (breaks rollback)Senior move (rollback-safe)
Rename a columnRENAME COLUMN in the same release as the codeAdd the new column, write to both, backfill, switch reads, drop the old one a release later
Drop a columnDrop it alongside the code changeStop using it first, then drop it in a later release
Add a required columnNOT NULL with no defaultAdd it nullable or with a default, backfill, then enforce

It feels slower. It's also the only reason "just roll back" works at 2 AM.


Common Pitfalls

These are the ones that get everyone at least once. Learn from others' pain:

1. denied: permission_denied when pushing to GHCR

Either the job is missing permissions: packages: write, or the package was first created by pushing with a personal token, so it isn't linked to this repository. Fix the permissions block, then open the package's Package settings → Manage Actions access and add the repository with the Write role.

2. repository name must be lowercase

github.repository is YourName/MyApp, and Docker refuses uppercase image names. metadata-action lowercases it for you. If you build tag strings by hand, lowercase them yourself.

3. exec format error when the container starts

You built on an Apple Silicon Mac (arm64) and deployed to an x86 server (amd64), or the other way around. Let CI build for the server's architecture, or build both with platforms: linux/amd64,linux/arm64 plus docker/setup-qemu-action.

4. The container is unhealthy, but the app works

Usually one of two things. The healthcheck calls a tool that isn't in the image (there's no curl in slim or distroless images, so use python -c, node -e, or Alpine's wget). Or it calls localhost, which on Alpine can resolve to IPv6 ::1 while your app only listens on IPv4. Use 127.0.0.1 explicitly.

5. The app can't reach the database on localhost:5432

Inside a container, localhost is the container itself. Use the service name: db:5432.

6. 502s right after a deploy that vanish when you restart nginx

nginx resolved app once at startup and is still sending traffic to the old container's IP. Use the resolver 127.0.0.11 + variable pattern from the nginx config above.

7. Waiting for a runner to pick up this job...

The job's runs-on labels don't match your runner's labels (compare them under Settings → Actions → Runners), or the runner is offline. Runners update themselves, but if one hasn't updated in 30 days, GitHub stops sending it jobs.

8. The server's disk fills up

Two usual suspects: old images (the prune step in deploy.sh handles those) and container logs. Docker's default json-file log driver never rotates. Cap it in /etc/docker/daemon.json, then restart Docker (the limits apply to newly created containers):

{
  "log-driver": "json-file",
  "log-opts": { "max-size": "10m", "max-file": "3" }
}

9. Secrets show up in docker history

Anything passed through ARG or ENV at build time is stored in the image, readable by anyone who can pull it. Runtime secrets belong in the server's .env. If the build itself needs a secret (say, a token for a private package registry), use a BuildKit secret mount (RUN --mount=type=secret,...), which never lands in a layer.


Final State Overview

PieceIts jobWhere it lives
DockerfileTurns your app, whatever the stack, into an imageYour repo
GitHub ActionsTests, builds once, tags by commit, pushesGitHub
GHCRStores every version, private by defaultghcr.io/your-org/myapp
compose.yaml + .envDeclares what runs, and with which configThe server
deploy.shPulls, migrates, swaps, verifies, rolls back if unhealthyThe server
nginxThe only door to the internetThe server

Every merge to main becomes an immutable, traceable image. The server needs nothing but Docker. Staging and production run the same bytes with different .env files. And a rollback is one command.

The Senior Checklist

Stick this next to whatever monitor you deploy from:

  • The same image goes to every environment. Only the .env changes.
  • Deploy by commit SHA (or digest). Never :latest.
  • Nothing secret in the image: .dockerignore in place, no secrets in ARG or ENV.
  • Multi-stage builds, pinned base images, a non-root user, exec-form CMD.
  • Every service has a healthcheck, and deploys wait for it.
  • Only the reverse proxy publishes ports.
  • Migrations are backward compatible: expand first, contract later.
  • Rollback is one command, and you've actually practiced it once.
  • Logs rotate, and old images get pruned.

Frequently Asked Questions

Is GHCR free?

Public images are free. For private images, GitHub's billing docs currently say container storage and bandwidth for the Container registry are free, and promise at least a month's notice before that changes. It's worth a quick re-check before you plan a budget around it.

Do I need Kubernetes for this?

No. One VM running Docker Compose handles a surprising amount of traffic, and everything here (immutable images, SHA tags, health-gated deploys, rollbacks) carries over unchanged when you outgrow it. Kubernetes changes where your images run, not how you build and tag them.

Should I use GHCR or Docker Hub?

If your code lives on GitHub, use GHCR: CI pushes with GITHUB_TOKEN, and access follows your repository permissions. Docker Hub shines for public images you want the whole world to discover, but mind its free-tier pull rate limits.

Does this work with GitLab or Bitbucket?

Yes. Swap GHCR for GitLab's built-in container registry (or any other registry) and GitHub Actions for GitLab CI or Bitbucket Pipelines. Build once, tag by commit, pull on the server: the concept doesn't change.

Why not just run git pull from the CI job?

Because then your server is still the one building your app: compilers on production, dependency drift between servers, and no instant rollback. The whole point is that production never builds anything.


That's the whole trick. Build once, tag it with the commit, push it to a registry, and let your servers do the only thing they should ever do: pull and run.

It doesn't matter if your team writes Django on Monday, Go on Tuesday, and on Friday inherits a Spring Boot service nobody wants to own. The pipeline doesn't care. The crane just moves boxes.

And the next time something breaks at 2 AM, you won't be SSH'd into production reinstalling dependencies. You'll run one command, go back to sleep, and fix it properly in the morning.

Happy Shipping!!


Related Reading

References