deployer

Health checks & zero downtime

Traffic only moves to a new version after it answered a health check. A version that never becomes healthy is stopped, and visitors keep using the old one.

Serve a cheap health endpoint

  • Answer GET /up (any path you like) with 200 as soon as the app can serve requests, including its database connection. Keep it cheap: it is called during every deploy.
  • Configure the path in deployer.toml: [web] healthcheck = "/up". With a path, only 2xx and 3xx count as healthy. Without one, the deployer probes / and accepts any answer below 500.
  • A redirect (3xx) counts as success and is not followed: an HTTP→HTTPS redirect means the app is up.

How the probe looks

The deployer probes the new container from the server, so your image needs no curl:

GET /up HTTP/1.1
Host: <your primary domain>
X-Forwarded-Proto: https
User-Agent: deployer-healthcheck/1

The real host name and https scheme mean framework host allow-lists (Rails, Django, …) accept the probe without special cases. Each request may take up to healthcheck_timeout (3s by default); the container must be healthy within deploy_timeout (1m by default, up to 15m).

A compose healthcheck is honoured too

If the web service has a compose healthcheck, the container must pass it first; then the HTTP probe runs. Services that others depend on with condition: service_healthy need one.

The switch

  1. The new version starts next to the old one (blue/green), on its own local port.
  2. It passes the health checks.
  3. nginx is reloaded gracefully to send new requests to it; requests in flight finish on the old version.
  4. For a short watch period the deployer keeps checking the new version; if it fails, traffic goes back automatically and the deploy is marked rolled back.
  5. The old version gets drain_timeout (30s) to finish open requests, then stops (SIGTERM, then SIGKILL after stop_timeout).

Your app therefore briefly runs twice. Things that must never run twice (a scheduler, a singleton consumer) belong in a separate service, or set [web] strategy = "recreate" and accept a short downtime.

Migrations

Run them with [release] command: once, in the new image, before traffic moves. Write migrations that the old version can live with during the switch (add columns first, remove them in a later release), so that a rollback stays possible.