{"id":760,"date":"2026-10-06T05:02:51","date_gmt":"2026-10-06T05:02:51","guid":{"rendered":"https:\/\/abrarqasim.com\/blog\/docker-compose-healthcheck-the-depends-on-that-only-waited-for-start\/"},"modified":"2026-10-06T05:02:51","modified_gmt":"2026-10-06T05:02:51","slug":"docker-compose-healthcheck-the-depends-on-that-only-waited-for-start","status":"publish","type":"post","link":"https:\/\/abrarqasim.com\/blog\/docker-compose-healthcheck-the-depends-on-that-only-waited-for-start\/","title":{"rendered":"Docker Compose Healthcheck: Stop Trusting depends_on"},"content":{"rendered":"<p>Short version for the impatient: <code>depends_on<\/code> on its own only waits for a container to start, not for the thing inside it to be ready. If your app boots before Postgres accepts connections, that is the bug, and a healthcheck plus <code>condition: service_healthy<\/code> is the fix.<\/p>\n<p>I have shipped this bug more than once. The pattern is always the same. Everything works on my laptop because the database container is warm from yesterday. Then I run <code>docker compose up<\/code> on a fresh VPS, the API container starts in two seconds, Postgres takes eight to run its init scripts, and the API dies with &ldquo;connection refused&rdquo; and a stack trace that looks like it is my fault. With <code>restart: unless-stopped<\/code> it eventually recovers, so I called it flaky and moved on. That was a mistake. A deploy that only works on the third attempt is a coin flip with extra steps.<\/p>\n<p>This post covers what the healthcheck options do, a Compose file I would actually run on a small server, and the part nobody warns you about: an unhealthy container does not get restarted by Compose. Everything below is checked against the Docker docs, which I link as I go.<\/p>\n<h2 id=\"what-depends_on-really-waits-for\">What depends_on really waits for<\/h2>\n<p>The short syntax is the one most of us write:<\/p>\n<pre><code class=\"language-yaml\">services:\n  api:\n    build: .\n    depends_on:\n      - db\n  db:\n    image: postgres:18\n<\/code><\/pre>\n<p>Compose starts <code>db<\/code> first, then <code>api<\/code>. That is the whole contract. &ldquo;Started&rdquo; means the container process exists. Postgres may still be replaying WAL or running your init SQL, and Compose does not care. The <a href=\"https:\/\/docs.docker.com\/compose\/how-tos\/startup-order\/\" rel=\"nofollow noopener\" target=\"_blank\">Docker startup order guide<\/a> says this directly: a database needs to start its own services before it can handle incoming connections, and you detect that ready state with the <code>condition<\/code> attribute.<\/p>\n<p>There are three conditions. <code>service_started<\/code> is what the short syntax gives you. <code>service_healthy<\/code> waits until the dependency&rsquo;s healthcheck passes. <code>service_completed_successfully<\/code> waits for a one-shot container to exit with code 0, which is the right tool for migrations. I will come back to that last one, because it is the underrated option.<\/p>\n<h2 id=\"writing-a-healthcheck-that-tells-the-truth\">Writing a healthcheck that tells the truth<\/h2>\n<p>A healthcheck is a command Docker runs inside the container on a timer. Exit code 0 means healthy, 1 means unhealthy. The <a href=\"https:\/\/docs.docker.com\/reference\/compose-file\/services\/\" rel=\"nofollow noopener\" target=\"_blank\">Compose services reference<\/a> lists the knobs: <code>test<\/code>, <code>interval<\/code>, <code>timeout<\/code>, <code>retries<\/code>, <code>start_period<\/code> and <code>start_interval<\/code>.<\/p>\n<p>For Postgres, the image ships with <code>pg_isready<\/code>, so there is nothing to install:<\/p>\n<pre><code class=\"language-yaml\">  db:\n    image: postgres:18\n    environment:\n      POSTGRES_USER: app\n      POSTGRES_PASSWORD: ${DB_PASSWORD}\n      POSTGRES_DB: app\n    healthcheck:\n      test: [&quot;CMD-SHELL&quot;, &quot;pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}&quot;]\n      interval: 10s\n      timeout: 5s\n      retries: 5\n      start_period: 30s\n<\/code><\/pre>\n<p>That <code>$$<\/code> is not a typo. Compose treats a single <code>$<\/code> as its own variable interpolation and would try to fill it from your host environment, usually with an empty string. Doubling it passes a literal <code>$<\/code> through, so the shell inside the container expands the variable. I lost a long evening to a healthcheck that ran <code>pg_isready -U  -d<\/code> because I wrote one dollar sign.<\/p>\n<p>The <code>start_period<\/code> option is the one I misunderstood for a long time. It is a grace window. According to the Docker docs, failures during that window do not count against <code>retries<\/code>. If a check succeeds during the window, the container counts as started and later failures start counting. So a slow first boot does not burn your five retries before Postgres has even finished initialising.<\/p>\n<p>The <code>start_interval<\/code> option, added in Compose 2.20.2 per the reference, lets you probe more often during the start period. Mine is usually <code>2s<\/code>, so the dependent service starts a few seconds sooner on a fast machine. It is optional and I skip it more often than not.<\/p>\n<h2 id=\"the-curl-trap-in-slim-images\">The curl trap in slim images<\/h2>\n<p>For your own web service, the instinct is <code>curl -f http:\/\/localhost:3000\/healthz<\/code>. It works until you switch to a slim or distroless base image and <code>curl<\/code> is no longer there. The check fails, the container goes unhealthy, and the logs of the app itself look perfectly fine. I stared at a healthy-looking Node process for twenty minutes before I ran <code>docker inspect<\/code> and read the health log.<\/p>\n<p>Use whatever the image already has. Alpine images carry BusyBox <code>wget<\/code>, and a language runtime can probe itself:<\/p>\n<pre><code class=\"language-yaml\">  api:\n    build: .\n    healthcheck:\n      test: [&quot;CMD&quot;, &quot;wget&quot;, &quot;-qO-&quot;, &quot;http:\/\/localhost:3000\/healthz&quot;]\n      interval: 15s\n      timeout: 3s\n      retries: 3\n      start_period: 20s\n<\/code><\/pre>\n<p>Keep the endpoint cheap. My <code>\/healthz<\/code> returns 200 if the process can answer HTTP, and that is all. I used to make it query the database too, on the theory that &ldquo;healthy&rdquo; should mean &ldquo;fully working&rdquo;. Then a ten-second database stall turned every API container unhealthy at once, and the dependent services lined up behind them. A shallow check answers one question, which is whether this process is alive. Your monitoring can ask the deeper questions separately.<\/p>\n<h2 id=\"a-compose-file-i-would-run-on-a-small-server\">A Compose file I would run on a small server<\/h2>\n<p>Here is the shape I use for a typical app with Postgres, a migration step and the API:<\/p>\n<pre><code class=\"language-yaml\">services:\n  db:\n    image: postgres:18\n    restart: unless-stopped\n    volumes:\n      - pgdata:\/var\/lib\/postgresql\n    healthcheck:\n      test: [&quot;CMD-SHELL&quot;, &quot;pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}&quot;]\n      interval: 10s\n      retries: 5\n      start_period: 30s\n\n  migrate:\n    build: .\n    command: [&quot;.\/bin\/migrate&quot;]\n    depends_on:\n      db:\n        condition: service_healthy\n    restart: &quot;no&quot;\n\n  api:\n    build: .\n    restart: unless-stopped\n    depends_on:\n      db:\n        condition: service_healthy\n        restart: true\n      migrate:\n        condition: service_completed_successfully\n    healthcheck:\n      test: [&quot;CMD&quot;, &quot;wget&quot;, &quot;-qO-&quot;, &quot;http:\/\/localhost:3000\/healthz&quot;]\n      interval: 15s\n      retries: 3\n      start_period: 20s\n\nvolumes:\n  pgdata:\n<\/code><\/pre>\n<p>The <code>migrate<\/code> service is the underrated condition from earlier. It runs once, exits 0, and <code>api<\/code> does not start until that has happened. No entrypoint script that loops on <code>pg_isready<\/code>, no <code>sleep 10<\/code> hidden in a Dockerfile. I deleted both of those from old projects and nothing was lost.<\/p>\n<p>The <code>restart: true<\/code> flag under the <code>db<\/code> dependency is a newer option. Per the reference it needs Compose 2.17.0, and it restarts the dependent service when you restart <code>db<\/code> through a Compose operation such as <code>docker compose restart<\/code>. The docs are explicit that this does not cover automatic restarts by the container runtime after a crash. Keep that distinction in mind, because it leads straight into the next problem.<\/p>\n<p>To bring the whole thing up and know whether it worked, use the <code>--wait<\/code> flag from the <a href=\"https:\/\/docs.docker.com\/reference\/cli\/docker\/compose\/up\/\" rel=\"nofollow noopener\" target=\"_blank\">docker compose up reference<\/a>:<\/p>\n<pre><code class=\"language-bash\">docker compose up -d --wait --wait-timeout 120\n<\/code><\/pre>\n<p><code>--wait<\/code> implies detached mode and blocks until services are running or healthy. It exits non-zero if they are not, which makes it usable in a deploy script or a CI job. That one line replaced a block of <code>sleep<\/code> and <code>curl<\/code> retries in my deploy scripts. If you are weighing deploy tooling more broadly, I compared two self-hosted options in <a href=\"https:\/\/abrarqasim.com\/blog\/coolify-vs-dokploy-what-actually-differs-after-migrating-a-client\" rel=\"noopener\">Coolify vs Dokploy after migrating a client<\/a>, which is a different question from this one but the same kind of server.<\/p>\n<h2 id=\"unhealthy-does-not-mean-restarted\">Unhealthy does not mean restarted<\/h2>\n<p>Here is the part I got wrong the longest. I assumed that an unhealthy container would be restarted. In plain Docker Compose it is not. The restart policy reacts to a container exiting. A container whose process is alive but whose healthcheck keeps failing just sits there labelled <code>unhealthy<\/code>, and traffic keeps going to it if nothing in front of it checks that label.<\/p>\n<p>So the healthcheck gives you three things: ordering at startup, a status in <code>docker ps<\/code>, and a signal for <code>--wait<\/code>. It does not give you self-healing. If you want automatic recovery on a single host, you need something that watches the status and acts. A tiny sidecar like <code>willfarrell\/autoheal<\/code> is a common choice, and a reverse proxy that routes only to healthy upstreams covers the traffic side. Alternatively, make the app exit on its own when it detects it is broken, and let <code>restart: unless-stopped<\/code> do its job.<\/p>\n<p>I have not settled on one answer here. For my own small deployments I let the app crash loudly and rely on the restart policy, plus an uptime monitor that pings from outside. For anything with real customers I add the proxy-level check as well. If your setup is bigger than a couple of containers on one VPS, this is a sign to look at an orchestrator, because Swarm and Kubernetes treat health as an input to scheduling and Compose does not.<\/p>\n<h2 id=\"what-to-change-this-week\">What to change this week<\/h2>\n<p>Open your current <code>compose.yaml<\/code> and look for every <code>depends_on<\/code> that has no <code>condition<\/code>. For each database or cache, add a healthcheck and switch the dependency to <code>service_healthy<\/code>. Move your migration command into its own service with <code>service_completed_successfully<\/code>. Then run <code>docker compose up -d --wait<\/code> on a machine with an empty volume and time it. If it passes cold, you have fixed the bug. If you want to see how I set up this kind of infrastructure for client projects, my work is on <a href=\"https:\/\/abrarqasim.com\" rel=\"noopener\">abrarqasim.com<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>depends_on only waits for a container to start. Here is how Docker Compose healthcheck, service_healthy and up &#8211;wait fix cold-start failures, plus the restart gap.<\/p>\n","protected":false},"author":2,"featured_media":759,"comment_status":"","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"","rank_math_description":"depends_on only waits for a container to start. Here is how Docker Compose healthcheck, service_healthy and up --wait fix cold-start failures, plus the restart gap.","rank_math_focus_keyword":"docker compose healthcheck","rank_math_canonical_url":"","rank_math_robots":"","footnotes":""},"categories":[302],"tags":[80,125,303,8],"class_list":["post-760","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-devops","tag-devops","tag-docker","tag-docker-compose","tag-self-hosting"],"_links":{"self":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/posts\/760","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/comments?post=760"}],"version-history":[{"count":0,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/posts\/760\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/media\/759"}],"wp:attachment":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/media?parent=760"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/categories?post=760"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/tags?post=760"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}