CloudOpsGuide
kubernetes

Kubernetes Probes Done Right: Liveness, Readiness, and Startup

Intermediate
10 minutes
October 2026
CloudOpsGuide Team

Kubernetes Probes Done Right: Liveness, Readiness, and Startup

Probes are how Kubernetes decides whether your container is alive, ready for traffic, or still starting up. Misconfigure them and you get either zombie pods serving errors or restart loops that never converge. Here's the correct mental model and working config for each.

Table of Contents

The three probes, in plain terms

  • livenessProbe — "should this container be restarted?" Fail it enough times and the kubelet kills and restarts the container.
  • readinessProbe — "should this pod get traffic?" Fail and the pod drops out of the Service endpoints — it stays running but gets no requests.
  • startupProbe — "has this slow starter finished booting?" Disables the other two probes until it succeeds.

The classic incident: a liveness probe that fails during a slow cold-start kills the pod, which restarts, which fails again — CrashLoop forever. startupProbe exists precisely for this.

Probe mechanisms

Four ways to check:

# HTTP — status 200–399 = success
livenessProbe:
  httpGet:
    path: /healthz
    port: 8080
    httpHeaders:
      - name: X-Health
        value: "1"

# TCP — connection open = success
readinessProbe:
  tcpSocket:
    port: 5432

# exec — exit code 0 = success
livenessProbe:
  exec:
    command: ["cat", "/tmp/healthy"]

# gRPC — for gRPC apps with the health-check protocol
livenessProbe:
  grpc:
    port: 50051
    service: myapp.Health

Timing fields: the knobs that matter

livenessProbe:
  httpGet: { path: /healthz, port: 8080 }
  initialDelaySeconds: 5    # wait before first probe
  periodSeconds: 10         # check every 10s
  timeoutSeconds: 2         # fail if slower than 2s
  failureThreshold: 3       # 3 consecutive failures = dead
  successThreshold: 1       # successes needed to be healthy again

The correct production recipe

startupProbe:
  httpGet: { path: /healthz, port: 8080 }
  periodSeconds: 5
  failureThreshold: 24      # up to 2 minutes to boot

livenessProbe:
  httpGet: { path: /healthz, port: 8080 }
  periodSeconds: 15
  failureThreshold: 3

readinessProbe:
  httpGet: { path: /readyz, port: 8080 }
  periodSeconds: 5
  failureThreshold: 2

With startupProbe, use periodSeconds × failureThreshold as your boot budget — no initialDelaySeconds guessing game.

Readiness ≠ liveness — the distinction that prevents outages

  • /healthz (liveness) should only check that the process is fundamentally alive — deadlock-free, event loop responsive. Keep dependencies OUT of it: if your liveness check fails when the database is down, every pod restarts pointlessly during a DB outage.
  • /readyz (readiness) SHOULD check dependencies — DB connections, cache pools. A pod that can't serve shouldn't get traffic, but it also shouldn't be killed.

Common mistakes

  • No probes at all — dead pods keep receiving traffic.
  • Liveness and readiness identical — you get kills where you wanted pauses.
  • Probing a heavyweight endpoint — keep checks cheap; they run forever, on every pod.
  • exec probes for simple HTTP apps — they fork a process per check; HTTP is lighter.

Probes are the contract between your app and the orchestrator. Write real /healthz and /readyz endpoints, protect slow starters with startupProbe, and Kubernetes starts working for you instead of around you.

Related Articles


Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 10 minutes