Kubernetes Probes Done Right: Liveness, Readiness, and Startup
Kubernetes Probes Done Right: Liveness, Readiness, and Startup
Probes are how Kubernetes decides whether your container is alive, ready for traffic, or still starting up. Misconfigure them and you get either zombie pods serving errors or restart loops that never converge. Here's the correct mental model and working config for each.
Table of Contents
- The three probes, in plain terms
- Probe mechanisms
- Timing fields: the knobs that matter
- The correct production recipe
- Readiness ≠ liveness — the distinction that prevents outages
- Common mistakes
The three probes, in plain terms
- livenessProbe — "should this container be restarted?" Fail it enough times and the kubelet kills and restarts the container.
- readinessProbe — "should this pod get traffic?" Fail and the pod drops out of the Service endpoints — it stays running but gets no requests.
- startupProbe — "has this slow starter finished booting?" Disables the other two probes until it succeeds.
The classic incident: a liveness probe that fails during a slow cold-start kills the pod, which restarts, which fails again — CrashLoop forever. startupProbe exists precisely for this.
Probe mechanisms
Four ways to check:
# HTTP — status 200–399 = success
livenessProbe:
httpGet:
path: /healthz
port: 8080
httpHeaders:
- name: X-Health
value: "1"
# TCP — connection open = success
readinessProbe:
tcpSocket:
port: 5432
# exec — exit code 0 = success
livenessProbe:
exec:
command: ["cat", "/tmp/healthy"]
# gRPC — for gRPC apps with the health-check protocol
livenessProbe:
grpc:
port: 50051
service: myapp.Health
Timing fields: the knobs that matter
livenessProbe:
httpGet: { path: /healthz, port: 8080 }
initialDelaySeconds: 5 # wait before first probe
periodSeconds: 10 # check every 10s
timeoutSeconds: 2 # fail if slower than 2s
failureThreshold: 3 # 3 consecutive failures = dead
successThreshold: 1 # successes needed to be healthy again
The correct production recipe
startupProbe:
httpGet: { path: /healthz, port: 8080 }
periodSeconds: 5
failureThreshold: 24 # up to 2 minutes to boot
livenessProbe:
httpGet: { path: /healthz, port: 8080 }
periodSeconds: 15
failureThreshold: 3
readinessProbe:
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 5
failureThreshold: 2
With startupProbe, use periodSeconds × failureThreshold as your boot budget — no initialDelaySeconds guessing game.
Readiness ≠ liveness — the distinction that prevents outages
- /healthz (liveness) should only check that the process is fundamentally alive — deadlock-free, event loop responsive. Keep dependencies OUT of it: if your liveness check fails when the database is down, every pod restarts pointlessly during a DB outage.
- /readyz (readiness) SHOULD check dependencies — DB connections, cache pools. A pod that can't serve shouldn't get traffic, but it also shouldn't be killed.
Common mistakes
- No probes at all — dead pods keep receiving traffic.
- Liveness and readiness identical — you get kills where you wanted pauses.
- Probing a heavyweight endpoint — keep checks cheap; they run forever, on every pod.
- exec probes for simple HTTP apps — they fork a process per check; HTTP is lighter.
Probes are the contract between your app and the orchestrator. Write real /healthz and /readyz endpoints, protect slow starters with startupProbe, and Kubernetes starts working for you instead of around you.
Related Articles
- Kubernetes for Beginners: Pods, Deployments, and Services
- Kubernetes Ingress and TLS: Exposing Services Properly
Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 10 minutes