CloudOpsGuide
kubernetes

Fix Kubernetes CrashLoopBackOff: 10 Common Solutions

Intermediate
12 minutes
October 2026
CloudOpsGuide Team

Fix Kubernetes CrashLoopBackOff: 10 Common Solutions

Comprehensive troubleshooting guide for CrashLoopBackOff errors with real-world solutions and preventive measures.

Table of Contents

Understanding CrashLoopBackOff

CrashLoopBackOff occurs when a container repeatedly crashes and Kubernetes restarts it, but it continues to fail.

Lifecycle States

  1. Pending: Pod is waiting to be scheduled
  2. Running: Container is running
  3. Succeeded: Container exited successfully
  4. Failed: Container exited with error
  5. CrashLoopBackOff: Container keeps crashing

When It Happens

# Typical pod status
NAME                      READY   STATUS             RESTARTS   AGE
my-app-5d7f9d6d8b-abc123   0/1     CrashLoopBackOff   5          5m

Common Causes

  1. Application errors - Code bugs or exceptions
  2. Missing dependencies - Required files or libraries
  3. Configuration errors - Wrong environment variables
  4. Resource limits - Insufficient memory/CPU
  5. Health check failures - Liveness/readiness probe issues
  6. Database connection - Cannot connect to services
  7. File permissions - Cannot access required files
  8. Port conflicts - Port already in use
  9. Image pull errors - Cannot pull container image
  10. Startup commands - Invalid entrypoint or command

10 Solutions

Solution 1: Check Pod Logs

# Get recent logs
kubectl logs <pod-name>

# Get logs from previous attempts
kubectl logs <pod-name> --previous

# Follow logs in real-time
kubectl logs -f <pod-name>

# Get logs from specific container
kubectl logs <pod-name> -c <container-name>

What to look for:

  • Stack traces
  • Error messages
  • Exception details
  • Failed assertions

Solution 2: Describe Pod for Details

kubectl describe pod <pod-name>

Key sections to check:

  • Events: Recent pod events and warnings
  • State: Current container state
  • Last State: Previous container state
  • Reason: Exit reason code

Solution 3: Verify Container Image

# Check if image exists
docker pull <image-name>

# Test image locally
docker run -it <image-name> /bin/bash

# Check image pull policy
kubectl get pod <pod-name> -o yaml | grep imagePullPolicy

Common issues:

  • Wrong image tag
  • Private registry authentication
  • Image doesn't exist
  • Registry unavailable

Solution 4: Check Resource Limits

# Check pod resource usage
kubectl top pod <pod-name>

# View resource limits
kubectl describe pod <pod-name> | grep -A 5 "Limits\|Requests"

Adjust resources if needed:

resources:
  requests:
    memory: "256Mi"
    cpu: "250m"
  limits:
    memory: "512Mi"
    cpu: "500m"

Solution 5: Verify Environment Variables

# List environment variables
kubectl exec -it <pod-name> -- env

# Check specific variable
kubectl exec -it <pod-name> -- printenv VARIABLE_NAME

Common issues:

  • Missing required variables
  • Wrong variable values
  • Encoding issues with secrets

Solution 6: Test Health Probes

# Check probe configuration
kubectl describe pod <pod-name> | grep -A 10 "Liveness\|Readiness"

# Test endpoint manually
kubectl exec -it <pod-name> -- curl localhost:8080/health

Fix probe configuration:

livenessProbe:
  httpGet:
    path: /health
    port: 8080
  initialDelaySeconds: 30
  periodSeconds: 10
  timeoutSeconds: 5
  failureThreshold: 3

Solution 7: Check Database Connections

# Test database connectivity
kubectl exec -it <pod-name> -- nc -zv <db-host> 5432

# Check DNS resolution
kubectl exec -it <pod-name> -- nslookup <db-host>

Common fixes:

  • Verify service name
  • Check network policies
  • Ensure database is running
  • Verify credentials

Solution 8: Verify File Permissions

# Check file permissions
kubectl exec -it <pod-name> -- ls -la /path/to/file

# Test file access
kubectl exec -it <pod-name> -- cat /path/to/file

Fix permissions:

securityContext:
  runAsUser: 1000
  fsGroup: 1000

Solution 9: Check Port Configuration

# Listen on correct port
kubectl exec -it <pod-name> -- netstat -tlnp

# Verify port binding
kubectl exec -it <pod-name> -- ss -tlnp

Ensure containerPort matches:

ports:
- containerPort: 8080
  protocol: TCP

Solution 10: Validate Startup Command

# Check entrypoint and command
kubectl describe pod <pod-name> | grep -A 5 "Command\|Args"

# Test command manually
kubectl exec -it <pod-name> -- <your-command>

Common issues:

  • Invalid command syntax
  • Missing dependencies
  • Wrong working directory

Prevention Strategies

1. Use Init Containers

initContainers:
- name: check-db
  image: busybox
  command: ['sh', '-c', 'until nc -z db 5432; do echo waiting for db; sleep 2; done']

2. Implement Graceful Shutdown

lifecycle:
  preStop:
    exec:
      command: ["/bin/sh", "-c", "sleep 10"]

3. Add Resource Monitoring

# Install metrics-server
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

# Monitor resource usage
kubectl top pods

4. Use Probes Correctly

readinessProbe:
  httpGet:
    path: /ready
    port: 8080
  initialDelaySeconds: 5
  periodSeconds: 5

livenessProbe:
  httpGet:
    path: /health
    port: 8080
  initialDelaySeconds: 30
  periodSeconds: 10

5. Implement Retry Logic

env:
- name: RETRY_ATTEMPTS
  value: "5"
- name: RETRY_DELAY
  value: "5"

Advanced Debugging

Debug with Ephemeral Containers

# Add debug container
kubectl debug -it <pod-name> --image=nicolaka/netshoot --target=<container-name>

# Use debug tools
kubectl exec -it <pod-name> -c debug -- /bin/bash

Check Node Logs

# Get node name
kubectl get pod <pod-name> -o jsonpath='{.spec.nodeName}'

# Check node logs
kubectl logs -n kube-system -l k8s-app=kubelet <node-name>

Use kubectl Events

# Watch events
kubectl get events --sort-by='.lastTimestamp'

# Filter by pod
kubectl get events --field-selector involvedObject.name=<pod-name>

Network Debugging

# Test pod-to-pod connectivity
kubectl exec -it <pod-1> -- ping <pod-2-ip>

# Test service connectivity
kubectl exec -it <pod> -- curl <service-name>:<port>

Quick Reference

IssueCommand
Check logskubectl logs <pod>
Describe podkubectl describe pod <pod>
Execute commandkubectl exec -it <pod> -- <command>
Get eventskubectl get events
Check resourceskubectl top pod <pod>

Common Exit Codes

CodeMeaning
0Success
1General error
125Docker daemon error
126Command not executable
127Command not found
137Killed (SIGKILL)
139Segmentation fault
143Terminated (SIGTERM)

Best Practices

  1. Always check logs first - Most errors are logged
  2. Use descriptive pod names - Easier to identify
  3. Set appropriate resource limits - Prevent OOM kills
  4. Implement health checks - Catch issues early
  5. Use liveness and readiness probes - Separate concerns
  6. Monitor pod restarts - Set up alerts
  7. Test images locally - Before deploying
  8. Use init containers - For dependencies
  9. Implement graceful shutdown - For clean exits
  10. Document common issues - For faster resolution

Related Articles


Last Updated: October 2026
Author: CloudOpsGuide Team
Difficulty: Intermediate
Estimated Reading Time: 12 minutes