kubernetes
Fix Kubernetes CrashLoopBackOff: 10 Common Solutions
Intermediate
12 minutes
October 2026
CloudOpsGuide Team
Fix Kubernetes CrashLoopBackOff: 10 Common Solutions
Comprehensive troubleshooting guide for CrashLoopBackOff errors with real-world solutions and preventive measures.
Table of Contents
Understanding CrashLoopBackOff
CrashLoopBackOff occurs when a container repeatedly crashes and Kubernetes restarts it, but it continues to fail.
Lifecycle States
- Pending: Pod is waiting to be scheduled
- Running: Container is running
- Succeeded: Container exited successfully
- Failed: Container exited with error
- CrashLoopBackOff: Container keeps crashing
When It Happens
# Typical pod status
NAME READY STATUS RESTARTS AGE
my-app-5d7f9d6d8b-abc123 0/1 CrashLoopBackOff 5 5m
Common Causes
- Application errors - Code bugs or exceptions
- Missing dependencies - Required files or libraries
- Configuration errors - Wrong environment variables
- Resource limits - Insufficient memory/CPU
- Health check failures - Liveness/readiness probe issues
- Database connection - Cannot connect to services
- File permissions - Cannot access required files
- Port conflicts - Port already in use
- Image pull errors - Cannot pull container image
- Startup commands - Invalid entrypoint or command
10 Solutions
Solution 1: Check Pod Logs
# Get recent logs
kubectl logs <pod-name>
# Get logs from previous attempts
kubectl logs <pod-name> --previous
# Follow logs in real-time
kubectl logs -f <pod-name>
# Get logs from specific container
kubectl logs <pod-name> -c <container-name>
What to look for:
- Stack traces
- Error messages
- Exception details
- Failed assertions
Solution 2: Describe Pod for Details
kubectl describe pod <pod-name>
Key sections to check:
- Events: Recent pod events and warnings
- State: Current container state
- Last State: Previous container state
- Reason: Exit reason code
Solution 3: Verify Container Image
# Check if image exists
docker pull <image-name>
# Test image locally
docker run -it <image-name> /bin/bash
# Check image pull policy
kubectl get pod <pod-name> -o yaml | grep imagePullPolicy
Common issues:
- Wrong image tag
- Private registry authentication
- Image doesn't exist
- Registry unavailable
Solution 4: Check Resource Limits
# Check pod resource usage
kubectl top pod <pod-name>
# View resource limits
kubectl describe pod <pod-name> | grep -A 5 "Limits\|Requests"
Adjust resources if needed:
resources:
requests:
memory: "256Mi"
cpu: "250m"
limits:
memory: "512Mi"
cpu: "500m"
Solution 5: Verify Environment Variables
# List environment variables
kubectl exec -it <pod-name> -- env
# Check specific variable
kubectl exec -it <pod-name> -- printenv VARIABLE_NAME
Common issues:
- Missing required variables
- Wrong variable values
- Encoding issues with secrets
Solution 6: Test Health Probes
# Check probe configuration
kubectl describe pod <pod-name> | grep -A 10 "Liveness\|Readiness"
# Test endpoint manually
kubectl exec -it <pod-name> -- curl localhost:8080/health
Fix probe configuration:
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
Solution 7: Check Database Connections
# Test database connectivity
kubectl exec -it <pod-name> -- nc -zv <db-host> 5432
# Check DNS resolution
kubectl exec -it <pod-name> -- nslookup <db-host>
Common fixes:
- Verify service name
- Check network policies
- Ensure database is running
- Verify credentials
Solution 8: Verify File Permissions
# Check file permissions
kubectl exec -it <pod-name> -- ls -la /path/to/file
# Test file access
kubectl exec -it <pod-name> -- cat /path/to/file
Fix permissions:
securityContext:
runAsUser: 1000
fsGroup: 1000
Solution 9: Check Port Configuration
# Listen on correct port
kubectl exec -it <pod-name> -- netstat -tlnp
# Verify port binding
kubectl exec -it <pod-name> -- ss -tlnp
Ensure containerPort matches:
ports:
- containerPort: 8080
protocol: TCP
Solution 10: Validate Startup Command
# Check entrypoint and command
kubectl describe pod <pod-name> | grep -A 5 "Command\|Args"
# Test command manually
kubectl exec -it <pod-name> -- <your-command>
Common issues:
- Invalid command syntax
- Missing dependencies
- Wrong working directory
Prevention Strategies
1. Use Init Containers
initContainers:
- name: check-db
image: busybox
command: ['sh', '-c', 'until nc -z db 5432; do echo waiting for db; sleep 2; done']
2. Implement Graceful Shutdown
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 10"]
3. Add Resource Monitoring
# Install metrics-server
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
# Monitor resource usage
kubectl top pods
4. Use Probes Correctly
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
5. Implement Retry Logic
env:
- name: RETRY_ATTEMPTS
value: "5"
- name: RETRY_DELAY
value: "5"
Advanced Debugging
Debug with Ephemeral Containers
# Add debug container
kubectl debug -it <pod-name> --image=nicolaka/netshoot --target=<container-name>
# Use debug tools
kubectl exec -it <pod-name> -c debug -- /bin/bash
Check Node Logs
# Get node name
kubectl get pod <pod-name> -o jsonpath='{.spec.nodeName}'
# Check node logs
kubectl logs -n kube-system -l k8s-app=kubelet <node-name>
Use kubectl Events
# Watch events
kubectl get events --sort-by='.lastTimestamp'
# Filter by pod
kubectl get events --field-selector involvedObject.name=<pod-name>
Network Debugging
# Test pod-to-pod connectivity
kubectl exec -it <pod-1> -- ping <pod-2-ip>
# Test service connectivity
kubectl exec -it <pod> -- curl <service-name>:<port>
Quick Reference
| Issue | Command |
|---|---|
| Check logs | kubectl logs <pod> |
| Describe pod | kubectl describe pod <pod> |
| Execute command | kubectl exec -it <pod> -- <command> |
| Get events | kubectl get events |
| Check resources | kubectl top pod <pod> |
Common Exit Codes
| Code | Meaning |
|---|---|
| 0 | Success |
| 1 | General error |
| 125 | Docker daemon error |
| 126 | Command not executable |
| 127 | Command not found |
| 137 | Killed (SIGKILL) |
| 139 | Segmentation fault |
| 143 | Terminated (SIGTERM) |
Best Practices
- Always check logs first - Most errors are logged
- Use descriptive pod names - Easier to identify
- Set appropriate resource limits - Prevent OOM kills
- Implement health checks - Catch issues early
- Use liveness and readiness probes - Separate concerns
- Monitor pod restarts - Set up alerts
- Test images locally - Before deploying
- Use init containers - For dependencies
- Implement graceful shutdown - For clean exits
- Document common issues - For faster resolution
Related Articles
- Kubernetes Pod Not Ready: Debugging Guide
- Kubernetes Resource Limits and Requests
- Kubernetes Liveness and Readiness Probes
Last Updated: October 2026
Author: CloudOpsGuide Team
Difficulty: Intermediate
Estimated Reading Time: 12 minutes