azure
Azure AKS Troubleshooting: Common Issues and Solutions
Intermediate
15 minutes
October 2026
CloudOpsGuide Team
Azure AKS Troubleshooting: Common Issues and Solutions
Diagnose and fix common Azure Kubernetes Service issues with kubectl commands and Azure CLI tools.
Table of Contents
- Common AKS Issues
- Connectivity Issues
- Pod Issues
- Node Issues
- Networking Issues
- Storage Issues
- Performance Issues
- Azure Specific Issues
Common AKS Issues
Issue 1: Cannot Connect to Cluster
Symptoms:
kubectlcommands timeout- "Unable to connect to the server" error
- Connection refused
Solutions:
# Get credentials
az aks get-credentials \
--resource-group <resource-group> \
--name <cluster-name>
# Verify connection
kubectl get nodes
# Check kubeconfig
kubectl config view
# Test with verbose output
kubectl get nodes --v=9
If still failing:
# Check if cluster is running
az aks show \
--resource-group <resource-group> \
--name <cluster-name> \
--query provisioningState
# Restart cluster if needed
az aks start \
--resource-group <resource-group> \
--name <cluster-name>
Issue 2: Pods Stuck in Pending State
Symptoms:
- Pods never start
- Status remains "Pending"
- No events showing
Solutions:
# Check pod events
kubectl describe pod <pod-name>
# Check node resources
kubectl top nodes
# Check for taints
kubectl describe nodes | grep -A 5 Taint
# Check quotas
kubectl describe quota
Common causes:
- Insufficient resources:
# Check resource requests
kubectl get pod <pod-name> -o yaml | grep -A 10 resources
# Scale up node pool
az aks nodepool scale \
--resource-group <resource-group> \
--cluster-name <cluster-name> \
--name <nodepool-name> \
--node-count 5
- Tolerations needed:
spec:
tolerations:
- key: "dedicated"
operator: "Equal"
value: "special"
effect: "NoSchedule"
Issue 3: CrashLoopBackOff
Symptoms:
- Pod keeps restarting
- Status: CrashLoopBackOff
- High restart count
Solutions:
# Check logs
kubectl logs <pod-name>
# Check previous logs
kubectl logs <pod-name> --previous
# Describe pod
kubectl describe pod <pod-name>
# Check events
kubectl get events --sort-by='.lastTimestamp'
Common fixes:
- Resource limits:
resources:
requests:
memory: "256Mi"
cpu: "250m"
limits:
memory: "512Mi"
cpu: "500m"
- Health probes:
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
Connectivity Issues
Issue 4: Cannot Access Services
Symptoms:
- Service not accessible
- Connection timeout
- DNS resolution fails
Solutions:
# Check service
kubectl get svc <service-name>
# Check service endpoints
kubectl get endpoints <service-name>
# Test from pod
kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never
nslookup <service-name>
curl http://<service-name>:<port>
DNS Issues:
# Check CoreDNS
kubectl get pods -n kube-system -l k8s-app=kube-dns
# Check CoreDNS logs
kubectl logs -n kube-system -l k8s-app=kube-dns
# Restart CoreDNS
kubectl rollout restart deployment/coredns -n kube-system
Issue 5: Ingress Not Working
Symptoms:
- Ingress rules not routing
- 502/503 errors
- Certificate issues
Solutions:
# Check ingress
kubectl get ingress
# Check ingress controller
kubectl get pods -n ingress-nginx
# Check ingress logs
kubectl logs -n ingress-nginx -l app.kubernetes.io/name=ingress-nginx
# Test ingress
kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never
curl -H "Host: example.com" http://<ingress-ip>
Common fixes:
- Missing annotations:
annotations:
nginx.ingress.kubernetes.io/rewrite-target: /
nginx.ingress.kubernetes.io/ssl-redirect: "true"
- Wrong service name:
backend:
service:
name: correct-service-name
port:
number: 80
Node Issues
Issue 6: Nodes Not Ready
Symptoms:
- Nodes in NotReady state
- Nodes cordoned
- Nodes unreachable
Solutions:
# Check node status
kubectl get nodes
# Describe node
kubectl describe node <node-name>
# Check node conditions
kubectl get nodes -o yaml | grep -A 5 conditions
# Uncordon node
kubectl uncordon <node-name>
Common causes:
- Disk pressure:
# Check disk usage
kubectl exec -it <node-name> -- df -h
# Clean up
kubectl delete pods --field-selector=status.phase=Succeeded
- Memory pressure:
# Check memory
kubectl top nodes
# Evict pods
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data
Issue 7: Node Pool Scaling Issues
Symptoms:
- Node pool not scaling
- Scale timeout
- Insufficient capacity
Solutions:
# Check node pool status
az aks nodepool list \
--resource-group <resource-group> \
--cluster-name <cluster-name>
# Check autoscaler
kubectl get configmap -n kube-system cluster-autoscaler-status
# Manual scale
az aks nodepool scale \
--resource-group <resource-group> \
--cluster-name <cluster-name> \
--name <nodepool-name> \
--node-count 5
Enable cluster autoscaler:
az aks nodepool update \
--resource-group <resource-group> \
--cluster-name <cluster-name> \
--name <nodepool-name> \
--enable-cluster-autoscaler \
--min-count 3 \
--max-count 10
Networking Issues
Issue 8: Pod-to-Pod Communication Fails
Symptoms:
- Cannot connect between pods
- Network policy blocking
- CNI issues
Solutions:
# Check network policies
kubectl get networkpolicies --all-namespaces
# Test connectivity
kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never
ping <target-pod-ip>
nc -zv <target-pod-ip> <port>
# Check CNI pods
kubectl get pods -n kube-system -l k8s-app=azure-cni
Azure CNI specific:
# Check Azure CNI version
az aks show \
--resource-group <resource-group> \
--name <cluster-name> \
--query networkProfile.networkPlugin
# Restart Azure CNI
kubectl delete pods -n kube-system -l k8s-app=azure-cni
Issue 9: External Access Blocked
Symptoms:
- Cannot access from outside
- LoadBalancer issues
- NAT problems
Solutions:
# Check LoadBalancer
kubectl get svc
# Check Azure load balancer
az network lb list --resource-group <resource-group>
# Check NSG rules
az network nsg list --resource-group <resource-group>
# Test from external
curl http://<external-ip>:<port>
Storage Issues
Issue 10: Persistent Volume Claims Pending
Symptoms:
- PVC stuck in Pending
- Cannot mount volumes
- Storage class issues
Solutions:
# Check PVC
kubectl get pvc
# Describe PVC
kubectl describe pvc <pvc-name>
# Check storage classes
kubectl get storageclass
# Check PV
kubectl get pv
Common fixes:
- Wrong storage class:
spec:
storageClassName: "managed-premium"
accessModes: [ "ReadWriteOnce" ]
resources:
requests:
storage: 10Gi
- Azure disk issues:
# Check Azure disks
az disk list --resource-group <resource-group>
# Check disk attachment
az disk show --name <disk-name> --resource-group <resource-group>
Performance Issues
Issue 11: High CPU/Memory Usage
Symptoms:
- Cluster slow
- High resource usage
- Pod evictions
Solutions:
# Check resource usage
kubectl top nodes
kubectl top pods
# Check resource limits
kubectl get pods -o jsonpath='{.items[*].spec.containers[*].resources}'
# Identify resource hogs
kubectl top pods --sort-by=cpu
kubectl top pods --sort-by=memory
Optimization:
- Set appropriate limits:
resources:
requests:
memory: "256Mi"
cpu: "250m"
limits:
memory: "512Mi"
cpu: "500m"
- Enable HPA:
kubectl autoscale deployment <deployment-name> \
--cpu-percent=70 \
--min=3 \
--max=10
Issue 12: Slow Pod Startup
Symptoms:
- Pods take long to start
- Image pull delays
- Init container issues
Solutions:
# Check image pull
kubectl describe pod <pod-name> | grep -A 5 ImagePull
# Check init containers
kubectl describe pod <pod-name> | grep -A 10 Init
# Use image pull policy
imagePullPolicy: IfNotPresent
Optimization:
- Use ACR caching:
az acr import \
--name <acr-name> \
--source <external-image> \
--image <image-name>:tag
- Pre-pull images:
initContainers:
- name: image-puller
image: busybox
command: ['sh', '-c', 'docker pull myimage:latest']
Azure Specific Issues
Issue 13: Azure Authentication
Symptoms:
- Cannot authenticate
- Token expired
- Permission denied
Solutions:
# Refresh credentials
az login
az aks get-credentials \
--resource-group <resource-group> \
--name <cluster-name>
# Check token
kubectl config view --raw
# Check Azure AD integration
az aks show \
--resource-group <resource-group> \
--name <cluster-name> \
--query aadProfile
Issue 14: Azure Disk Attachment Fails
Symptoms:
- Disk not attaching
- Attachment timeout
- I/O errors
Solutions:
# Check disk status
az disk show --name <disk-name> --resource-group <resource-group>
# Detach and reattach
az disk detach --name <disk-name> --resource-group <resource-group>
# Check disk quotas
az disk list --resource-group <resource-group} --query "[].diskSizeGb"
Quick Reference
| Issue | Command |
|---|---|
| Cluster connection | az aks get-credentials |
| Pod status | kubectl get pods |
| Pod logs | kubectl logs <pod> |
| Describe pod | kubectl describe pod <pod> |
| Node status | kubectl get nodes |
| Service status | kubectl get svc |
| Events | kubectl get events |
| Resource usage | kubectl top pods |
Monitoring Tools
Azure Monitor
# Enable Container Insights
az aks enable-addons \
--resource-group <resource-group> \
--name <cluster-name> \
--addons monitoring \
--workspace-resource-id <workspace-id>
Kubectl Plugins
# Install krew
kubectl krew install krew
# Install useful plugins
kubectl krew install get-all
kubectl krew install df-pv
kubectl krew install node-shell
Best Practices
- Monitor regularly - Set up alerts
- Use resource limits - Prevent resource exhaustion
- Implement health checks - Catch issues early
- Use node selectors - Control pod placement
- Enable autoscaling - Handle load changes
- Regular updates - Keep cluster patched
- Backup configurations - Use GitOps
- Document issues - Build knowledge base
- Test in staging - Before production
- Use managed services - Reduce complexity
Related Articles
Last Updated: October 2026
Author: CloudOpsGuide Team
Difficulty: Intermediate
Estimated Reading Time: 15 minutes