CloudOpsGuide
azure

Azure AKS Troubleshooting: Common Issues and Solutions

Intermediate
15 minutes
October 2026
CloudOpsGuide Team

Azure AKS Troubleshooting: Common Issues and Solutions

Diagnose and fix common Azure Kubernetes Service issues with kubectl commands and Azure CLI tools.

Table of Contents

Common AKS Issues

Issue 1: Cannot Connect to Cluster

Symptoms:

  • kubectl commands timeout
  • "Unable to connect to the server" error
  • Connection refused

Solutions:

# Get credentials
az aks get-credentials \
  --resource-group <resource-group> \
  --name <cluster-name>

# Verify connection
kubectl get nodes

# Check kubeconfig
kubectl config view

# Test with verbose output
kubectl get nodes --v=9

If still failing:

# Check if cluster is running
az aks show \
  --resource-group <resource-group> \
  --name <cluster-name> \
  --query provisioningState

# Restart cluster if needed
az aks start \
  --resource-group <resource-group> \
  --name <cluster-name>

Issue 2: Pods Stuck in Pending State

Symptoms:

  • Pods never start
  • Status remains "Pending"
  • No events showing

Solutions:

# Check pod events
kubectl describe pod <pod-name>

# Check node resources
kubectl top nodes

# Check for taints
kubectl describe nodes | grep -A 5 Taint

# Check quotas
kubectl describe quota

Common causes:

  1. Insufficient resources:
# Check resource requests
kubectl get pod <pod-name> -o yaml | grep -A 10 resources

# Scale up node pool
az aks nodepool scale \
  --resource-group <resource-group> \
  --cluster-name <cluster-name> \
  --name <nodepool-name> \
  --node-count 5
  1. Tolerations needed:
spec:
  tolerations:
  - key: "dedicated"
    operator: "Equal"
    value: "special"
    effect: "NoSchedule"

Issue 3: CrashLoopBackOff

Symptoms:

  • Pod keeps restarting
  • Status: CrashLoopBackOff
  • High restart count

Solutions:

# Check logs
kubectl logs <pod-name>

# Check previous logs
kubectl logs <pod-name> --previous

# Describe pod
kubectl describe pod <pod-name>

# Check events
kubectl get events --sort-by='.lastTimestamp'

Common fixes:

  1. Resource limits:
resources:
  requests:
    memory: "256Mi"
    cpu: "250m"
  limits:
    memory: "512Mi"
    cpu: "500m"
  1. Health probes:
livenessProbe:
  httpGet:
    path: /health
    port: 8080
  initialDelaySeconds: 30
  periodSeconds: 10

Connectivity Issues

Issue 4: Cannot Access Services

Symptoms:

  • Service not accessible
  • Connection timeout
  • DNS resolution fails

Solutions:

# Check service
kubectl get svc <service-name>

# Check service endpoints
kubectl get endpoints <service-name>

# Test from pod
kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never
nslookup <service-name>
curl http://<service-name>:<port>

DNS Issues:

# Check CoreDNS
kubectl get pods -n kube-system -l k8s-app=kube-dns

# Check CoreDNS logs
kubectl logs -n kube-system -l k8s-app=kube-dns

# Restart CoreDNS
kubectl rollout restart deployment/coredns -n kube-system

Issue 5: Ingress Not Working

Symptoms:

  • Ingress rules not routing
  • 502/503 errors
  • Certificate issues

Solutions:

# Check ingress
kubectl get ingress

# Check ingress controller
kubectl get pods -n ingress-nginx

# Check ingress logs
kubectl logs -n ingress-nginx -l app.kubernetes.io/name=ingress-nginx

# Test ingress
kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never
curl -H "Host: example.com" http://<ingress-ip>

Common fixes:

  1. Missing annotations:
annotations:
  nginx.ingress.kubernetes.io/rewrite-target: /
  nginx.ingress.kubernetes.io/ssl-redirect: "true"
  1. Wrong service name:
backend:
  service:
    name: correct-service-name
    port:
      number: 80

Node Issues

Issue 6: Nodes Not Ready

Symptoms:

  • Nodes in NotReady state
  • Nodes cordoned
  • Nodes unreachable

Solutions:

# Check node status
kubectl get nodes

# Describe node
kubectl describe node <node-name>

# Check node conditions
kubectl get nodes -o yaml | grep -A 5 conditions

# Uncordon node
kubectl uncordon <node-name>

Common causes:

  1. Disk pressure:
# Check disk usage
kubectl exec -it <node-name> -- df -h

# Clean up
kubectl delete pods --field-selector=status.phase=Succeeded
  1. Memory pressure:
# Check memory
kubectl top nodes

# Evict pods
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data

Issue 7: Node Pool Scaling Issues

Symptoms:

  • Node pool not scaling
  • Scale timeout
  • Insufficient capacity

Solutions:

# Check node pool status
az aks nodepool list \
  --resource-group <resource-group> \
  --cluster-name <cluster-name>

# Check autoscaler
kubectl get configmap -n kube-system cluster-autoscaler-status

# Manual scale
az aks nodepool scale \
  --resource-group <resource-group> \
  --cluster-name <cluster-name> \
  --name <nodepool-name> \
  --node-count 5

Enable cluster autoscaler:

az aks nodepool update \
  --resource-group <resource-group> \
  --cluster-name <cluster-name> \
  --name <nodepool-name> \
  --enable-cluster-autoscaler \
  --min-count 3 \
  --max-count 10

Networking Issues

Issue 8: Pod-to-Pod Communication Fails

Symptoms:

  • Cannot connect between pods
  • Network policy blocking
  • CNI issues

Solutions:

# Check network policies
kubectl get networkpolicies --all-namespaces

# Test connectivity
kubectl run -it --rm debug --image=nicolaka/netshoot --restart=Never
ping <target-pod-ip>
nc -zv <target-pod-ip> <port>

# Check CNI pods
kubectl get pods -n kube-system -l k8s-app=azure-cni

Azure CNI specific:

# Check Azure CNI version
az aks show \
  --resource-group <resource-group> \
  --name <cluster-name> \
  --query networkProfile.networkPlugin

# Restart Azure CNI
kubectl delete pods -n kube-system -l k8s-app=azure-cni

Issue 9: External Access Blocked

Symptoms:

  • Cannot access from outside
  • LoadBalancer issues
  • NAT problems

Solutions:

# Check LoadBalancer
kubectl get svc

# Check Azure load balancer
az network lb list --resource-group <resource-group>

# Check NSG rules
az network nsg list --resource-group <resource-group>

# Test from external
curl http://<external-ip>:<port>

Storage Issues

Issue 10: Persistent Volume Claims Pending

Symptoms:

  • PVC stuck in Pending
  • Cannot mount volumes
  • Storage class issues

Solutions:

# Check PVC
kubectl get pvc

# Describe PVC
kubectl describe pvc <pvc-name>

# Check storage classes
kubectl get storageclass

# Check PV
kubectl get pv

Common fixes:

  1. Wrong storage class:
spec:
  storageClassName: "managed-premium"
  accessModes: [ "ReadWriteOnce" ]
  resources:
    requests:
      storage: 10Gi
  1. Azure disk issues:
# Check Azure disks
az disk list --resource-group <resource-group>

# Check disk attachment
az disk show --name <disk-name> --resource-group <resource-group>

Performance Issues

Issue 11: High CPU/Memory Usage

Symptoms:

  • Cluster slow
  • High resource usage
  • Pod evictions

Solutions:

# Check resource usage
kubectl top nodes
kubectl top pods

# Check resource limits
kubectl get pods -o jsonpath='{.items[*].spec.containers[*].resources}'

# Identify resource hogs
kubectl top pods --sort-by=cpu
kubectl top pods --sort-by=memory

Optimization:

  1. Set appropriate limits:
resources:
  requests:
    memory: "256Mi"
    cpu: "250m"
  limits:
    memory: "512Mi"
    cpu: "500m"
  1. Enable HPA:
kubectl autoscale deployment <deployment-name> \
  --cpu-percent=70 \
  --min=3 \
  --max=10

Issue 12: Slow Pod Startup

Symptoms:

  • Pods take long to start
  • Image pull delays
  • Init container issues

Solutions:

# Check image pull
kubectl describe pod <pod-name> | grep -A 5 ImagePull

# Check init containers
kubectl describe pod <pod-name> | grep -A 10 Init

# Use image pull policy
imagePullPolicy: IfNotPresent

Optimization:

  1. Use ACR caching:
az acr import \
  --name <acr-name> \
  --source <external-image> \
  --image <image-name>:tag
  1. Pre-pull images:
initContainers:
- name: image-puller
  image: busybox
  command: ['sh', '-c', 'docker pull myimage:latest']

Azure Specific Issues

Issue 13: Azure Authentication

Symptoms:

  • Cannot authenticate
  • Token expired
  • Permission denied

Solutions:

# Refresh credentials
az login
az aks get-credentials \
  --resource-group <resource-group> \
  --name <cluster-name>

# Check token
kubectl config view --raw

# Check Azure AD integration
az aks show \
  --resource-group <resource-group> \
  --name <cluster-name> \
  --query aadProfile

Issue 14: Azure Disk Attachment Fails

Symptoms:

  • Disk not attaching
  • Attachment timeout
  • I/O errors

Solutions:

# Check disk status
az disk show --name <disk-name> --resource-group <resource-group>

# Detach and reattach
az disk detach --name <disk-name> --resource-group <resource-group>

# Check disk quotas
az disk list --resource-group <resource-group} --query "[].diskSizeGb"

Quick Reference

IssueCommand
Cluster connectionaz aks get-credentials
Pod statuskubectl get pods
Pod logskubectl logs <pod>
Describe podkubectl describe pod <pod>
Node statuskubectl get nodes
Service statuskubectl get svc
Eventskubectl get events
Resource usagekubectl top pods

Monitoring Tools

Azure Monitor

# Enable Container Insights
az aks enable-addons \
  --resource-group <resource-group> \
  --name <cluster-name> \
  --addons monitoring \
  --workspace-resource-id <workspace-id>

Kubectl Plugins

# Install krew
kubectl krew install krew

# Install useful plugins
kubectl krew install get-all
kubectl krew install df-pv
kubectl krew install node-shell

Best Practices

  1. Monitor regularly - Set up alerts
  2. Use resource limits - Prevent resource exhaustion
  3. Implement health checks - Catch issues early
  4. Use node selectors - Control pod placement
  5. Enable autoscaling - Handle load changes
  6. Regular updates - Keep cluster patched
  7. Backup configurations - Use GitOps
  8. Document issues - Build knowledge base
  9. Test in staging - Before production
  10. Use managed services - Reduce complexity

Related Articles


Last Updated: October 2026
Author: CloudOpsGuide Team
Difficulty: Intermediate
Estimated Reading Time: 15 minutes