Kubernetes Resource Limits and Requests Explained
Kubernetes Resource Limits and Requests Explained
Complete guide to CPU and memory requests/limits in Kubernetes — how they affect scheduling, throttling, and OOM kills.
Table of Contents
- Requests vs Limits
- How the Scheduler Uses Requests
- QoS Classes
- CPU Throttling vs OOM Kills
- How to Choose Values
- LimitRange and Quotas
Requests vs Limits
resources:
requests:
memory: "256Mi" # guaranteed — used for scheduling
cpu: "250m" # guaranteed CPU time
limits:
memory: "512Mi" # hard cap — exceed = OOMKill
cpu: "500m" # hard cap — exceed = throttled
| Requests | Limits | |
|---|---|---|
| Purpose | Scheduling guarantee | Maximum allowed |
| Exceeding it | Impossible (won't schedule) | CPU = throttle, Memory = kill |
| Unset | Pod is Burstable/BestEffort | Container can burst to node capacity |
Units
- CPU:
250m= 250 millicores = 0.25 of a CPU core.1= one full core. - Memory:
Mi(mebibytes),Gi(gibibytes).512Mi≈ 536 MB.
How the Scheduler Uses Requests
The scheduler sums all container requests on a pod and only places it on a node with enough allocatable capacity:
Node capacity: 4 CPU, 16 GB
Already requested: 3.2 CPU, 12 GB
Your pod requests: 1 CPU, 4 GB
→ Won't schedule → pod stays Pending
This is the #1 cause of Pending pods — not insufficient actual usage, but insufficient requested capacity.
QoS Classes
Kubernetes assigns every pod a Quality of Service class that determines eviction priority:
Guaranteed (highest protection)
# requests == limits for CPU AND memory, on every container
resources:
requests: { cpu: "500m", memory: "512Mi" }
limits: { cpu: "500m", memory: "512Mi" }
Burstable (most common)
# requests < limits
resources:
requests: { cpu: "100m", memory: "128Mi" }
limits: { cpu: "1", memory: "1Gi" }
BestEffort (evicted first)
# No requests or limits at all — first to be evicted under pressure
# Avoid in production
Check a pod's class:
kubectl get pod <name> -o jsonpath='{.status.qosClass}'
CPU Throttling vs OOM Kills
Two very different failure modes:
CPU — Throttling (soft)
Exceeding CPU limit → the process is throttled (slowed down), not killed. Shows as latency spikes.
# Check throttling metrics
kubectl top pod <name>
# In Prometheus: container_cpu_cfs_throttled_periods_total
Memory — OOMKill (hard)
Exceeding memory limit → kernel kills the container → OOMKilled → CrashLoopBackOff if it keeps happening.
kubectl describe pod <name> | grep -i oom
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137
How to Choose Values
Step 1: Measure First
# Deploy without limits in dev, observe real usage
kubectl top pods -n dev
Or use VPA in recommendation mode:
kubectl apply -f https://github.com/kubernetes/autoscaler/releases/latest/download/vertical-pod-autoscaler.yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: my-app-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Off" # recommend only, don't auto-apply
kubectl get vpa my-app-vpa -o yaml | grep -A 10 recommendation
Step 2: Apply the Common Formula
requests = observed p50 usage (typical load)
limits = observed p95–p99 × 1.2–1.5 (headroom for spikes)
Step 3: Iterate
Watch for:
- Throttling → raise CPU limit
- OOMKilled → raise memory limit (or fix the leak)
- Pod Pending → lower requests or scale the node pool
LimitRange and Quotas
Default Limits per Namespace
apiVersion: v1
kind: LimitRange
metadata:
name: limits
namespace: dev
spec:
limits:
- type: Container
default:
cpu: "500m"
memory: "512Mi"
defaultRequest:
cpu: "100m"
memory: "128Mi"
Every container in dev without explicit resources gets these defaults.
Namespace Quotas
apiVersion: v1
kind: ResourceQuota
metadata:
name: quota
namespace: dev
spec:
hard:
requests.cpu: "8"
requests.memory: "16Gi"
limits.memory: "32Gi"
pods: "50"
Common Mistakes
1. Limits without requests on CPU
# Subtle problem: guaranteed only 100m but throttled at 500m
requests: { cpu: "100m" }
limits: { cpu: "500m" }
Under node contention this pod gets just 100m and chokes. For latency-sensitive apps, consider requests == limits (Guaranteed QoS) — no throttling, first-class scheduling.
2. Copy-pasting limits across different workloads
A Go microservice and a JVM monolith have wildly different memory profiles. Measure each.
3. Memory limit < JVM heap + overhead
JVM apps need: heap + metaspace + thread stacks + native memory. If Xmx=1g, the limit must be >1.5g or you'll OOMKill constantly.
4. requests > node capacity
# Pod requests 8 CPU on 4-CPU nodes → forever Pending
kubectl describe pod <name> | grep "Insufficient cpu"
Quick Reference
# Current usage
kubectl top pods --all-namespaces
kubectl top nodes
# What a pod requested
kubectl get pod <name> -o jsonpath='{.spec.containers[*].resources}'
# QoS class
kubectl get pods -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass
# OOMKilled pods
kubectl get pods --all-namespaces | grep -i oom
kubectl get events --field-selector reason=OOMKilled
Related Articles
Last Updated: October 2026
Author: CloudOpsGuide Team
Difficulty: Intermediate
Estimated Reading Time: 13 minutes