CloudOpsGuide
kubernetes

Kubernetes Resource Limits and Requests Explained

Intermediate
13 minutes
October 2026
CloudOpsGuide Team

Kubernetes Resource Limits and Requests Explained

Complete guide to CPU and memory requests/limits in Kubernetes — how they affect scheduling, throttling, and OOM kills.

Table of Contents

Requests vs Limits

resources:
  requests:
    memory: "256Mi"   # guaranteed — used for scheduling
    cpu: "250m"       # guaranteed CPU time
  limits:
    memory: "512Mi"   # hard cap — exceed = OOMKill
    cpu: "500m"       # hard cap — exceed = throttled
RequestsLimits
PurposeScheduling guaranteeMaximum allowed
Exceeding itImpossible (won't schedule)CPU = throttle, Memory = kill
UnsetPod is Burstable/BestEffortContainer can burst to node capacity

Units

  • CPU: 250m = 250 millicores = 0.25 of a CPU core. 1 = one full core.
  • Memory: Mi (mebibytes), Gi (gibibytes). 512Mi ≈ 536 MB.

How the Scheduler Uses Requests

The scheduler sums all container requests on a pod and only places it on a node with enough allocatable capacity:

Node capacity: 4 CPU, 16 GB
Already requested: 3.2 CPU, 12 GB
Your pod requests: 1 CPU, 4 GB
→ Won't schedule → pod stays Pending

This is the #1 cause of Pending pods — not insufficient actual usage, but insufficient requested capacity.

QoS Classes

Kubernetes assigns every pod a Quality of Service class that determines eviction priority:

Guaranteed (highest protection)

# requests == limits for CPU AND memory, on every container
resources:
  requests: { cpu: "500m", memory: "512Mi" }
  limits:   { cpu: "500m", memory: "512Mi" }

Burstable (most common)

# requests < limits
resources:
  requests: { cpu: "100m", memory: "128Mi" }
  limits:   { cpu: "1", memory: "1Gi" }

BestEffort (evicted first)

# No requests or limits at all — first to be evicted under pressure
# Avoid in production

Check a pod's class:

kubectl get pod <name> -o jsonpath='{.status.qosClass}'

CPU Throttling vs OOM Kills

Two very different failure modes:

CPU — Throttling (soft)

Exceeding CPU limit → the process is throttled (slowed down), not killed. Shows as latency spikes.

# Check throttling metrics
kubectl top pod <name>
# In Prometheus: container_cpu_cfs_throttled_periods_total

Memory — OOMKill (hard)

Exceeding memory limit → kernel kills the container → OOMKilled → CrashLoopBackOff if it keeps happening.

kubectl describe pod <name> | grep -i oom
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137

How to Choose Values

Step 1: Measure First

# Deploy without limits in dev, observe real usage
kubectl top pods -n dev

Or use VPA in recommendation mode:

kubectl apply -f https://github.com/kubernetes/autoscaler/releases/latest/download/vertical-pod-autoscaler.yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: my-app-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app
  updatePolicy:
    updateMode: "Off"   # recommend only, don't auto-apply
kubectl get vpa my-app-vpa -o yaml | grep -A 10 recommendation

Step 2: Apply the Common Formula

requests = observed p50 usage  (typical load)
limits   = observed p95–p99 × 1.2–1.5  (headroom for spikes)

Step 3: Iterate

Watch for:

  • Throttling → raise CPU limit
  • OOMKilled → raise memory limit (or fix the leak)
  • Pod Pending → lower requests or scale the node pool

LimitRange and Quotas

Default Limits per Namespace

apiVersion: v1
kind: LimitRange
metadata:
  name: limits
  namespace: dev
spec:
  limits:
  - type: Container
    default:
      cpu: "500m"
      memory: "512Mi"
    defaultRequest:
      cpu: "100m"
      memory: "128Mi"

Every container in dev without explicit resources gets these defaults.

Namespace Quotas

apiVersion: v1
kind: ResourceQuota
metadata:
  name: quota
  namespace: dev
spec:
  hard:
    requests.cpu: "8"
    requests.memory: "16Gi"
    limits.memory: "32Gi"
    pods: "50"

Common Mistakes

1. Limits without requests on CPU

# Subtle problem: guaranteed only 100m but throttled at 500m
requests: { cpu: "100m" }
limits:   { cpu: "500m" }

Under node contention this pod gets just 100m and chokes. For latency-sensitive apps, consider requests == limits (Guaranteed QoS) — no throttling, first-class scheduling.

2. Copy-pasting limits across different workloads

A Go microservice and a JVM monolith have wildly different memory profiles. Measure each.

3. Memory limit < JVM heap + overhead

JVM apps need: heap + metaspace + thread stacks + native memory. If Xmx=1g, the limit must be >1.5g or you'll OOMKill constantly.

4. requests > node capacity

# Pod requests 8 CPU on 4-CPU nodes → forever Pending
kubectl describe pod <name> | grep "Insufficient cpu"

Quick Reference

# Current usage
kubectl top pods --all-namespaces
kubectl top nodes

# What a pod requested
kubectl get pod <name> -o jsonpath='{.spec.containers[*].resources}'

# QoS class
kubectl get pods -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass

# OOMKilled pods
kubectl get pods --all-namespaces | grep -i oom
kubectl get events --field-selector reason=OOMKilled

Related Articles


Last Updated: October 2026
Author: CloudOpsGuide Team
Difficulty: Intermediate
Estimated Reading Time: 13 minutes