Infrastructure & Operations › Kubernetes & Orchestration
Requests and Limits
How much CPU and memory a pod reserves and is allowed to use.
Also known as: requests and limits, resource requests, resource limits
Requests and limits are how you tell Kubernetes how much CPU and memory a container needs. They look similar but do different jobs, and mixing them up causes some of the most common cluster problems.
- A request is what the container is guaranteed and what the scheduler uses to find a node with room. Two pods requesting 1 CPU each need a node with 2 CPUs free.
- A limit is the ceiling the container may not exceed. CPU above the limit is throttled; memory above the limit gets the container OOMKilled.
resources:
requests: { cpu: "250m", memory: "256Mi" }
limits: { cpu: "1", memory: "512Mi" }
From these, Kubernetes assigns a QoS class:
- Guaranteed — requests equal limits on every container; the most protected.
- Burstable — requests below limits; can use spare capacity but is first in line to be squeezed.
- BestEffort — no requests or limits; the first to be evicted under pressure.
The classic mistakes:
- Setting no requests. The scheduler can’t place pods well, nodes get overcommitted, and autoscalers that compute utilization from requests can’t work. Requests are effectively mandatory for anything real.
- A memory limit too low. The container is killed — usually
OOMKilled— under load it would have survived without the cap. Memory limits protect the node, so set them sensibly above real peak. - A CPU limit that strangles. CPU limits cause throttling even when the node has idle capacity, which hurts latency-sensitive services. Many teams set CPU requests but leave CPU limits off, to avoid artificial throttling.
- Requests far above real usage. Space is reserved and never used, wasting node capacity and money (see capacity planning).
- Guessing once and never revisiting. Real usage changes. Measure and adjust, or let the Vertical Pod Autoscaler recommend values.
Requests and limits underpin almost everything else — scheduling, autoscaling, quotas and priority. Getting them roughly right is the difference between a cluster that runs well and one that’s simultaneously throttled and under-utilised.