What requests and limits actually control
requests.cpu is what the scheduler uses: it places the pod on a node with enough allocatable CPU and acts as the weight when the node is oversubscribed. requests.memory works the same way, but memory cannot be compressed, so it is also the real occupancy counted against node memory pressure.
limits.cpu is a hard ceiling, yet exceeding it does not kill the container. CFS throttling kicks in: the process keeps running but only receives the CPU time granted for its quota within each 100ms period, and the excess is descheduled. Guaranteed share only exists when requests equal limits.
limits.memory is an impassable wall. The moment container memory reaches the limit, the kernel OOM Killer terminates the process, the pod restarts, and the exit code is 137. This differs fundamentally from CPU, which is why you cannot paper over a memory leak by nudging the limit upward.
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: 500m
memory: 1GiThe three QoS classes and eviction priority
QoS is computed by the kubelet, not declared by you. Guaranteed requires every container to have CPU and memory limits equal to requests; a single missing or unequal value makes the pod Burstable; setting nothing makes it BestEffort.
Under node memory pressure, eviction proceeds in reverse QoS order: BestEffort goes first, Burstable is ordered by how far actual usage exceeds requests, and Guaranteed goes last. This is why marking a critical service as Guaranteed survives pressure while a pod with no requests is always the first to be pushed out.
The most misunderstood setting is CPU with limits equal to requests: the container does hold that quota exclusively, but if the app only uses 100m the remaining 400m sits idle and wasted. Decide based on node pressure events and real container CPU usage rather than intuition.
apiVersion: v1
kind: LimitRange
metadata:
name: default-container-limits
spec:
limits:
- type: Container
default:
cpu: 500m
memory: 512Mi
defaultRequest:
cpu: 100m
memory: 128Mi
max:
cpu: "4"
memory: 8Gi
min:
cpu: 50m
memory: 64MiAttributing an OOMKilled container correctly
First separate the two kinds of OOM. A container OOMKilled means the container exceeded its memory limit; the reason field reads OOMKilled and the exit code is 137. A node OOM is the host kernel killing a process, usually logged as Out of memory: Killed process, and other pods on the node may be hit at the same time.
The standard path for a container OOM is: `kubectl describe pod` to confirm reason OOMKilled, then `kubectl logs --previous` for business logs just before the crash, and finally check whether the container has **no** memory limit at all. A pod with no limit that gets evicted under node pressure is also killed by the kernel and is easily misread as OOMKilled.
The correct way to set a memory ceiling is to profile first and then choose the number. If the cache is simply too large, express the bound in application settings such as -XX:MaxRAMPercentage or an explicit heap cap, rather than pushing limits.memory up until the OOM can never happen, which only exports the failure to the node.
kubectl get pod <pod> -o jsonpath="{.status.containerStatuses[*].lastState.terminated.reason}"
kubectl describe pod <pod> | grep -A5 "Last State"
kubectl logs <pod> --previous --tail=200
kubectl top pod <pod> --containersObserving CPU throttling
Problems caused by a CPU limit never show up as restarts, only as rising latency, which makes them easy to miss. The reliable signal is the ratio of container_cpu_cfs_throttled_seconds_total to container_cpu_cfs_periods_total.
Query that ratio in Prometheus: a persistently high share of throttled periods means the quota is insufficient. Typical culprits are bursty compute during startup, batch jobs, and GC-heavy applications, whose P99 latency gets stretched by whole throttling periods.
A workable adjustment order: first compare requests against real usage to confirm scheduling waste, then decide whether to raise limits.cpu or optimize the code. Remember that raising limits lowers the QoS guarantee and increases node pressure, so prefer explicit Guaranteed for latency-sensitive load and looser limits for throughput-oriented background work.
Remember the resource semantics of init containers. They run sequentially and their peak memory is accounted separately, and their requirements add into the effective pod request, so an oversized init spec pushes the whole pod onto a larger node. Multiple init containers are combined by taking the maximum rather than the sum, but the interaction with the app containers still has to be estimated from real peaks.
container_cpu_cfs_throttled_seconds_total / container_cpu_cfs_periods_total
kubectl describe node <node> | sed -n "/Allocated resources:/,/Events:/p"Deriving requests from monitoring data
A sound requests value comes from measurement, not intuition. Use the P95 or P99 of CPU usage over a representative window as the starting point: most of the time no throttling occurs, yet you avoid reserving so much for rare spikes that nodes waste capacity.
Memory must be sized on peak rather than average. Memory is not reclaimable, so limits.memory should sit 1.3 to 1.5 times above the observed peak to absorb GC swings and request bursts. Deriving requests from average memory gets pods OOMKilled on bursts, while setting limits to average memory causes constant reclamation under load.
Validate the numbers with a pressure test afterwards: deploy beyond the total requests of the target nodes and see whether the pods still schedule. If the node cannot even fit the requests, the scheduler keeps pods Pending and the real fix is lowering the granularity to container level. If they fit but the node is still saturated, the cluster needs more nodes rather than different numbers on a single pod.
# 用 metrics-server 抓当前用量分布,再据此确定 requests
kubectl top pod -l app=api --containers
# 确认调度可行性:Pending 通常是节点 requests 装不下
kubectl get pods -l app=api -o jsonpath="{.items[*].status.conditions[?(@.type=='PodScheduled')].status}"
kubectl describe pod <pending-pod> | sed -n "/Events:/,$ p"