Three layers of automation: who decides replica count, pod size, and node count
Separate these three jobs and most confusion dissolves. HPA changes replica count — whether a Deployment runs 3 pods or 30. VPA changes the requests/limits of an individual container. Cluster Autoscaler or Karpenter changes node count. They act at different layers and in principle compose, but they read overlapping signals, which is where the conflict originates.
HPA is reactive: it adjusts future replica count based on current load. VPA is corrective: it observes historical usage and moves requests closer to reality. Cluster Autoscaler is also reactive, but its decision includes pod requests — it asks "can the current nodes still fit these pods" — and pod sizing is exactly what VPA mutates.
A useful counter-intuition: shrinking requests makes pods easier to place on existing nodes, which makes Cluster Autoscaler conclude no new node is needed, so it stops provisioning. VPA saving resources can therefore quietly suppress cluster growth; raising requests does the opposite. Compose all three and behaviour routinely exceeds what any single component implies.
The pragmatic path is layered adoption: use HPA for "not enough replicas", keep requests fixed for scheduling sanity, and introduce VPA only where resource profiles genuinely differ, with its blast radius constrained (for example auto mode limited to non-critical services).
# 现状盘点
kubectl get hpa -A
kubectl top pods -A --sort-by=cpu | head -20 # 需要 metrics-server
kubectl get vpa -A 2>/dev/null
# 节点层面:requests 总量 vs 实际分配量
kubectl describe nodes | grep -A6 "Allocated resources:"What HPA measures, and why metrics-server is a hard dependency
The most common HPA uses `type: Resource` with CPU or memory. "Utilisation" here is not absolute usage but a ratio against requests: `usage / requests`. That gives requests a second job — they are both the scheduling input and the baseline for scaling math. Overly generous requests make scaling visibly sluggish; overly small ones cause runaway scaling well before pressure is real.
The chain depends on metrics-server: kubelet collects cAdvisor metrics, metrics-server aggregates them into the API server’s `metrics.k8s.io` API, and HPA reads from there. If any link is missing, HPA sets its target to `unknown` and displays `<unknown>/<unknown>`.
Many clusters ship without metrics-server, especially self-built or minimal installs. Then the HPA object is created successfully, its status stays unknown, and scaling never happens. The diagnostic is the TARGETS line of `kubectl describe hpa`: if you see `<unknown>`, the fault is in the metrics pipeline, not in your HPA spec.
Beyond Resource, HPA supports `type: Pods`, `type: Object`, and `type: External`. Pods and Object are fetched by the HPA controller directly from the metrics API; External requires an adapter (Prometheus Adapter, KEDA, …) that projects custom metrics into `external.metrics.k8s.io`. Custom metrics suit signals that actually move with business load, such as queue depth or concurrent sessions.
A few details matter. An HPA with `minReplicas: 1` has a documented scale-up delay of roughly 30-60 seconds; the `behavior` field lets you set `scaleDown` stabilization windows and policies to damp oscillation. And any manifest that does not pin `replicas` in the Deployment should treat HPA as the sole owner of replica count, otherwise the two overwrite each other.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-server
namespace: prod
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-server
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: AverageValue
averageValue: 512Mi
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Pods
value: 2
periodSeconds: 60VPA vs HPA resource contention: know exactly which fields they fight over
The contention is very concrete: when HPA computes CPU utilisation from the metrics API, the denominator is the pod’s CPU requests — and requests are exactly what VPA adjusts. If VPA observes low usage and lowers requests, HPA’s denominator shrinks, the same CPU usage yields a higher ratio, and unnecessary scale-out follows. This is the VPA-induced HPA oscillation.
The upstream guidance is to let only one autoscaler manage resources for a given workload. The common workable split is HPA for replica count and VPA for requests, but confined to components with genuinely distinct resource profiles (batch workers, offline indexers), with explicit acceptance of the oscillation risk that shared metrics imply.
If you do run both, several practices damp the conflict. Give VPA a sensible resource floor so it never recommends an unrealistically low request. Let VPA manage only memory, not CPU — VPA supports per-resource update policies. Or move the scaling signal to Custom Metrics so HPA stops using CPU requests as a denominator, cutting the coupling at its source.
VPA has three update modes. `Off` only records recommendations in `vpa-status` and changes nothing — the right mode for an observation period. `Initial` injects values only when a pod is created. `Auto` evicts and recreates running pods to apply new sizing, which breaks long-lived connections and local caches, so evaluate business tolerance before enabling it.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: batch-worker
namespace: prod
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: batch-worker
updatePolicy:
updateMode: "Off" # 先只观察,不自动改
resourcePolicy:
containerPolicies:
- containerName: worker
minAllowed:
cpu: 500m
memory: 512Mi
maxAllowed:
cpu: "4"
memory: 8Gi
controlledResources:
- cpu
- memoryWhen the right move is more nodes, not more replicas
When HPA fails to scale, there are only a handful of causes and each has its own fix. First, the per-node ceiling: pods already saturate a node, `maxReplicas` is not reached, yet new pods stay Pending. That calls for more nodes, not a higher cap. The tell is in events — `Insufficient cpu` on `kubectl describe pod` means the node is full, not that HPA is idle.
Second, the metrics pipeline is broken and the target reads `<unknown>`. Third, the cap itself is too conservative (small `maxReplicas`, or an API QPS limit in the way). Fourth, the application cannot scale horizontally at all — every replica shares a single stateful resource, the same local directory, the same distributed lock, or one connection-pool ceiling.
Cluster auto-scaling has two main tracks. Cluster Autoscaler matches pods against existing or creatable node groups, bounded by the groups you defined. Karpenter launches nodes sized to the actual pod requests, giving finer granularity, better behaviour on mixed workloads, and more aggressive consolidation of idle capacity. Pick by workload diversity: homogeneous fleets are fine with Cluster Autoscaler; widely varying shapes pay off with Karpenter.
Whichever you choose, add guards: a PodDisruptionBudget so scale-down cannot drop available replicas below a threshold, headroom between limits and requests, and hard caps plus budget alerts on node provisioning. The real danger of autoscaling nodes is not being too slow — it is spinning up a fleet at 3am and running up the bill, or consolidating away the only hot replica.
# 1) 副本扩不动时先看事件,而不是先调 HPA
kubectl describe pod <pending-pod> | sed -n '/Events/,$p'
kubectl get hpa -n prod -o wide
# 2) 目标值是 <unknown>?查指标链路
kubectl top nodes
kubectl -n kube-system get deploy metrics-server
# 3) 缩容保护
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: api-server-pdb
namespace: prod
spec:
minAvailable: 2
selector:
matchLabels:
app: api-server