The one thing to remember: no policy means everything is allowed
Kubernetes gives every pod a flat, cluster-wide network identity, and any pod can send packets to any other pod. NetworkPolicy is not a default-deny firewall: it is the opposite. Traffic is evaluated only in a direction where at least one policy selects it; traffic that is not selected by any policy is allowed through.
Put differently, a namespace with zero NetworkPolicy objects has both ingress and egress fully open, and the cluster behaves like one big flat network to every pod. Teams often assume "no policy yet means safe by default" — the reality is the exact opposite, and it is one of the easiest security review items to miss.
Policy combination is a union, not an override. If a packet is selected by three policies and any one of them allows it, the packet passes; it is only dropped when every selecting policy requires a match that fails. When writing policies you have to reason about what other authors put in the same direction rather than assuming you are the only writer.
Policies target pods through `podSelector`, not Services. The labels you match must be on the pod template; matching Service labels does nothing. Also note that `podSelector: {}` is a legal empty selector meaning "all pods in this namespace", not "select nothing".
# 现状盘点:哪些命名空间/对象真的有策略
kubectl get networkpolicy -A
# 看某条策略到底选中了哪些 pod(最容易出错的一步)
kubectl get networkpolicy allow-dns -n prod -o yaml
# 逐条确认:selector 与 pod 模板 label 是否一致
kubectl get pods -n prod --show-labelsIngress and egress are two independent switches
A NetworkPolicy has two top-level fields that work independently. A policy containing only `ingress` leaves egress completely open, which is the root of the most common "but I restricted outbound" misunderstanding. Many teams write only ingress rules and assume pods can no longer phone home — they can.
Ingress means "who may connect to me". Rules match the source pod labels (`from`) plus the destination port (`ports`). When you use `ipBlock` instead, matching is by IP CIDR and ignores pod labels, which becomes counter-intuitive as soon as NAT or host networking is involved.
Egress means "who may I connect to". Rules match destination pod labels (`to`) plus destination port. Here is a classic trap: the moment you write any egress policy, every outbound flow — DNS included — must be explicitly allowed by some rule, otherwise the pod cannot even resolve names and presents as blanket timeouts.
The production pattern is a baseline policy declaring `policyTypes: [Ingress, Egress]` and listing what is allowed, with narrower policies layered on top. Apply ingress first and egress second: a broken ingress policy shows up as an obvious outage, while a broken egress policy usually shows up as intermittent timeouts.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny
namespace: prod
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
---
# 只允许 DNS 出站 —— 缺了这段,所有 pod 都会解析失败
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns-egress
namespace: prod
spec:
podSelector: {}
policyTypes:
- Egress
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
podSelector:
matchLabels:
k8s-app: kube-dns
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53Prerequisite: policy is enforced by the CNI, not by kube-proxy
A NetworkPolicy is a pure API object. Its mere existence changes no data path; the actual enforcer is the cluster CNI plugin. With a plugin that does not support NetworkPolicy, `kubectl get networkpolicy` still lists your objects and `kubectl apply` still returns success, while cluster behaviour is completely unchanged — the number one reason a correctly written policy has no effect.
Plugins with policy support include Calico, Cilium, Antrea and several vendor implementations. Verify two things before trusting it: the plugin declares NetworkPolicy support in its official compatibility matrix, and policy enforcement is enabled by default rather than gated behind an extra step (early Calico needed extra per-Pod-CIDR configuration, while Cilium additionally offers L7 policy).
Validation does not come from re-reading your YAML; it comes from the data path. Run a busybox pod and `wget`/`nc` against another pod by Service name or Pod IP: success when allowed, and a hanging connection (timeout) rather than an immediate refusal when denied. Hangs rather than rejections are the signature of policy drop, because the packet is discarded by iptables/eBPF rules and no RST is ever returned.
Pods must also be assigned a Pod CIDR. Policy enforcement needs IP-based matching and tagging, and some CNI modes allocate no Pod CIDR — in which case policies are inert too. This is the second most frequent explanation for "the object exists but nothing changes".
# 1) 确认 CNI 插件与策略模式
kubectl -n kube-system get pods -o wide | grep -Ei "calico|cilium|antrea|weave|flannel"
# 2) 实测:从被允许的源连被拒绝的目标
kubectl run nettest --rm -it --image=busybox:1.36 --restart=Never -- sh
# 在 nettest 里:
# nslookup kube-dns.kube-system.svc.cluster.local # DNS 策略生效检查
# wget -T 3 http://<target-pod-ip>:8080/ # 通 => 允许,timeout => 被策略丢弃
# 3) Calcio 场景查看策略端点(可选插件不同命令不同)
calicoctl get policy --scope=global 2>/dev/null | head -20Six reasons a policy you wrote is not enforced
First, mismatched podSelector labels. The policy sits on the Deployment template labels, but live pods also carry controller-added fields such as `pod-template-hash`; or you labelled only the Service. Compare character by character with `kubectl get pods --show-labels` rather than from memory.
Second, you wrote the Service port instead of the container port. NetworkPolicy matches the port the destination pod actually receives, which is normally containerPort; the Service port is often a different number, so the rule never matches. Confirm with `kubectl get pod -o jsonpath`.
Third, the direction is backwards. Writing ingress rules on the client pod to restrict access to a server pod is a no-op — ingress only governs "who connects to me". Constraining client behaviour requires an egress policy on the client pod.
Fourth, namespaceSelector is same-namespace by default. To reference another namespace you must state it explicitly, using the `kubernetes.io/metadata.name` label or a `matchLabels` selector; a bare `podSelector` is interpreted as pods inside the policy’s own namespace.
Fifth, NodePort and hostNetwork traffic can bypass policy. hostNetwork pods use the node IP and are typically not covered by NetworkPolicy, and traffic arriving from outside via NodePort is often not matched depending on implementation. When you accept isolation, verify where the test traffic actually originates, or you will wrongly conclude policy is broken.
Sixth, you configured only ingress and assumed egress was restricted. The safest pattern: as soon as a namespace has any policy, ship default-deny for both directions and then allow deliberately, so "the default" lives in cluster configuration instead of in someone’s memory.
# 逐项对照:策略 selector vs 真实 pod label
kubectl get netpol -n prod -o jsonpath='{range .items[*]}{.metadata.name}{" -> "}{.spec.podSelector}{"\n"}{end}'
kubectl get pods -n prod -o custom-columns=NAME:.metadata.name,LABELS:.metadata.labels
# 确认 containerPort 到底是多少(策略要写这个值)
kubectl get pod <pod> -n prod -o jsonpath='{.spec.containers[*].ports[*].containerPort}''
# 确认命名空间自带标签
kubectl get ns prod --show-labelsTreating policy as engineering: layering and naming
The long-term cost of policy is not writing it, it is maintaining it. Settle on three layers: one `default-deny` for the global floor, `allow-<role>` for bulk permission by role (frontend, backend, worker), and `allow-<role>-to-<dep>` for point exceptions. With a consistent naming prefix, `kubectl get netpol` reveals intent at a glance and code review has an obvious subject.
Application labels must be stable. Policies key off pod labels, and labels move with releases. Keep a dedicated set of semantic labels (for example `app`, `tier`, `team`) purely for policy, and never reuse volatile ones like `version` or a hash; renaming a label during an ordinary deploy silently deletes your security boundary.
Every change needs a verification script covering at least three directions: the allowed destination must succeed, the denied destination must time out, and an allowed source must connect in. Encoding those three assertions as a shell loop inside a pod and running it in CI beats a human clicking once. Policy changes are security changes and belong in the same review and rollback process as code.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-api-to-db
namespace: prod
spec:
podSelector:
matchLabels:
app: postgres # 目标:数据库 pod
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: api-server # 来源:同命名空间的 API pod
ports:
- protocol: TCP
port: 5432
# 出站不受本策略影响:这里的 podSelector 只影响 ingress 方向