Image scanning and threat modeling protect you before deploy. Admission policy protects you at the moment workloads enter the cluster. Runtime monitoring is the camera for everything that happens after.
Three layers of cluster defense
| Layer | Role | Tools |
|---|---|---|
| Pre-admission | Lint, scan image, verify signature | Trivy, Cosign |
| Admission | Block workloads that violate policy | Kyverno, OPA Gatekeeper |
| Runtime | Detect abnormal behavior, network policy | Falco, Cilium Tetragon, NetworkPolicy |
Kyverno vs OPA Gatekeeper — which one?
- Kyverno: write policies in plain YAML, short, easy to learn. Supports mutate, generate, verifyImages keyless. Fits most cases.
- OPA Gatekeeper: written in Rego, powerful for complex logic. Use it when you already have an OPA ecosystem (API gateway, microservice authz).
Recommendation: start with Kyverno for new clusters. Migrating between the two later is not too hard — policies are declarative.
Baseline policies you should have
- Apply Pod Security Standards: restricted on application namespaces.
- Block pods using
privileged: true,hostNetwork,hostPID. - Require images from internal registries (allowlist).
- Require
resources.requests/limitson every container. - Require standard labels (team, env, cost-center) for cost tracking and IR.
- Verify Cosign signatures on production images.
Roll out safely: Audit → Fix → Enforce
Never apply validationFailureAction: Enforce immediately on a running cluster. A reference workflow:
- Apply policies with
validationFailureAction: Audit. - Measure violations via the Kyverno PolicyReport CRD over 1-2 weeks.
- File fix tickets per offending workload, with an owner.
- When violations reach 0 in staging, switch to Enforce in dev → staging → prod.
Have a clear exception process: special workloads (e.g., a privileged debug pod) must carry an annotated reason, expiry and owner.
Default-deny network policy
Each application namespace should start with a "deny all" policy and open only what is actually needed:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: payments
spec:
podSelector: {}
policyTypes: ["Ingress", "Egress"]
Then open each connection: app → DB, app → service mesh, egress to external APIs through an egress gateway. With Cilium, you can write L7 policy (HTTP path, gRPC method, Kafka topic) — much stronger than port/IP only.
Falco: runtime detection for containers
Falco hooks into the kernel (eBPF or kernel module) to observe syscalls. Highest-value rules:
- Shell in container (
shell_in_container). Production rarely has a legitimate need for one. - Modifying binaries in a runtime image (sign of compromise).
- Unusual outbound connections to IPs/domains outside the allowlist.
- Reading sensitive files like
/etc/shadowor kubeconfig. - Mounting sensitive paths like
/var/run/docker.sock,/proc.
Stream Falco events to Slack/PagerDuty for high-severity rules and to your SIEM for later analysis.
Cilium Tetragon: low overhead, eBPF-assisted
Tetragon uses eBPF similarly but is performance-tuned for large clusters and can enforce, not just detect (e.g., kill processes when a syscall violates policy). Useful when in-kernel response time matters.
When an alert fires — short IR workflow
- Triage within 15 minutes: classify true vs false positive by severity.
- Isolate the pod with a
NetworkPolicyblocking egress; do NOT delete it (preserve evidence). - Capture snapshots:
kubectl debug, dump memory, copy logs, export the audit trail. - Rotate any credentials the pod could have touched: ServiceAccount tokens, mounted secrets.
- After investigation: write a blameless post-mortem and add a new Falco rule if needed.
Conclusion
Defending Kubernetes is more than RBAC and a firewall. You need three layers: pre-admission (image scan, sign), admission (Kyverno), runtime (Falco/Tetragon, NetworkPolicy). Begin with baseline policies in audit mode, gradually enforce, integrate Falco with your SIEM — that is a solid configuration for most production clusters in 2026.
