Chuyển đến nội dung chính

Kubernetes Admission Policy & Runtime Defense with Kyverno and Falco

Duy Tran9 min
Kubernetes Admission Policy & Runtime Defense with Kyverno and Falco
Image scanning and threat modeling protect you before deploy. Admission policy protects you at the moment workloads enter the cluster. Runtime monitoring is the camera for everything that happens after.

Three layers of cluster defense

LayerRoleTools
Pre-admissionLint, scan image, verify signatureTrivy, Cosign
AdmissionBlock workloads that violate policyKyverno, OPA Gatekeeper
RuntimeDetect abnormal behavior, network policyFalco, Cilium Tetragon, NetworkPolicy

Kyverno vs OPA Gatekeeper — which one?

  • Kyverno: write policies in plain YAML, short, easy to learn. Supports mutate, generate, verifyImages keyless. Fits most cases.
  • OPA Gatekeeper: written in Rego, powerful for complex logic. Use it when you already have an OPA ecosystem (API gateway, microservice authz).

Recommendation: start with Kyverno for new clusters. Migrating between the two later is not too hard — policies are declarative.

Baseline policies you should have

  1. Apply Pod Security Standards: restricted on application namespaces.
  2. Block pods using privileged: true, hostNetwork, hostPID.
  3. Require images from internal registries (allowlist).
  4. Require resources.requests/limits on every container.
  5. Require standard labels (team, env, cost-center) for cost tracking and IR.
  6. Verify Cosign signatures on production images.

Roll out safely: Audit → Fix → Enforce

Never apply validationFailureAction: Enforce immediately on a running cluster. A reference workflow:

  1. Apply policies with validationFailureAction: Audit.
  2. Measure violations via the Kyverno PolicyReport CRD over 1-2 weeks.
  3. File fix tickets per offending workload, with an owner.
  4. When violations reach 0 in staging, switch to Enforce in dev → staging → prod.

Have a clear exception process: special workloads (e.g., a privileged debug pod) must carry an annotated reason, expiry and owner.

Default-deny network policy

Each application namespace should start with a "deny all" policy and open only what is actually needed:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-all
  namespace: payments
spec:
  podSelector: {}
  policyTypes: ["Ingress", "Egress"]

Then open each connection: app → DB, app → service mesh, egress to external APIs through an egress gateway. With Cilium, you can write L7 policy (HTTP path, gRPC method, Kafka topic) — much stronger than port/IP only.

Falco: runtime detection for containers

Falco hooks into the kernel (eBPF or kernel module) to observe syscalls. Highest-value rules:

  • Shell in container (shell_in_container). Production rarely has a legitimate need for one.
  • Modifying binaries in a runtime image (sign of compromise).
  • Unusual outbound connections to IPs/domains outside the allowlist.
  • Reading sensitive files like /etc/shadow or kubeconfig.
  • Mounting sensitive paths like /var/run/docker.sock, /proc.

Stream Falco events to Slack/PagerDuty for high-severity rules and to your SIEM for later analysis.

Cilium Tetragon: low overhead, eBPF-assisted

Tetragon uses eBPF similarly but is performance-tuned for large clusters and can enforce, not just detect (e.g., kill processes when a syscall violates policy). Useful when in-kernel response time matters.

When an alert fires — short IR workflow

  1. Triage within 15 minutes: classify true vs false positive by severity.
  2. Isolate the pod with a NetworkPolicy blocking egress; do NOT delete it (preserve evidence).
  3. Capture snapshots: kubectl debug, dump memory, copy logs, export the audit trail.
  4. Rotate any credentials the pod could have touched: ServiceAccount tokens, mounted secrets.
  5. After investigation: write a blameless post-mortem and add a new Falco rule if needed.

Conclusion

Defending Kubernetes is more than RBAC and a firewall. You need three layers: pre-admission (image scan, sign), admission (Kyverno), runtime (Falco/Tetragon, NetworkPolicy). Begin with baseline policies in audit mode, gradually enforce, integrate Falco with your SIEM — that is a solid configuration for most production clusters in 2026.