Chuyển đến nội dung chính
Bảo mật

Detection Engineering & Incident Response trong DevSecOps

Phòng thủ tốt cần ba thứ: log có cấu trúc, detection rule map theo ATT&CK, và IR runbook đã diễn tập. Bài viết tổng hợp cách build chương trình detection-as-code và post-mortem blameless cho team DevSecOps.

Detection Engineering & Incident Response trong DevSecOps
Có log không có nghĩa là có detection. Có alert không có nghĩa là có response. Detection engineering là quá trình biến raw event thành tín hiệu hành động được, có owner, có MTTD/MTTR đo được.

Logging trước khi nói detection

Log không có cấu trúc thì không scale. Bốn nguyên tắc:

  • Structured JSON với field chuẩn: timestamp, service, env, request_id, user_id (đã hash/redact), action, result.
  • Correlation ID đi xuyên service, gateway, queue — dùng W3C Trace Context.
  • Đừng log raw secret/PII. Redact ở SDK chung trước khi xuất stdout.
  • Audit log riêng cho hành động nhạy cảm (admin action, key access, data export) — schema khác, retention dài hơn (12-36 tháng).

Nguồn log quan trọng cần stream về SIEM

NguồnGiá trị detection
Cloud audit (CloudTrail, Activity Log)IAM bất thường, key creation, region lạ
K8s audit log (Metadata level)RBAC change, exec into pod, secret access
Identity provider (Okta, Azure AD)Brute force, impossible travel, MFA bypass
Application audit logPrivilege escalation, data export khối lớn
Falco / TetragonHành vi runtime container
WAF, CDN, gatewayBrute force, scraping, anomaly

Sigma: detection-as-code

Sigma là format YAML chuẩn để mô tả detection rule độc lập với SIEM (chuyển đổi sang Splunk SPL, ELK Lucene, Sentinel KQL...). Ví dụ:

title: K8s exec into production pod
id: 1f0e3aa8-...
status: stable
logsource:
  product: kubernetes
  service: audit
detection:
  selection:
    verb: create
    objectRef.subresource: exec
    objectRef.namespace|startswith: prod-
  condition: selection
level: high
tags:
  - attack.execution
  - attack.t1609   # container administration command

Lưu rule trong git, review qua PR, có CI test (positive/negative case). Đó là detection-as-code.

Map theo MITRE ATT&CK

ATT&CK chia attacker behavior thành tactic (mục tiêu) và technique (cách thực hiện). Lợi ích khi map detection theo ATT&CK:

  • Biết mình thiếu detection ở tactic nào (vd: tốt ở Initial Access nhưng yếu ở Lateral Movement, Exfiltration).
  • So với threat intel: nhóm APT đang nhắm tới ngành mình dùng technique gì → ưu tiên detection tương ứng.
  • Báo cáo coverage rõ ràng cho leadership: "phủ 65% technique relevant cho stack chúng ta".

IR runbook và tabletop exercise

Vòng ứng phó NIST PICERL: Preparation → Identification → Containment → Eradication → Recovery → Lessons learned. Mỗi runbook nên có:

  • Severity matrix: SEV1/2/3 với SLA response.
  • On-call rotation rõ ràng, escalation path.
  • Communication template: status page, customer notice, regulator notice (NĐ 13: 72h).
  • Evidence preservation: snapshot disk/memory, copy log, export audit — TRƯỚC khi xoá/khôi phục.
  • Containment playbook theo loại sự cố: secret leak, account compromise, ransomware, data exfil.

Tổ chức tabletop exercise 1-2 lần/quý, mỗi lần 60-90 phút, với 1 scenario thực tế. Đo: thời gian phát hiện (MTTD), thời gian ngăn chặn (MTTC), thời gian khôi phục (MTTR).

Post-mortem blameless

Mục tiêu không phải tìm người để phạt, mà tìm điều kiện hệ thống cho phép sự cố xảy ra. Template tham khảo:

  • Tóm tắt (3-5 dòng).
  • Timeline: ai, làm gì, lúc nào, với UTC timestamp.
  • Impact: user, data, financial.
  • Root cause: contributing factor (thường nhiều, không một).
  • What went well: thừa nhận cái tốt để củng cố.
  • Action item: có owner và deadline thực tế, vào sprint backlog.

Văn hoá blameless cần lãnh đạo bảo vệ — nếu người báo cáo bị phạt, lần sau không ai dám kể trung thực.

Bug bounty và purple team

Bổ sung cho detection nội bộ:

  • Responsible disclosure via security.txt + email security@: rẻ và miễn phí.
  • Bug bounty qua HackerOne/Intigriti hoặc nội địa: scope rõ, payout rõ.
  • Purple team exercise: red team chạy 1 technique, blue team đo có detect không, sau đó cùng tune rule. Học chéo, không tính điểm.

Metric cần đo

  • MTTD theo loại sự cố.
  • MTTR vuln theo severity và % đúng SLA.
  • Tỷ lệ alert true positive (phòng cháy alert fatigue).
  • Coverage ATT&CK theo tactic.
  • Tần suất tabletop, % action item từ post-mortem hoàn thành đúng hạn.

Kết luận

Detection engineering và IR là nơi DevSecOps gặp SOC. Coi rule như code (Sigma + git + CI), coi runbook như product (review, version, đo MTTR), và bảo vệ văn hoá blameless. Khi đó mỗi sự cố trở thành nhiên liệu cải tiến hệ thống thay vì sự kiện gây hoảng loạn.

DUY TRAN
Tác giả

DUY TRAN

Pursuing an AI-first mindset and intelligent system architecture. I build solutions by combining technology, creativity, and the ability to see structure in chaos — the foundation for becoming a Solution Architect.

Bình luận

Bài viết liên quan