Chuyển đến nội dung chính

Detection Engineering & Incident Response in DevSecOps

Duy Tran10 min
Detection Engineering & Incident Response in DevSecOps
Having logs is not having detection. Having alerts is not having response. Detection engineering is the discipline of turning raw events into actionable signals with owners, MTTD and MTTR you can measure.

Logging before detection

Unstructured logs do not scale. Four principles:

  • Structured JSON with standard fields: timestamp, service, env, request_id, user_id (hashed/redacted), action, result.
  • Correlation IDs propagated across services, gateways, queues — use W3C Trace Context.
  • Never log raw secrets/PII. Redact in a shared SDK before stdout.
  • Separate audit logs for sensitive actions (admin actions, key access, data exports) — different schema, longer retention (12-36 months).

Log sources to stream into the SIEM

SourceDetection value
Cloud audit (CloudTrail, Activity Log)Unusual IAM, key creation, foreign region
K8s audit log (Metadata level)RBAC change, exec into pod, secret access
Identity provider (Okta, Azure AD)Brute force, impossible travel, MFA bypass
Application audit logPrivilege escalation, large data export
Falco / TetragonContainer runtime behavior
WAF, CDN, gatewayBrute force, scraping, anomalies

Sigma: detection-as-code

Sigma is a YAML format for SIEM-agnostic detection rules (convertible to Splunk SPL, ELK Lucene, Sentinel KQL, ...). Example:

title: K8s exec into production pod
id: 1f0e3aa8-...
status: stable
logsource:
  product: kubernetes
  service: audit
detection:
  selection:
    verb: create
    objectRef.subresource: exec
    objectRef.namespace|startswith: prod-
  condition: selection
level: high
tags:
  - attack.execution
  - attack.t1609   # container administration command

Store rules in git, review through PRs, with CI tests (positive/negative cases). That is detection-as-code.

Map to MITRE ATT&CK

ATT&CK breaks attacker behavior into tactics (goals) and techniques (means). Mapping detections to ATT&CK gives you:

  • Visibility into where you lack coverage (e.g., strong on Initial Access, weak on Lateral Movement, Exfiltration).
  • Alignment with threat intel: what techniques are targeting your industry — prioritise detections accordingly.
  • Clear coverage reporting for leadership: "65% of techniques relevant to our stack are covered".

IR runbooks and tabletop exercises

NIST PICERL: Preparation → Identification → Containment → Eradication → Recovery → Lessons learned. Each runbook should have:

  • Severity matrix with SLAs (SEV1/2/3).
  • Clear on-call rotation and escalation path.
  • Communication templates: status page, customer notice, regulator notice (Vietnam Decree 13 requires 72h).
  • Evidence preservation: snapshot disks/memory, copy logs, export audit trails — BEFORE deleting/restoring.
  • Containment playbooks per incident type: secret leak, account compromise, ransomware, data exfil.

Run a tabletop exercise 1-2 times per quarter, 60-90 minutes each, with one realistic scenario. Measure MTTD, MTTC and MTTR.

Blameless post-mortem

The goal is not to find someone to blame but to find the system conditions that allowed the incident. A reference template:

  • Summary (3-5 lines).
  • Timeline: who did what, when (UTC).
  • Impact: users, data, financial.
  • Root cause: usually multiple contributing factors, not one.
  • What went well: acknowledge what worked to reinforce it.
  • Action items: with owners and realistic deadlines, in the sprint backlog.

Blameless culture must be protected by leadership — if reporters get punished, no one will share honestly next time.

Bug bounty and purple team

  • Responsible disclosure via security.txt + security@ email: cheap, free.
  • Bug bounty on HackerOne/Intigriti or local platforms: clear scope, clear payouts.
  • Purple team exercises: red team runs one technique, blue team measures whether it gets detected, then both tune the rule together. Cross-learning, no scoreboard.

Metrics to track

  • MTTD by incident type.
  • Vuln MTTR by severity and SLA hit rate.
  • Alert true-positive rate (to prevent fatigue).
  • ATT&CK coverage by tactic.
  • Tabletop frequency, % of action items closed on time.

Conclusion

Detection engineering and IR are where DevSecOps meets the SOC. Treat rules as code (Sigma + git + CI), runbooks as products (review, version, measure MTTR), and protect the blameless culture. Then every incident becomes fuel for system improvement instead of a panic event.