Having logs is not having detection. Having alerts is not having response. Detection engineering is the discipline of turning raw events into actionable signals with owners, MTTD and MTTR you can measure.
Logging before detection
Unstructured logs do not scale. Four principles:
- Structured JSON with standard fields:
timestamp,service,env,request_id,user_id(hashed/redacted),action,result. - Correlation IDs propagated across services, gateways, queues — use W3C Trace Context.
- Never log raw secrets/PII. Redact in a shared SDK before stdout.
- Separate audit logs for sensitive actions (admin actions, key access, data exports) — different schema, longer retention (12-36 months).
Log sources to stream into the SIEM
| Source | Detection value |
|---|---|
| Cloud audit (CloudTrail, Activity Log) | Unusual IAM, key creation, foreign region |
| K8s audit log (Metadata level) | RBAC change, exec into pod, secret access |
| Identity provider (Okta, Azure AD) | Brute force, impossible travel, MFA bypass |
| Application audit log | Privilege escalation, large data export |
| Falco / Tetragon | Container runtime behavior |
| WAF, CDN, gateway | Brute force, scraping, anomalies |
Sigma: detection-as-code
Sigma is a YAML format for SIEM-agnostic detection rules (convertible to Splunk SPL, ELK Lucene, Sentinel KQL, ...). Example:
title: K8s exec into production pod
id: 1f0e3aa8-...
status: stable
logsource:
product: kubernetes
service: audit
detection:
selection:
verb: create
objectRef.subresource: exec
objectRef.namespace|startswith: prod-
condition: selection
level: high
tags:
- attack.execution
- attack.t1609 # container administration command
Store rules in git, review through PRs, with CI tests (positive/negative cases). That is detection-as-code.
Map to MITRE ATT&CK
ATT&CK breaks attacker behavior into tactics (goals) and techniques (means). Mapping detections to ATT&CK gives you:
- Visibility into where you lack coverage (e.g., strong on Initial Access, weak on Lateral Movement, Exfiltration).
- Alignment with threat intel: what techniques are targeting your industry — prioritise detections accordingly.
- Clear coverage reporting for leadership: "65% of techniques relevant to our stack are covered".
IR runbooks and tabletop exercises
NIST PICERL: Preparation → Identification → Containment → Eradication → Recovery → Lessons learned. Each runbook should have:
- Severity matrix with SLAs (SEV1/2/3).
- Clear on-call rotation and escalation path.
- Communication templates: status page, customer notice, regulator notice (Vietnam Decree 13 requires 72h).
- Evidence preservation: snapshot disks/memory, copy logs, export audit trails — BEFORE deleting/restoring.
- Containment playbooks per incident type: secret leak, account compromise, ransomware, data exfil.
Run a tabletop exercise 1-2 times per quarter, 60-90 minutes each, with one realistic scenario. Measure MTTD, MTTC and MTTR.
Blameless post-mortem
The goal is not to find someone to blame but to find the system conditions that allowed the incident. A reference template:
- Summary (3-5 lines).
- Timeline: who did what, when (UTC).
- Impact: users, data, financial.
- Root cause: usually multiple contributing factors, not one.
- What went well: acknowledge what worked to reinforce it.
- Action items: with owners and realistic deadlines, in the sprint backlog.
Blameless culture must be protected by leadership — if reporters get punished, no one will share honestly next time.
Bug bounty and purple team
- Responsible disclosure via
security.txt+security@email: cheap, free. - Bug bounty on HackerOne/Intigriti or local platforms: clear scope, clear payouts.
- Purple team exercises: red team runs one technique, blue team measures whether it gets detected, then both tune the rule together. Cross-learning, no scoreboard.
Metrics to track
- MTTD by incident type.
- Vuln MTTR by severity and SLA hit rate.
- Alert true-positive rate (to prevent fatigue).
- ATT&CK coverage by tactic.
- Tabletop frequency, % of action items closed on time.
Conclusion
Detection engineering and IR are where DevSecOps meets the SOC. Treat rules as code (Sigma + git + CI), runbooks as products (review, version, measure MTTR), and protect the blameless culture. Then every incident becomes fuel for system improvement instead of a panic event.
