Incident Response Playbook
Provide a repeatable response flow for security and production incidents with clear decision gates and evidence standards.
Purpose
Provide a repeatable response flow for security and production incidents with clear decision gates and evidence standards.
Authority Model
- Approval authority: single owner (explicit approval or direct manual execution by owner)
- Emergency decisions requiring approval:
- Outbound SSH exception creation beyond default policy
- Temporary reduction of security controls for break-glass recovery
- Production rollback vs forward-fix under active incident
Severity Levels
- Sev 1: active compromise or major outage with customer impact
- Sev 2: high-risk degradation or suspicious activity without confirmed compromise
- Sev 3: isolated fault with low blast radius
Response Timeline
0 to 5 minutes: Contain
- Stabilize exposed surfaces first (for example block outbound SSH where needed).
- Keep evidence-preserving actions first; avoid destructive cleanup before capture.
- Start incident note with UTC timestamp and responder identity.
5 to 30 minutes: Collect
- Capture process, sockets, active sessions, firewall state, and auth matrix.
- Capture logs and command outputs to an evidence folder under Incidents.
- Record approved containment actions and command references.
30 to 120 minutes: Eradicate and Recover
- Remove confirmed malicious artifacts.
- Rotate credentials and trusted keys if compromise indicators exist.
- Validate policy state (UFW, sshd effective settings, fail2ban status).
- Restore service with controlled verification checks.
Same day: Revalidate
- Re-capture outbound/inbound socket snapshots.
- Confirm expected access policy behavior from allowed and denied paths.
- Archive all outputs in a single incident package.
Required Evidence Package
- Incident summary
- Timeline
- Process and socket attribution
- Firewall and sshd effective configuration
- Authentication and key inventory snapshots
- Revalidation snapshots after containment
- Provider communication record (if applicable)
Store bundles at Incidents/<date-time>-<incident-name>/ with markdown + raw evidence.
Communication Rules
- Internal status updates should include: impact, current control state, next gate, and owner decision.
- Provider responses should be evidence-backed and stored with the same incident bundle.
Closure Criteria
- Root cause or best-supported cause is documented.
- All critical controls are restored and revalidated.
- Outstanding risk items are converted into tracked next-steps items.
- Owner signs off on close status.