From df20493f56f1ae93e3858a499ec89c122b140807 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 16:51:15 +0900 Subject: [PATCH 01/11] docs(incident-management): add endpoint compromise runbook The runbooks covered infrastructure and key compromise but not the contributor workstation, which is where infostealer and fake-recruiter intrusions actually land. malware.mdx covers the victim side; this is the responder side. Severity derives from the access tier the user holds rather than the malware family, since access is observable in the first twenty minutes and impact is not. Escalators cover ungated pipelines, sub-threshold signers, and public channel control. The credential revocation table is ordered because the common failures are sequencing errors: rotating an SSH key before killing the live session lets the attacker re-add it, and changing a password before revoking sessions leaves the stolen cookie valid. Containment forks on capability rather than doctrine, and handles the already-powered-off device, which is the state malware.mdx produces. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 325 ++++++++++++++++++ .../runbooks/index.mdx | 1 + .../runbooks/overview.mdx | 12 +- .../incident-management/playbooks/malware.mdx | 2 + vocs.config.ts | 1 + 5 files changed, 336 insertions(+), 5 deletions(-) create mode 100644 docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx new file mode 100644 index 000000000..f06177b60 --- /dev/null +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -0,0 +1,325 @@ +--- +title: "Runbook: Endpoint Compromise | Security Alliance" +description: "Responder-side runbook for a compromised contributor workstation: contain the device, scope its access, revoke credentials in order, and rebuild." +tags: + - Security Specialist + - Operations & Strategy + - DevOps +contributors: + - role: wrote + users: [] + - role: reviewed + users: [] + - role: fact-checked + users: [] +--- + +import { TagList, AttributionList, ContributeFooter, Checklist } from '../../../../../components' + +# Runbook: Endpoint Compromise + + + + +> 🔑 **Key Takeaway**: A compromised workstation is a credential incident, not a hardware +> incident. Contain the device, inventory what it could reach, revoke in the right order, +> and rebuild rather than clean. + +**This is an example runbook.** Review and customize for your organization before use. Fill in your +isolation procedure, credential inventory, and named approvers. + +This runbook covers a contributor workstation (laptop or desktop, macOS/Windows/Linux) that is suspected +or confirmed compromised: infostealer malware, a malicious dependency or fake job "test task", a trojanized +meeting client, or hands-on access. Servers and mobile devices are out of scope, since server containment +usually means snapshot and quarantine. + +The person on the other end of this runbook is a colleague, and they may be frightened. Hand them +[Malware Infection](/incident-management/playbooks/malware) for the victim-facing side of the same incident. + +## Quick Reference + +| Field | Value | +| ------- | ------- | +| **Typical Severity** | Varies by access tier (see below) | +| **Primary Responder** | Security SME | +| **Last Updated** | [Date] | +| **Owner** | [Name] | + +Severity follows what the device could reach. Access is observable +within the first twenty minutes and impact is not, so this table stands in for the impact bands defined in +the [Incident Response Policy](/incident-management/incident-response-template/incident-response-policy). + +| Access tier held by the user | Severity | What the user holds | What the attacker gains | +| ------------------------------ | ---------- | --------------------- | ------------------------- | +| Signing, deployer, or production admin | P1 | Multisig signer key, contract upgrade or owner key, deployer EOA, hot wallet, cloud root, domain registrar, exchange withdrawal credentials | Direct fund movement or a privileged transaction, meeting the policy's fund-loss condition on the spot | +| Broad internal access | P2 | CI/CD and package publish tokens, repository write, cloud IAM write, internal admin panels, identity provider admin, official announcement channels | Reach production through a build, change what users are served, or pivot toward a P1 holder | +| Limited or no privileged access | P3 | Email, chat, shared documents, read-only dashboards | Impersonate a trusted colleague against the rows above, plus everything the user could read | + +Escalate a tier when any of these hold. The first two amount to holding a deployer key: a pipeline with no +approval gate deploys on the attacker's behalf, and a threshold only delays execution once a trusted signer +is proposing. The third turns one laptop into a campaign against users. + + +- [ ] Pipeline or repo write reaches production without approval (P1, see + [Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise)) +- [ ] Multisig signer below threshold (P1) +- [ ] Controls official social, announcements, or docs publishing (P2 minimum) + + +Until the inventory in [Credential Revocation Order](#credential-revocation-order) is complete, treat the +incident at the higher tier. Access is least understood during the first twenty minutes, which is exactly +when severity gets set. + +This runbook stops at P3. The policy also defines P4 and P5, but their response times ("can be scheduled" +and "no immediate action") do not fit a confirmed endpoint compromise, where a stolen session cookie stays +valid until someone revokes it. + +## Identification + +### Symptoms + + +- [ ] Unexpected script, installer, or "fix" for a broken meeting link was run +- [ ] Job assessment, trading bot, or repo arrived from a new contact +- [ ] Browser or editor extension from a repo recommendation, sent file, or sideload +- [ ] Wallet or node plugin installed from a link rather than the vendor +- [ ] Unfamiliar login alerts, sessions, or unprompted 2FA challenges +- [ ] Wallet drain or sweeper activity on a user-controlled address +- [ ] Unexpected persistence (login items, launch agents, scheduled tasks) +- [ ] Outbound traffic to unknown hosts +- [ ] Third-party notification (exchange, partner, another protocol's security team) + + +Extension and plugin installs deserve their own check: a malicious one reads every session the browser +holds, and a repository can pressure the install through its own recommendations. See +[Integrated Development Environments](/devsecops/integrated-development-environments). + +**If you have EDR:** a single quarantined detection still triggers this runbook, since quarantine only +proves the product caught one payload. Treat the signals below as confirmation, and remember that an +unmanaged device produces none of them. + + +- [ ] Agent stopped reporting, disabled, or tamper alert +- [ ] Credential-store access (keychain, cookie database, process memory) +- [ ] Shell or interpreter spawned from a browser, archive, or downloads folder +- [ ] Unsigned binary executed from a user-writable path +- [ ] Further activity after a detection was blocked or auto-resolved +- [ ] New extension or plugin making outbound calls + + +### Differentiation + +If a wallet drained but the device shows no sign of compromise, the user most likely signed a malicious +approval. Use [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) and +the [Drainer playbook](/incident-management/playbooks/hacked-drainer) instead. If several contributors are +affected at once with no shared file or link, suspect the identity provider or the supply chain rather than +a single endpoint. + +## Immediate Actions + +### Step 1: Reach the user out of band + +**Why:** Their chat, email, and phone accounts may already be in the attacker's hands. + + +- [ ] Call on a channel not tied to the suspect device +- [ ] Confirm identity by voice, never by text alone +- [ ] Tell them to stop using it and leave it powered as-is +- [ ] Say plainly that blame is not the point + + +If the user does not answer, keep going. Session revocation, identity provider lockout, and cloud access +cuts all work without them, and the account is the urgent part. Leave a message with one instruction: +leave the device alone. + +> **If the user may be the threat actor** (fake contributor, suspected DPRK IT worker, hostile departure): +> do not call them and do not reveal that you detected anything. Preserve access logs and device state, cut +> access quietly, and engage +> [Legal](/incident-management/incident-response-template/contacts#legal--communications) before contact. +> See [Mitigating DPRK IT Workers](/dprk-it-workers/mitigating-dprk-it-workers). + +### Step 2: Contain the device + +**Why:** Stop live session abuse and further exfiltration without destroying what you need to scope the +incident. + +Pick the one row that matches your capability: + +| Situation | Action | +| ----------- | -------- | +| Device already powered off | Leave it off, and do not boot it to "check something" | +| Endpoint tooling with network containment | Contain the host in the console, leave it powered for memory capture | +| No tooling, memory capture possible within [X hours] | Disconnect the network, do not sleep or shut down | +| No tooling, no memory capture possible | Disconnect the network, then power off fully | + +Expect the first row often. [Malware Infection](/incident-management/playbooks/malware) tells the affected +person to power off immediately, so a user who followed that guidance, or who simply panicked, hands you a +cold machine. Memory is already gone at that point and booting the device only runs the malware's +persistence again, so the rest of this runbook proceeds from the credential side. + +Powering off destroys running processes, injected code, and secrets decrypted in memory, which is the +evidence that tells you what actually ran and therefore what must be rotated. It also stops an active +thief. Choose speed when nobody is positioned to capture memory: a machine left running on a desk "for +forensics" that nobody ever images is the worst of both options. + + +- [ ] Network access cut, by console containment or physical disconnect +- [ ] Device not wiped, not "cleaned", and not casually restarted +- [ ] Hardware wallets, security keys, and external drives unplugged and set aside +- [ ] Time and method of containment recorded in the + [Incident Log](/incident-management/incident-response-template/templates/incident-log-template) + + +### Step 3: Scope what the device could reach + +**Why:** Severity, revocation order, and blast radius all depend on this list, and nobody recalls it +reliably under pressure. + + +- [ ] Accounts held (identity provider, chat, email, code, cloud, treasury) +- [ ] Keys held, and signer roles the user can approve +- [ ] Paired devices (2FA apps, security keys, hardware wallets) +- [ ] Plaintext credentials on disk (.env files, cloud and SSH config, notes apps) +- [ ] Severity set from the access tier table, defaulting high + + +**If you have endpoint management:** pull the installed application list, the extension inventory, and +recent process history instead of relying on the user's memory. + +### Step 4: Start the revocation clock + +Move straight to [Credential Revocation Order](#credential-revocation-order). Session revocation is time +sensitive: a stolen session cookie keeps working until the session itself is killed, no matter how many +passwords change. + +## Credential Revocation Order + +**Why order matters:** the usual failures are sequencing errors. Rotating an SSH key before killing the +attacker's live session lets them add the new key themselves. Changing a password before revoking sessions +leaves the stolen cookie valid. Re-enrolling 2FA from a cloud backup restores the stolen seed. + +Work top to bottom. Each row links to whoever owns the procedure. + +| # | Credential class | Where | Owner | Notes | +| --- | ------------------ | ------- | ------- | ------- | +| 1 | Active sessions | Identity provider, chat, email, code host | [Name] | Kill all sessions first; cookie theft is the default infostealer outcome | +| 2 | Signing and deployer keys | Multisig, contracts, treasury | [Name] | Run in parallel at P1 per [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) | +| 3 | Password manager | Vault account | [Name] | Assume vault contents exposed; rotate master password and re-enroll 2FA from a clean device | +| 4 | MFA enrollments | Identity provider, exchanges | [Name] | Re-enroll on a new device; never restore an authenticator from cloud backup | +| 5 | SSH and git access | Code host, servers | [Name] | Remove old public keys and look for attacker-added keys and deploy keys | +| 6 | Cloud and CI/CD tokens | Cloud IAM, CI secrets, registries | [Name] | See [Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise) | +| 7 | API and exchange keys | Exchanges, data providers | [Name] | Withdrawal-capable keys first | +| 8 | Passwords | Everything above | [Name] | Last, after sessions are dead, from a clean device | +| 9 | Recovery channels | Mail rules, recovery codes, phone, OAuth grants | [Name] | Attackers persist here; check rules and grants, not just credentials | + + +- [ ] Every row above assigned to a named person +- [ ] All revocation done from a known-clean device +- [ ] Persistence checked (mail rules, OAuth grants, deploy keys, new signers) +- [ ] Anything not revocable within [X hours] escalated + + +## Investigation + +### Key Questions + + +- [ ] What was the initial access (link, file, dependency, installer, physical)? +- [ ] When did it start, and what happened before containment? +- [ ] What was exfiltrated, and what is only assumed exfiltrated? +- [ ] Did the attacker move laterally, or stop at the device? +- [ ] Are other contributors exposed to the same lure? + + +### Evidence to Collect + +| Data | Source | +| ------ | -------- | +| Memory and disk images | The device, before rebuild | +| Process, persistence, and install history | Endpoint tooling, local logs | +| Authentication and session logs | Identity provider, code host, cloud | +| Outbound network telemetry | DNS, proxy, and firewall logs, if retained | +| The lure itself | Message, repository, installer, meeting link | +| On-chain activity | Block explorer, for addresses the user controls | + +Preserve evidence before rebuilding. See [Forensic Readiness](/incident-management/forensic-readiness) for +chain of custody and retention design, which only helps if it is already in place when the incident starts. + +If the same lure reached other contributors, treat that wider exposure as its own incident and send the +team the specific indicators to look for. + +If the investigation shows the lure arrived but was never executed, and no credential exposure is +confirmed, downgrade the incident and close it. Record what ruled execution out in the +[Incident Log](/incident-management/incident-response-template/templates/incident-log-template). A P1 that +nobody closes buries the next real one. + +## Recovery and Prevention + +### Rebuild, do not clean + +Wipe and rebuild, or replace the hardware. Do not restore the user profile from a backup taken after the +compromise window opened, and do not accept "the antivirus removed it" as an outcome. You cannot prove an +infostealer is gone, and the cost of being wrong is handing back signing authority. + + +- [ ] Device wiped and reinstalled from trusted media, or replaced +- [ ] Data moved file by file, never a whole user folder +- [ ] Browser profiles, extensions, and shell config rebuilt, not migrated +- [ ] Device retained if forensic or regulatory obligations require it + + +### Return-to-service gate + +Privileged access returns only when a named approver confirms all of the following: + + +- [ ] Device rebuilt or replaced, and enrolled in management +- [ ] Every credential class above reissued from a clean device +- [ ] MFA re-enrolled fresh, not restored from a backup +- [ ] Recovery channels re-verified (mail rules, recovery codes, phone, OAuth grants) +- [ ] Signing and deployer roles restored last, after [X days, 30 suggested] +- [ ] Approver: [Name] + + +### Prevention + +Device tiers belong to [Endpoint Security](/opsec/endpoint/overview): signers and production admins on +managed hardware, contractors on VDI, everyone else behind a managed browser. + + +- [ ] Signing devices separate from daily-driver workstations +- [ ] Device tier matched to role risk, EDR and disk encryption on P1 holders +- [ ] Extensions allowlisted, repo-recommended ones declined by default +- [ ] Untrusted code run only in a sandbox or disposable VM +- [ ] Phishing-resistant MFA on high-value accounts +- [ ] Shortened session lifetimes for privileged tooling +- [ ] Contributors know where to report, including out of hours + + +## Escalation + + +- [ ] [Decision Makers](/incident-management/incident-response-template/contacts#decision-makers) - + immediately for any P1 access tier +- [ ] [Security Partners](/incident-management/incident-response-template/contacts#security-partners) - for + forensics, or whenever lateral movement is suspected +- [ ] [Legal](/incident-management/incident-response-template/contacts#legal--communications) - if funds + were stolen, data was exposed, or insider involvement is suspected +- [ ] [SEAL 911](https://t.me/seal_911_bot) - for crypto-specific containment help beyond your team + + +## Further reading + +- [Runbooks overview](/incident-management/incident-response-template/runbooks/overview): how the runbooks in this section fit together +- [Malware Infection](/incident-management/playbooks/malware): the victim-facing guide to hand the + affected person +- [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise): rotation + procedures for signer and deployer keys +- [Forensic Readiness](/incident-management/forensic-readiness): preserving evidence before you rebuild +- [Endpoint Security](/opsec/endpoint/overview): device tiers, EDR, and MDM that limit the blast radius +- [Developer-targeted intrusions](/devsecops/developer-targeted-intrusions/overview): how the lure reaches + a contributor in the first place + +--- + + diff --git a/docs/pages/incident-management/incident-response-template/runbooks/index.mdx b/docs/pages/incident-management/incident-response-template/runbooks/index.mdx index 6d9b31a44..8cbd801e5 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/index.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/index.mdx @@ -14,6 +14,7 @@ title: "Runbooks" - [Runbooks](/incident-management/incident-response-template/runbooks/overview) - [Runbook: Smart Contract Exploit](/incident-management/incident-response-template/runbooks/smart-contract-exploit) - [Runbook: Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) +- [Runbook: Endpoint Compromise](/incident-management/incident-response-template/runbooks/endpoint-compromise) - [Runbook: Frontend Compromise](/incident-management/incident-response-template/runbooks/frontend-compromise) - [Runbook: DNS Hijack](/incident-management/incident-response-template/runbooks/dns-hijack) - [Runbook: CDN/Hosting Compromise](/incident-management/incident-response-template/runbooks/cdn-hosting-compromise) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/overview.mdx b/docs/pages/incident-management/incident-response-template/runbooks/overview.mdx index 3ef2ec215..27d8171bc 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/overview.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/overview.mdx @@ -40,15 +40,17 @@ ensure consistent response. active exploit or critical vulnerability. 2. [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise): private key or signer compromise. -3. [Frontend Compromise](/incident-management/incident-response-template/runbooks/frontend-compromise): +3. [Endpoint Compromise](/incident-management/incident-response-template/runbooks/endpoint-compromise): + contributor workstation compromise; severity follows the access tier the user holds. +4. [Frontend Compromise](/incident-management/incident-response-template/runbooks/frontend-compromise): website/UI compromise (routes to specialized runbooks below). -4. [DNS Hijack](/incident-management/incident-response-template/runbooks/dns-hijack): domain/DNS +5. [DNS Hijack](/incident-management/incident-response-template/runbooks/dns-hijack): domain/DNS compromise. -5. [CDN/Hosting Compromise](/incident-management/incident-response-template/runbooks/cdn-hosting-compromise): +6. [CDN/Hosting Compromise](/incident-management/incident-response-template/runbooks/cdn-hosting-compromise): CDN or hosting provider compromise. -6. [Dependency Attack](/incident-management/incident-response-template/runbooks/dependency-attack): +7. [Dependency Attack](/incident-management/incident-response-template/runbooks/dependency-attack): npm/package supply chain attack. -7. [Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise): +8. [Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise): CI/CD compromise. ### High/Moderate (P2-P3) diff --git a/docs/pages/incident-management/playbooks/malware.mdx b/docs/pages/incident-management/playbooks/malware.mdx index cc7b3a7b7..cbc116eba 100644 --- a/docs/pages/incident-management/playbooks/malware.mdx +++ b/docs/pages/incident-management/playbooks/malware.mdx @@ -185,6 +185,8 @@ Here are some guides specifically for securing your: ## Further reading - [Playbooks overview](/incident-management/playbooks/overview): how the playbooks in this section fit together +- [Endpoint Compromise runbook](/incident-management/incident-response-template/runbooks/endpoint-compromise): + the responder-side process if this happened to a colleague - [Endpoint Security](/opsec/endpoint/overview): device hardening that limits the blast radius - [Drainer playbook](/incident-management/playbooks/hacked-drainer): response if wallet access followed - [SEAL 911 War Room Guidelines](/incident-management/playbooks/seal-911-war-room-guidelines): reaching outside help fast diff --git a/vocs.config.ts b/vocs.config.ts index dc1e691b4..fd0abfcc6 100644 --- a/vocs.config.ts +++ b/vocs.config.ts @@ -293,6 +293,7 @@ const config = { { text: 'Overview', link: '/incident-management/incident-response-template/runbooks/overview' }, { text: 'Smart Contract Exploit', link: '/incident-management/incident-response-template/runbooks/smart-contract-exploit' }, { text: 'Key Compromise', link: '/incident-management/incident-response-template/runbooks/key-compromise' }, + { text: 'Endpoint Compromise', link: '/incident-management/incident-response-template/runbooks/endpoint-compromise' }, { text: 'Frontend Compromise', link: '/incident-management/incident-response-template/runbooks/frontend-compromise' }, { text: 'DNS Hijack', link: '/incident-management/incident-response-template/runbooks/dns-hijack' }, { text: 'CDN/Hosting Compromise', link: '/incident-management/incident-response-template/runbooks/cdn-hosting-compromise' }, From e0e1e09cd85408a1bbb9a583cc8149d66ce1421e Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 16:55:45 +0900 Subject: [PATCH 02/11] docs(incident-management): close NIST and CISA gaps in endpoint runbook Checked the runbook against SP 800-61r3, SP 800-86, and the CISA incident and vulnerability response playbooks. Four things were missing. Evidence collection now runs in order of volatility per SP 800-86, so a responder holding a live machine knows memory outranks disk rather than working down an unordered list. Eradication is gated on finishing the foothold inventory first, which is the CISA sequencing rule. Rebuilding the laptop while an OAuth grant survives moves the attacker off the device and leaves them in the accounts, and the rebuild destroys the evidence that would have found them. The runbook had no way to close an incident. SP 800-61r3 makes improve a standing function, and the repo's own runbook template asks for a resolution checklist, so there is now a closing section covering the incident log, post-mortem, indicator sharing, and the detection gap. External notification was absent on the responder side even though malware.mdx carries it for victims. Legal owns the decision, but the clock starts at discovery, so it belongs in escalation. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 62 +++++++++++++++---- 1 file changed, 51 insertions(+), 11 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index f06177b60..33513480f 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -233,17 +233,25 @@ Work top to bottom. Each row links to whoever owns the procedure. ### Evidence to Collect -| Data | Source | -| ------ | -------- | -| Memory and disk images | The device, before rebuild | -| Process, persistence, and install history | Endpoint tooling, local logs | -| Authentication and session logs | Identity provider, code host, cloud | -| Outbound network telemetry | DNS, proxy, and firewall logs, if retained | -| The lure itself | Message, repository, installer, meeting link | -| On-chain activity | Block explorer, for addresses the user controls | - -Preserve evidence before rebuilding. See [Forensic Readiness](/incident-management/forensic-readiness) for -chain of custody and retention design, which only helps if it is already in place when the incident starts. +Collect in order of volatility, following [NIST SP +800-86](https://csrc.nist.gov/pubs/sp/800/86/final): the most perishable evidence first, because each step +down the list survives longer without you. If the device is still running, memory outranks everything. + +| Order | Data | Source | +| ------- | ------ | -------- | +| 1 | Memory image | The live device, before any shutdown | +| 2 | Live network state and running processes | The live device | +| 3 | Disk image | The device, before rebuild | +| 4 | Persistence and install history | Endpoint tooling, local logs | +| 5 | Authentication and session logs | Identity provider, code host, cloud | +| 6 | Outbound network telemetry | DNS, proxy, and firewall logs, if retained | +| 7 | The lure itself | Message, repository, installer, meeting link | +| 8 | On-chain activity | Block explorer, for addresses the user controls | + +Rows 1 and 2 are gone the moment the device powers off, which is why Step 2 forks on whether anyone can +capture them. Rows 4 through 6 depend on retention windows you set long before the incident. See +[Forensic Readiness](/incident-management/forensic-readiness) for chain of custody and retention design, +which only helps if it is already in place when the incident starts. If the same lure reached other contributors, treat that wider exposure as its own incident and send the team the specific indicators to look for. @@ -257,6 +265,10 @@ nobody closes buries the next real one. ### Rebuild, do not clean +Finish identifying every foothold before you start removing any of them. Rebuilding the laptop while an +OAuth grant, mail rule, or deploy key survives just moves the attacker off the device and leaves them in +the accounts, and the rebuild destroys the evidence that would have found them. + Wipe and rebuild, or replace the hardware. Do not restore the user profile from a backup taken after the compromise window opened, and do not accept "the antivirus removed it" as an outcome. You cannot prove an infostealer is gone, and the cost of being wrong is handing back signing authority. @@ -296,6 +308,23 @@ managed hardware, contractors on VDI, everyone else behind a managed browser. - [ ] Contributors know where to report, including out of hours +### Closing the incident + +An incident is not closed because the laptop was rebuilt. [NIST SP +800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats improvement as a standing function rather +than a meeting held once the noise dies down, so the work below is the last step of this runbook. + + +- [ ] Timeline written up in the + [Incident Log](/incident-management/incident-response-template/templates/incident-log-template) +- [ ] Return-to-service gate signed off by the named approver +- [ ] Post-mortem scheduled if the compromise reached P1 or P2 + ([template](/incident-management/incident-response-template/templates/post-mortem-template)) +- [ ] Indicators shared with the team, and with partners if the lure was reused against them +- [ ] Prevention items above converted into owned work, not left in this checklist +- [ ] Detection gap reviewed: what would have caught this a day earlier + + ## Escalation @@ -308,6 +337,17 @@ managed hardware, contractors on VDI, everyone else behind a managed browser. - [ ] [SEAL 911](https://t.me/seal_911_bot) - for crypto-specific containment help beyond your team +Legal owns external notification, but the clock starts at discovery and runs while you are still +containing. Raise it early enough that a deadline stays a choice. + + +- [ ] Disclosure obligations checked against your jurisdictions and [X hours] reporting windows +- [ ] Affected users or customers notified if their data or funds were reachable +- [ ] Law enforcement report filed if funds were stolen (see + [Malware Infection](/incident-management/playbooks/malware) for the reporting venues) +- [ ] Partners and exchanges told which addresses and accounts to watch + + ## Further reading - [Runbooks overview](/incident-management/incident-response-template/runbooks/overview): how the runbooks in this section fit together From ca235d74cfefe86e760acc0ab63d8847647a4e49 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 17:03:34 +0900 Subject: [PATCH 03/11] docs(incident-management): tighten prose in endpoint runbook The closing section opened by negating ("an incident is not closed because the laptop was rebuilt") and letting the next sentence carry the correction, which is the same rhetorical move the key takeaway already spends. Leads with the claim instead. Also drops a filler "just", replaces an unclear line about evidence surviving "without you" with what the rows actually do, and unpacks "a deadline stays a choice" into plain terms. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 18 +++++++++--------- 1 file changed, 9 insertions(+), 9 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index 33513480f..f71738a41 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -233,9 +233,9 @@ Work top to bottom. Each row links to whoever owns the procedure. ### Evidence to Collect -Collect in order of volatility, following [NIST SP -800-86](https://csrc.nist.gov/pubs/sp/800/86/final): the most perishable evidence first, because each step -down the list survives longer without you. If the device is still running, memory outranks everything. +Collect in order of volatility, following [NIST SP 800-86](https://csrc.nist.gov/pubs/sp/800/86/final). +The top rows vanish on their own; the lower ones keep. If the device is still running, memory outranks +everything. | Order | Data | Source | | ------- | ------ | -------- | @@ -266,8 +266,8 @@ nobody closes buries the next real one. ### Rebuild, do not clean Finish identifying every foothold before you start removing any of them. Rebuilding the laptop while an -OAuth grant, mail rule, or deploy key survives just moves the attacker off the device and leaves them in -the accounts, and the rebuild destroys the evidence that would have found them. +OAuth grant, mail rule, or deploy key survives moves the attacker off the device and leaves them in the +accounts, and the rebuild destroys the evidence that would have found them. Wipe and rebuild, or replace the hardware. Do not restore the user profile from a backup taken after the compromise window opened, and do not accept "the antivirus removed it" as an outcome. You cannot prove an @@ -310,9 +310,9 @@ managed hardware, contractors on VDI, everyone else behind a managed browser. ### Closing the incident -An incident is not closed because the laptop was rebuilt. [NIST SP -800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats improvement as a standing function rather -than a meeting held once the noise dies down, so the work below is the last step of this runbook. +The rebuild is the midpoint. [NIST SP 800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats +improvement as a standing function rather than a meeting held once the noise dies down, so the work below +closes the runbook. - [ ] Timeline written up in the @@ -338,7 +338,7 @@ than a meeting held once the noise dies down, so the work below is the last step Legal owns external notification, but the clock starts at discovery and runs while you are still -containing. Raise it early enough that a deadline stays a choice. +containing. Raise it early, while the timing is still yours to pick. - [ ] Disclosure obligations checked against your jurisdictions and [X hours] reporting windows From 3c514ba62db9ca1bf0f7384e6bdaa742a678d0b0 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 17:07:31 +0900 Subject: [PATCH 04/11] docs(incident-management): wire endpoint runbook to existing pages The symptoms list describes lures this repo already documents with real message screenshots, but nothing connected the two. Differentiation now routes a suspicious meeting-link domain or fake recruiter task to the DPRK playbook, and a Zoom screen-share request to ELUSIVE COMET, since naming the campaign tells the responder which follow-on accounts to check. Both playbooks link back. ELUSIVE COMET previously ended at the malware playbook, which is victim guidance, leaving no route to the org-side process. External notification points at the communications template and communication strategies, which own wording, cadence, and spokesperson while this runbook only covers the legal clock. Closing the incident points at Lessons Learned for the review questions and blameless format. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 14 ++++++++++++-- .../incident-management/playbooks/hacked-dprk.mdx | 2 ++ .../playbooks/hacked-elusive-comet.mdx | 2 ++ 3 files changed, 16 insertions(+), 2 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index f71738a41..f58c07956 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -115,6 +115,13 @@ the [Drainer playbook](/incident-management/playbooks/hacked-drainer) instead. I affected at once with no shared file or link, suspect the identity provider or the supply chain rather than a single endpoint. +Two of the lures above are documented campaigns, with real message screenshots worth comparing against what +the user received. A meeting link that resolved to an unfamiliar domain, or a fake recruiter task, matches +[North Korea (DPRK) Attack](/incident-management/playbooks/hacked-dprk). A Zoom call where the caller asked +for screen share or remote control matches +[ELUSIVE COMET](/incident-management/playbooks/hacked-elusive-comet). Naming the campaign tells you which +follow-on accounts to check, so it is worth the two minutes. + ## Immediate Actions ### Step 1: Reach the user out of band @@ -312,7 +319,8 @@ managed hardware, contractors on VDI, everyone else behind a managed browser. The rebuild is the midpoint. [NIST SP 800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats improvement as a standing function rather than a meeting held once the noise dies down, so the work below -closes the runbook. +closes the runbook. [Lessons Learned](/incident-management/lessons-learned) carries the review questions and +the blameless format. - [ ] Timeline written up in the @@ -338,7 +346,9 @@ closes the runbook. Legal owns external notification, but the clock starts at discovery and runs while you are still -containing. Raise it early, while the timing is still yours to pick. +containing. Raise it early, while the timing is still yours to pick. For wording, cadence, and who speaks, +use [Communications](/incident-management/incident-response-template/communications) and +[Communication Strategies](/incident-management/communication-strategies). - [ ] Disclosure obligations checked against your jurisdictions and [X hours] reporting windows diff --git a/docs/pages/incident-management/playbooks/hacked-dprk.mdx b/docs/pages/incident-management/playbooks/hacked-dprk.mdx index a169a2f75..b9ea38e85 100644 --- a/docs/pages/incident-management/playbooks/hacked-dprk.mdx +++ b/docs/pages/incident-management/playbooks/hacked-dprk.mdx @@ -94,6 +94,8 @@ file you were expecting, so that you do not suspect you were infected. - [Playbooks overview](/incident-management/playbooks/overview): how the playbooks in this section fit together - [DPRK IT Workers](/dprk-it-workers/overview): who the actor is and how they get hired - [Mitigating DPRK IT Workers](/dprk-it-workers/mitigating-dprk-it-workers): hardening and post-discovery steps +- [Endpoint Compromise runbook](/incident-management/incident-response-template/runbooks/endpoint-compromise): + what the security team runs if this reached a contributor's machine - [SEAL 911 War Room Guidelines](/incident-management/playbooks/seal-911-war-room-guidelines): reaching outside help fast --- diff --git a/docs/pages/incident-management/playbooks/hacked-elusive-comet.mdx b/docs/pages/incident-management/playbooks/hacked-elusive-comet.mdx index 3870d84b4..9db948bf8 100644 --- a/docs/pages/incident-management/playbooks/hacked-elusive-comet.mdx +++ b/docs/pages/incident-management/playbooks/hacked-elusive-comet.mdx @@ -81,6 +81,8 @@ accounts and sending out phishing messages to more people. - [Playbooks overview](/incident-management/playbooks/overview): how the playbooks in this section fit together - [Zoom Hardening](/guides/endpoint-security/zoom-hardening): closing the vector this attack uses - [Malware playbook](/incident-management/playbooks/malware): response once code has run on the device +- [Endpoint Compromise runbook](/incident-management/incident-response-template/runbooks/endpoint-compromise): + what the security team runs if this reached a contributor's machine - [SEAL 911 War Room Guidelines](/incident-management/playbooks/seal-911-war-room-guidelines): reaching outside help fast --- From 5f94a2da0185174fb261287d34538f5b3d01b070 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 17:08:49 +0900 Subject: [PATCH 05/11] docs(incident-management): cut vague endorsement from differentiation The paragraph added with the playbook cross-links leaned on "worth" twice in four sentences and closed on a generic recommendation rather than a reason. Says what the screenshots are for and stops there. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index f58c07956..1f5c6afeb 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -115,12 +115,13 @@ the [Drainer playbook](/incident-management/playbooks/hacked-drainer) instead. I affected at once with no shared file or link, suspect the identity provider or the supply chain rather than a single endpoint. -Two of the lures above are documented campaigns, with real message screenshots worth comparing against what -the user received. A meeting link that resolved to an unfamiliar domain, or a fake recruiter task, matches +Two of the lures above are documented campaigns. Both playbooks carry screenshots of the messages +attackers sent, so you can put them in front of the user and ask whether theirs matched. A meeting link +that resolved to an unfamiliar domain, or a fake recruiter task, matches [North Korea (DPRK) Attack](/incident-management/playbooks/hacked-dprk). A Zoom call where the caller asked for screen share or remote control matches [ELUSIVE COMET](/incident-management/playbooks/hacked-elusive-comet). Naming the campaign tells you which -follow-on accounts to check, so it is worth the two minutes. +follow-on accounts to check. ## Immediate Actions From 2913978732e1615d1ab2d5405b7545bebe53edc3 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 17:28:56 +0900 Subject: [PATCH 06/11] docs(incident-management): cut repetition from endpoint runbook "A stolen session cookie stays valid until revoked" was stated three times: in the P4/P5 note, in step 4, and in row 1 of the revocation table. Step 4 carried nothing else, since it only pointed at the section immediately below it, so it is gone. Two lines contradicted each other on the same phrase. Quick Reference said access is observable within the first twenty minutes, then twenty lines later said access is least understood during the first twenty minutes. The second now makes its point about defaulting high without re-litigating when access becomes knowable. Step 2 explained the memory tradeoff twice in consecutive paragraphs. General principle first, then the already-powered-off case as a short consequence. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 30 +++++++------------ 1 file changed, 11 insertions(+), 19 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index 1f5c6afeb..b9cd1211d 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -67,8 +67,8 @@ is proposing. The third turns one laptop into a campaign against users. Until the inventory in [Credential Revocation Order](#credential-revocation-order) is complete, treat the -incident at the higher tier. Access is least understood during the first twenty minutes, which is exactly -when severity gets set. +incident at the higher tier. Severity gets set before anyone has finished listing what the user could +reach. This runbook stops at P3. The policy also defines P4 and P5, but their response times ("can be scheduled" and "no immediate action") do not fit a confirmed endpoint compromise, where a stolen session cookie stays @@ -160,16 +160,15 @@ Pick the one row that matches your capability: | No tooling, memory capture possible within [X hours] | Disconnect the network, do not sleep or shut down | | No tooling, no memory capture possible | Disconnect the network, then power off fully | -Expect the first row often. [Malware Infection](/incident-management/playbooks/malware) tells the affected -person to power off immediately, so a user who followed that guidance, or who simply panicked, hands you a -cold machine. Memory is already gone at that point and booting the device only runs the malware's -persistence again, so the rest of this runbook proceeds from the credential side. - Powering off destroys running processes, injected code, and secrets decrypted in memory, which is the evidence that tells you what actually ran and therefore what must be rotated. It also stops an active thief. Choose speed when nobody is positioned to capture memory: a machine left running on a desk "for forensics" that nobody ever images is the worst of both options. +Expect the first row often, since [Malware Infection](/incident-management/playbooks/malware) tells the +affected person to power off immediately. Booting the device only runs the persistence again, so that case +proceeds from the credential side. + - [ ] Network access cut, by console containment or physical disconnect - [ ] Device not wiped, not "cleaned", and not casually restarted @@ -194,12 +193,6 @@ reliably under pressure. **If you have endpoint management:** pull the installed application list, the extension inventory, and recent process history instead of relying on the user's memory. -### Step 4: Start the revocation clock - -Move straight to [Credential Revocation Order](#credential-revocation-order). Session revocation is time -sensitive: a stolen session cookie keeps working until the session itself is killed, no matter how many -passwords change. - ## Credential Revocation Order **Why order matters:** the usual failures are sequencing errors. Rotating an SSH key before killing the @@ -223,7 +216,7 @@ Work top to bottom. Each row links to whoever owns the procedure. - [ ] Every row above assigned to a named person - [ ] All revocation done from a known-clean device -- [ ] Persistence checked (mail rules, OAuth grants, deploy keys, new signers) +- [ ] Row 9 actually worked, not only credentials rotated - [ ] Anything not revocable within [X hours] escalated @@ -242,7 +235,7 @@ Work top to bottom. Each row links to whoever owns the procedure. ### Evidence to Collect Collect in order of volatility, following [NIST SP 800-86](https://csrc.nist.gov/pubs/sp/800/86/final). -The top rows vanish on their own; the lower ones keep. If the device is still running, memory outranks +The top rows vanish on their own; the rest wait for you. If the device is still running, memory outranks everything. | Order | Data | Source | @@ -256,10 +249,9 @@ everything. | 7 | The lure itself | Message, repository, installer, meeting link | | 8 | On-chain activity | Block explorer, for addresses the user controls | -Rows 1 and 2 are gone the moment the device powers off, which is why Step 2 forks on whether anyone can -capture them. Rows 4 through 6 depend on retention windows you set long before the incident. See -[Forensic Readiness](/incident-management/forensic-readiness) for chain of custody and retention design, -which only helps if it is already in place when the incident starts. +Rows 1 and 2 are why Step 2 forks on memory capture. Rows 4 through 6 depend on retention windows set long +before the incident. See [Forensic Readiness](/incident-management/forensic-readiness) for chain of custody +and retention design, which only helps if it is already in place when the incident starts. If the same lure reached other contributors, treat that wider exposure as its own incident and send the team the specific indicators to look for. From 9996ce4be5475ba4f56673abb67b85326babf443 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Tue, 22 Sep 2026 17:35:14 +0900 Subject: [PATCH 07/11] docs(incident-management): match heading case to sibling runbooks Six headings were sentence case while the ones inherited from the template were title case, so the file was inconsistent with itself and with all nine sibling runbooks, which title-case step headings too. Claude-Session: https://claude.ai/code/session_01CPQ1FcHGiBKWPd8waT7uGx --- .../runbooks/endpoint-compromise.mdx | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index b9cd1211d..0fc68d1a9 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -125,7 +125,7 @@ follow-on accounts to check. ## Immediate Actions -### Step 1: Reach the user out of band +### Step 1: Reach the User Out of Band **Why:** Their chat, email, and phone accounts may already be in the attacker's hands. @@ -146,7 +146,7 @@ leave the device alone. > [Legal](/incident-management/incident-response-template/contacts#legal--communications) before contact. > See [Mitigating DPRK IT Workers](/dprk-it-workers/mitigating-dprk-it-workers). -### Step 2: Contain the device +### Step 2: Contain the Device **Why:** Stop live session abuse and further exfiltration without destroying what you need to scope the incident. @@ -177,7 +177,7 @@ proceeds from the credential side. [Incident Log](/incident-management/incident-response-template/templates/incident-log-template) -### Step 3: Scope what the device could reach +### Step 3: Scope What the Device Could Reach **Why:** Severity, revocation order, and blast radius all depend on this list, and nobody recalls it reliably under pressure. @@ -263,7 +263,7 @@ nobody closes buries the next real one. ## Recovery and Prevention -### Rebuild, do not clean +### Rebuild, Do Not Clean Finish identifying every foothold before you start removing any of them. Rebuilding the laptop while an OAuth grant, mail rule, or deploy key survives moves the attacker off the device and leaves them in the @@ -280,7 +280,7 @@ infostealer is gone, and the cost of being wrong is handing back signing authori - [ ] Device retained if forensic or regulatory obligations require it -### Return-to-service gate +### Return-to-Service Gate Privileged access returns only when a named approver confirms all of the following: @@ -308,7 +308,7 @@ managed hardware, contractors on VDI, everyone else behind a managed browser. - [ ] Contributors know where to report, including out of hours -### Closing the incident +### Closing the Incident The rebuild is the midpoint. [NIST SP 800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats improvement as a standing function rather than a meeting held once the noise dies down, so the work below From 45c90a757c5125e4ccb01296ebddb5e53781e165 Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Wed, 7 Oct 2026 21:15:01 +0900 Subject: [PATCH 08/11] docs(incident-management): rework endpoint runbook after review Addresses the steward review on #647 plus a blind cold-read, a fact-check of the external citations, and a repo-wide duplication audit. Citation was wrong. The volatility ordering credited NIST SP 800-86, which never uses the phrase "order of volatility" and puts network connections and login sessions ahead of memory because connections time out. The term is RFC 3227's. Memory still leads here, since a full image captures connections, sessions, and the process table, but that is now stated as the reason rather than dressed up as NIST guidance. Two orderings were backwards. The insider-threat warning sat below the instruction to phone the user, so a responder working top to bottom had already tipped off a suspected DPRK IT worker. The foothold-inventory gate sat seventy lines after the revocation table it was meant to gate. Both moved ahead of what they govern. Session logs are now exported before revocation. Killing sessions can truncate the sign-in records that show where the attacker connected from, and the runbook previously collected them two sections later. Severity drops to P1 and P2. Judging a colleague's access tier from memory at 3am is how a signer gets filed as P3, and the review made the case that broad internal access can outrank a signing key. Unclear access now defaults to P1, and the separate escalator checklist folds into the P1 row. Step 1 splits into an org-side track that needs nothing from the user and a person-side track that does, because the account is the urgent part and the previous order made contact a prerequisite. Adds what to do when the user stays unreachable, which the old text left open after telling responders to keep going. Adds the step that confirms compromise, which the runbook assumed and never described, with a path for unmanaged devices that produce no EDR signal. Insider determination stays with Legal and HR. Overstated claims hedged: cookie survival is identity-provider dependent, Entra revokes on password change; powering off does not erase memory that hibernation and swap files still hold, so image the disk either way; cloud 2FA backup restores the same seed rather than one necessarily stolen. Also: wallet assets must move, not just the session, when a hot wallet is exposed; post-mortem gate follows the template's P1-P3 rather than narrowing it; every bracketed value carries a suggested default and a label, since three different ones all read "[X hours]"; known-clean device is defined where four steps depend on it; headings are sentence case per the style guide, which the sibling runbooks predate. The password manager guide told responders to wipe a device that is "suspected compromised", the opposite first action, which would destroy the evidence this runbook collects. Split from the lost-or-stolen case. --- .../password-manager-endpoint-hardening.mdx | 6 +- .../runbooks/endpoint-compromise.mdx | 358 ++++++++++-------- .../runbooks/index.mdx | 2 +- wordlist.txt | 1 + 4 files changed, 210 insertions(+), 157 deletions(-) diff --git a/docs/pages/guides/endpoint-security/password-manager-endpoint-hardening.mdx b/docs/pages/guides/endpoint-security/password-manager-endpoint-hardening.mdx index 4657ea1f3..982fff74b 100644 --- a/docs/pages/guides/endpoint-security/password-manager-endpoint-hardening.mdx +++ b/docs/pages/guides/endpoint-security/password-manager-endpoint-hardening.mdx @@ -85,7 +85,11 @@ endpoints that can unlock the vault accordingly. ## Lost or Stolen Device Response -Use these steps if a device with possible vault access is lost, stolen, or suspected compromised: +Use these steps if a device with possible vault access is lost or stolen. If the device is instead +suspected compromised and still in hand, do not wipe it: wiping destroys the memory and disk evidence that +determines what was taken. Preserve it and follow +[Endpoint Compromise](/incident-management/incident-response-template/runbooks/endpoint-compromise), which +also sets the order for steps 2 to 5 below. 1. Lock, mark lost, or wipe the device as quickly as possible. 2. Revoke the device or active sessions from the password manager or identity provider if the product supports it. diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index 0fc68d1a9..77967ca13 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -1,5 +1,5 @@ --- -title: "Runbook: Endpoint Compromise | Security Alliance" +title: "Runbook: Endpoint compromise | Security Alliance" description: "Responder-side runbook for a compromised contributor workstation: contain the device, scope its access, revoke credentials in order, and rebuild." tags: - Security Specialist @@ -7,7 +7,7 @@ tags: - DevOps contributors: - role: wrote - users: [] + users: [s1ns3nz0] - role: reviewed users: [] - role: fact-checked @@ -16,7 +16,7 @@ contributors: import { TagList, AttributionList, ContributeFooter, Checklist } from '../../../../../components' -# Runbook: Endpoint Compromise +# Runbook: Endpoint compromise @@ -25,54 +25,51 @@ import { TagList, AttributionList, ContributeFooter, Checklist } from '../../../ > incident. Contain the device, inventory what it could reach, revoke in the right order, > and rebuild rather than clean. -**This is an example runbook.** Review and customize for your organization before use. Fill in your -isolation procedure, credential inventory, and named approvers. +**This is an example runbook.** Review and customize for your organization before use. Fill in the +containment tools, the collection tools, the named owners in the revocation table, and the approver. +Decide every `[bracketed]` value below in advance; an incident is the wrong time to pick one. This runbook covers a contributor workstation (laptop or desktop, macOS/Windows/Linux) that is suspected or confirmed compromised: infostealer malware, a malicious dependency or fake job "test task", a trojanized -meeting client, or hands-on access. Servers and mobile devices are out of scope, since server containment -usually means snapshot and quarantine. +meeting client, or hands-on access. Servers are out of scope, since server containment usually means +snapshot and quarantine. A mobile device is out of scope as the compromised device, but a phone holding the +user's authenticator is in scope as a credential, in rows 4 and 9 of the revocation table. The person on the other end of this runbook is a colleague, and they may be frightened. Hand them -[Malware Infection](/incident-management/playbooks/malware) for the victim-facing side of the same incident. +[Malware Infection](/incident-management/playbooks/malware) after containment, as the victim-facing side of +the same incident. Do not hand it over while you still need their attention for Step 1. -## Quick Reference +## Quick reference | Field | Value | | ------- | ------- | -| **Typical Severity** | Varies by access tier (see below) | -| **Primary Responder** | Security SME | -| **Last Updated** | [Date] | +| **Typical severity** | P1 or P2, by the access the user holds | +| **Primary responder** | Security SME | +| **Last updated** | [Date] | | **Owner** | [Name] | -Severity follows what the device could reach. Access is observable -within the first twenty minutes and impact is not, so this table stands in for the impact bands defined in -the [Incident Response Policy](/incident-management/incident-response-template/incident-response-policy). - -| Access tier held by the user | Severity | What the user holds | What the attacker gains | -| ------------------------------ | ---------- | --------------------- | ------------------------- | -| Signing, deployer, or production admin | P1 | Multisig signer key, contract upgrade or owner key, deployer EOA, hot wallet, cloud root, domain registrar, exchange withdrawal credentials | Direct fund movement or a privileged transaction, meeting the policy's fund-loss condition on the spot | -| Broad internal access | P2 | CI/CD and package publish tokens, repository write, cloud IAM write, internal admin panels, identity provider admin, official announcement channels | Reach production through a build, change what users are served, or pivot toward a P1 holder | -| Limited or no privileged access | P3 | Email, chat, shared documents, read-only dashboards | Impersonate a trusted colleague against the rows above, plus everything the user could read | - -Escalate a tier when any of these hold. The first two amount to holding a deployer key: a pipeline with no -approval gate deploys on the attacker's behalf, and a threshold only delays execution once a trusted signer -is proposing. The third turns one laptop into a campaign against users. - - -- [ ] Pipeline or repo write reaches production without approval (P1, see - [Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise)) -- [ ] Multisig signer below threshold (P1) -- [ ] Controls official social, announcements, or docs publishing (P2 minimum) - +A **known-clean device** means one that was outside the compromise window and is not the machine under +investigation: a freshly imaged spare, an administrator workstation, or a phone that was never paired to the +suspect machine. Identify one in advance. Sourcing it mid-incident costs hours, and four steps below depend +on having it. + +Severity follows what the device could reach. Access is observable within the first twenty minutes and +impact is not, so this table maps access onto the impact bands in the +[Incident Response Policy](/incident-management/incident-response-template/incident-response-policy). + +| Access the user holds | Severity | Examples | What the attacker gains | +| ----------------------- | ---------- | ---------- | ------------------------- | +| Signing, deployer, or production admin, including any repository or pipeline path that reaches production without a human approval gate | P1 | Multisig signer key, contract upgrade or owner key, deployer EOA, hot wallet, cloud root, domain registrar, exchange withdrawal credentials, CI/CD secrets that deploy unreviewed | Move funds or land code in production, meeting the policy's fund-loss condition on the spot | +| Everything else | P2 | Repository write behind review, cloud IAM write, internal admin panels, identity provider admin, official announcement channels, email, chat, shared documents | Impersonate a trusted colleague against the row above, publish to users, or pivot toward a P1 holder | -Until the inventory in [Credential Revocation Order](#credential-revocation-order) is complete, treat the -incident at the higher tier. Severity gets set before anyone has finished listing what the user could -reach. +A multisig signer below threshold is still P1: one more compromised signer executes. A pipeline with no +approval gate is P1 because it deploys on the attacker's behalf, as in +[Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise). -This runbook stops at P3. The policy also defines P4 and P5, but their response times ("can be scheduled" -and "no immediate action") do not fit a confirmed endpoint compromise, where a stolen session cookie stays -valid until someone revokes it. +**If you cannot tell which row applies, use P1.** Severity is set before anyone has finished the access +inventory in [Step 3](#step-3-scope-what-the-device-could-reach), and in an organization of any size you will +not know a colleague's access from memory. There is no tier below P2 here: a stolen session cookie does not +wait for a scheduled response. ## Identification @@ -94,9 +91,13 @@ Extension and plugin installs deserve their own check: a malicious one reads eve holds, and a repository can pressure the install through its own recommendations. See [Integrated Development Environments](/devsecops/integrated-development-environments). -**If you have EDR:** a single quarantined detection still triggers this runbook, since quarantine only -proves the product caught one payload. Treat the signals below as confirmation, and remember that an -unmanaged device produces none of them. +### Confirming compromise + +One symptom is enough to start. This section is about whether you can stop. + +**With EDR**, work the console: process ancestry for anything spawned from a browser, archive, or downloads +folder; credential-store access; and the agent's own health. A single quarantined detection still runs this +runbook, since quarantine proves the product caught one payload. - [ ] Agent stopped reporting, disabled, or tamper alert @@ -107,79 +108,128 @@ unmanaged device produces none of them. - [ ] New extension or plugin making outbound calls +**Without EDR**, you will not get a clean confirmation, so do not wait for one. Check the four places +persistence lands and the two the user controls, from the device's own interface before it is disconnected, +or from the disk image afterwards. + + +- [ ] Login items and launch agents or daemons +- [ ] Scheduled tasks and cron entries +- [ ] Shell profile files and PATH overrides +- [ ] Browser and editor extensions, against what the user remembers installing +- [ ] Identity provider sign-in history for logins the user does not recognize +- [ ] What the user actually ran, in their words, with the file or link if they still have it + + +Absent evidence is not absence. If the user executed something and you cannot rule out what it did, treat +the device as compromised and continue. + ### Differentiation If a wallet drained but the device shows no sign of compromise, the user most likely signed a malicious -approval. Use [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) and -the [Drainer playbook](/incident-management/playbooks/hacked-drainer) instead. If several contributors are -affected at once with no shared file or link, suspect the identity provider or the supply chain rather than -a single endpoint. - -Two of the lures above are documented campaigns. Both playbooks carry screenshots of the messages -attackers sent, so you can put them in front of the user and ask whether theirs matched. A meeting link -that resolved to an unfamiliar domain, or a fake recruiter task, matches +approval. Stop here and switch to +[Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) and the +[Drainer playbook](/incident-management/playbooks/hacked-drainer). If several contributors are affected at +once with no shared file or link, suspect the identity provider or the supply chain rather than a single +endpoint. + +Two of the lures above are documented campaigns. Both playbooks carry screenshots of the messages attackers +sent, so you can put them in front of the user and ask whether theirs matched. A meeting link that resolved +to an unfamiliar domain, or a fake recruiter task, matches [North Korea (DPRK) Attack](/incident-management/playbooks/hacked-dprk). A Zoom call where the caller asked for screen share or remote control matches -[ELUSIVE COMET](/incident-management/playbooks/hacked-elusive-comet). Naming the campaign tells you which -follow-on accounts to check. +[ELUSIVE COMET](/incident-management/playbooks/hacked-elusive-comet). Read either alongside this runbook +rather than instead of it: they identify the campaign and tell you which follow-on accounts to check, while +the steps below stay the same. + +## Immediate actions -## Immediate Actions +### Step 1: Cut access and reach the user -### Step 1: Reach the User Out of Band +> **Before you call.** If the user may be the threat actor (fake contributor, suspected DPRK IT worker, +> hostile departure), do not call them and do not reveal that you detected anything. Preserve access logs +> and device state, cut access quietly, and engage +> [Legal](/incident-management/incident-response-template/contacts#legal--communications) first. Confirming +> that suspicion is a Legal and HR judgement rather than a responder's; see +> [Mitigating DPRK IT Workers](/dprk-it-workers/mitigating-dprk-it-workers). -**Why:** Their chat, email, and phone accounts may already be in the attacker's hands. +**Why:** The account is the urgent part, and the user's chat, email, and phone may already be in the +attacker's hands. - +Two tracks, two people, the same minute. The org-side track needs nothing from the user. + + +- [ ] Identity provider sign-in and session logs exported, before anything is revoked +- [ ] Sessions killed at identity provider, chat, email, and code host +- [ ] Cloud and production access suspended +- [ ] [Decision Makers](/incident-management/incident-response-template/contacts#decision-makers) paged if + the access is or may be P1 + + + - [ ] Call on a channel not tied to the suspect device - [ ] Confirm identity by voice, never by text alone -- [ ] Tell them to stop using it and leave it powered as-is +- [ ] Tell them to stop using the device and leave it powered as it is +- [ ] Tell them to unplug hardware wallets and security keys now, not later - [ ] Say plainly that blame is not the point -If the user does not answer, keep going. Session revocation, identity provider lockout, and cloud access -cuts all work without them, and the account is the urgent part. Leave a message with one instruction: -leave the device alone. +Export the logs before revoking: killing sessions can truncate the sign-in records that show where the +attacker connected from, and they do not come back. Revocation also tells the attacker they were seen, so +close fund-movement paths first and leave the noisier account cleanup for the table below. -> **If the user may be the threat actor** (fake contributor, suspected DPRK IT worker, hostile departure): -> do not call them and do not reveal that you detected anything. Preserve access logs and device state, cut -> access quietly, and engage -> [Legal](/incident-management/incident-response-template/contacts#legal--communications) before contact. -> See [Mitigating DPRK IT Workers](/dprk-it-workers/mitigating-dprk-it-workers). +If the user does not answer, keep going. Leave one instruction: leave the device alone. Rows 3, 4, and 8 of +the revocation table do need them, so if they are still unreachable after +[contact timeout, 4 hours suggested], suspend the account outright rather than leaving it half-rotated. -### Step 2: Contain the Device +### Step 2: Contain the device and capture what disappears -**Why:** Stop live session abuse and further exfiltration without destroying what you need to scope the -incident. +**Why:** Stop live session abuse without destroying what you need to scope the incident. -Pick the one row that matches your capability: +Find the row that matches your situation: | Situation | Action | | ----------- | -------- | | Device already powered off | Leave it off, and do not boot it to "check something" | -| Endpoint tooling with network containment | Contain the host in the console, leave it powered for memory capture | -| No tooling, memory capture possible within [X hours] | Disconnect the network, do not sleep or shut down | -| No tooling, no memory capture possible | Disconnect the network, then power off fully | +| Endpoint tooling with network containment | Contain the host in the console, leave it powered for capture | +| No tooling, someone can capture memory within [memory capture window, 2 hours suggested] | Disconnect the network, do not sleep or shut down | +| No tooling, nobody can capture memory | Disconnect the network, then power off fully | + +Disconnect means unplug the cable and turn off the wireless radio from a hardware switch where one exists. +Toggling Wi-Fi from inside the operating system means working on a live compromised machine. -Powering off destroys running processes, injected code, and secrets decrypted in memory, which is the -evidence that tells you what actually ran and therefore what must be rotated. It also stops an active -thief. Choose speed when nobody is positioned to capture memory: a machine left running on a desk "for -forensics" that nobody ever images is the worst of both options. +Powering off ends access to running processes, injected code, and secrets decrypted in memory, which is the +evidence that tells you what actually ran and therefore what must be rotated. It also stops an active thief. +Choose speed when nobody is positioned to capture memory: a machine left running on a desk "for forensics" +that nobody ever images is the worst of both options. Expect the first row often, since [Malware Infection](/incident-management/playbooks/malware) tells the -affected person to power off immediately. Booting the device only runs the persistence again, so that case -proceeds from the credential side. +affected person to power off immediately. Booting it runs the persistence again, so that case proceeds from +the credential side. Image the disk anyway: hibernation and swap files (`hiberfil.sys`, `pagefile.sys`, +macOS `sleepimage`) often still hold decrypted secrets and process state. + +While the device is still live, collect in this order. Decide the tools in advance; choosing one +mid-incident is how memory gets lost. + +| Order | Data | Source | Tool | +| ------- | ------ | -------- | ------ | +| 1 | Memory image | The live device, before any shutdown | [tool and command] | +| 2 | Live network state and running processes | The live device | [tool and command] | +| 3 | Disk image | The device, before rebuild | [tool and command] | + +Most volatile first, per [RFC 3227](https://www.rfc-editor.org/rfc/rfc3227). Memory leads because a full +image also captures connections, login sessions, and the process table. - [ ] Network access cut, by console containment or physical disconnect - [ ] Device not wiped, not "cleaned", and not casually restarted - [ ] Hardware wallets, security keys, and external drives unplugged and set aside -- [ ] Time and method of containment recorded in the - [Incident Log](/incident-management/incident-response-template/templates/incident-log-template) +- [ ] Time and method of containment recorded -### Step 3: Scope What the Device Could Reach +### Step 3: Scope what the device could reach -**Why:** Severity, revocation order, and blast radius all depend on this list, and nobody recalls it +**Why:** Severity, revocation order, and blast radius all depend on this inventory, and nobody recalls it reliably under pressure. @@ -187,42 +237,49 @@ reliably under pressure. - [ ] Keys held, and signer roles the user can approve - [ ] Paired devices (2FA apps, security keys, hardware wallets) - [ ] Plaintext credentials on disk (.env files, cloud and SSH config, notes apps) -- [ ] Severity set from the access tier table, defaulting high +- [ ] Severity set from the table above, P1 wherever the access is unclear **If you have endpoint management:** pull the installed application list, the extension inventory, and recent process history instead of relying on the user's memory. -## Credential Revocation Order +## Credential revocation order + +Finish the access inventory in [Step 3](#step-3-scope-what-the-device-could-reach) before removing +footholds. Revoking what you have found while an OAuth grant or mail rule you have not found survives moves +the attacker out of the device and leaves them in the accounts. **Why order matters:** the usual failures are sequencing errors. Rotating an SSH key before killing the attacker's live session lets them add the new key themselves. Changing a password before revoking sessions -leaves the stolen cookie valid. Re-enrolling 2FA from a cloud backup restores the stolen seed. +leaves the stolen cookie valid on any identity provider that does not revoke sessions on password change, +which is most of them. Re-enrolling 2FA from a cloud backup restores the same seed, which the attacker may +already hold. -Work top to bottom. Each row links to whoever owns the procedure. +Rows 1 and 2 run together at P1. Otherwise work top to bottom. Assign every `[Name]` before an incident: +an unassigned row is where revocation stalls at 3am. | # | Credential class | Where | Owner | Notes | | --- | ------------------ | ------- | ------- | ------- | -| 1 | Active sessions | Identity provider, chat, email, code host | [Name] | Kill all sessions first; cookie theft is the default infostealer outcome | -| 2 | Signing and deployer keys | Multisig, contracts, treasury | [Name] | Run in parallel at P1 per [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) | -| 3 | Password manager | Vault account | [Name] | Assume vault contents exposed; rotate master password and re-enroll 2FA from a clean device | +| 1 | Active sessions | Identity provider, chat, email, code host | [Name] | Started in Step 1; confirm every system is covered. Session revocation does not reach OAuth tokens, so revoke those separately | +| 2 | Signing and deployer keys | Multisig, contracts, treasury | [Name] | Run alongside row 1 at P1, per [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise). For a hot wallet, revoking the session is not enough: move the assets to a wallet generated on a known-clean device | +| 3 | Password manager | Vault account | [Name] | Assume vault contents exposed; rotate master password and re-enroll 2FA from a known-clean device | | 4 | MFA enrollments | Identity provider, exchanges | [Name] | Re-enroll on a new device; never restore an authenticator from cloud backup | -| 5 | SSH and git access | Code host, servers | [Name] | Remove old public keys and look for attacker-added keys and deploy keys | +| 5 | SSH and git access | Code host, servers | [Name] | Remove old public keys and look for attacker-added keys and deploy keys. Reissue per [SSH client and key management hardening](/guides/endpoint-security/ssh-client-and-key-management-hardening) | | 6 | Cloud and CI/CD tokens | Cloud IAM, CI secrets, registries | [Name] | See [Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise) | | 7 | API and exchange keys | Exchanges, data providers | [Name] | Withdrawal-capable keys first | -| 8 | Passwords | Everything above | [Name] | Last, after sessions are dead, from a clean device | +| 8 | Passwords | Everything above | [Name] | Last, after sessions are dead, from a known-clean device | | 9 | Recovery channels | Mail rules, recovery codes, phone, OAuth grants | [Name] | Attackers persist here; check rules and grants, not just credentials | - [ ] Every row above assigned to a named person - [ ] All revocation done from a known-clean device -- [ ] Row 9 actually worked, not only credentials rotated -- [ ] Anything not revocable within [X hours] escalated +- [ ] Row 9 re-checked after the rest, not only rotated +- [ ] Anything not revocable within [revocation deadline, 4 hours suggested] escalated ## Investigation -### Key Questions +### Key questions - [ ] What was the initial access (link, file, dependency, installer, physical)? @@ -232,42 +289,59 @@ Work top to bottom. Each row links to whoever owns the procedure. - [ ] Are other contributors exposed to the same lure? -### Evidence to Collect +### Evidence to collect -Collect in order of volatility, following [NIST SP 800-86](https://csrc.nist.gov/pubs/sp/800/86/final). -The top rows vanish on their own; the rest wait for you. If the device is still running, memory outranks -everything. +The volatile sources are in Step 2; they are gone by the time you reach this section. -| Order | Data | Source | -| ------- | ------ | -------- | -| 1 | Memory image | The live device, before any shutdown | -| 2 | Live network state and running processes | The live device | -| 3 | Disk image | The device, before rebuild | -| 4 | Persistence and install history | Endpoint tooling, local logs | -| 5 | Authentication and session logs | Identity provider, code host, cloud | -| 6 | Outbound network telemetry | DNS, proxy, and firewall logs, if retained | -| 7 | The lure itself | Message, repository, installer, meeting link | -| 8 | On-chain activity | Block explorer, for addresses the user controls | +| Data | Source | +| ------ | -------- | +| Persistence and install history | Endpoint tooling, local logs, or the disk image | +| Authentication and session logs | Exported in Step 1, before revocation | +| Outbound network telemetry | DNS, proxy, and firewall logs, if retained | +| The lure itself | Message, repository, installer, meeting link | +| On-chain activity | Block explorer, for addresses the user controls | -Rows 1 and 2 are why Step 2 forks on memory capture. Rows 4 through 6 depend on retention windows set long -before the incident. See [Forensic Readiness](/incident-management/forensic-readiness) for chain of custody -and retention design, which only helps if it is already in place when the incident starts. +Retention is decided long before an incident. See +[Forensic Readiness](/incident-management/forensic-readiness) for chain of custody and retention design, +which only helps if it is already in place when the incident starts. -If the same lure reached other contributors, treat that wider exposure as its own incident and send the -team the specific indicators to look for. +If the same lure reached other contributors, treat that wider exposure as its own incident and send the team +the specific indicators to look for. -If the investigation shows the lure arrived but was never executed, and no credential exposure is -confirmed, downgrade the incident and close it. Record what ruled execution out in the +If the lure arrived but was never executed, and no credential exposure is confirmed, downgrade the incident +and close it. Record what ruled execution out in the [Incident Log](/incident-management/incident-response-template/templates/incident-log-template). A P1 that nobody closes buries the next real one. -## Recovery and Prevention +## Escalation -### Rebuild, Do Not Clean + +- [ ] [Decision Makers](/incident-management/incident-response-template/contacts#decision-makers) - + immediately at P1, as started in Step 1 +- [ ] [Security Partners](/incident-management/incident-response-template/contacts#security-partners) - for + forensics, or whenever lateral movement is suspected +- [ ] [Legal](/incident-management/incident-response-template/contacts#legal--communications) - if funds + were stolen, data was exposed, or insider involvement is suspected +- [ ] [SEAL 911](https://t.me/seal_911_bot) - for crypto-specific containment help beyond your team + -Finish identifying every foothold before you start removing any of them. Rebuilding the laptop while an -OAuth grant, mail rule, or deploy key survives moves the attacker off the device and leaves them in the -accounts, and the rebuild destroys the evidence that would have found them. +Legal owns external notification, but the clock starts at discovery and runs while you are still containing. +Raise it early, while the timing is still yours to pick. For wording, cadence, and who speaks, use +[Communications](/incident-management/incident-response-template/communications) and +[Communication Strategies](/incident-management/communication-strategies). + + +- [ ] Disclosure obligations checked against your jurisdictions and + [reporting window, 72 hours suggested] +- [ ] Affected users or customers notified if their data or funds were reachable +- [ ] Law enforcement report filed if funds were stolen (see + [Malware Infection](/incident-management/playbooks/malware) for the reporting venues) +- [ ] Partners and exchanges told which addresses and accounts to watch + + +## Recovery and prevention + +### Rebuild, do not clean Wipe and rebuild, or replace the hardware. Do not restore the user profile from a backup taken after the compromise window opened, and do not accept "the antivirus removed it" as an outcome. You cannot prove an @@ -280,35 +354,34 @@ infostealer is gone, and the cost of being wrong is handing back signing authori - [ ] Device retained if forensic or regulatory obligations require it -### Return-to-Service Gate +### Return-to-service gate Privileged access returns only when a named approver confirms all of the following: -- [ ] Device rebuilt or replaced, and enrolled in management -- [ ] Every credential class above reissued from a clean device +- [ ] Device rebuilt or replaced, and enrolled in management where the organization has it +- [ ] Every credential class above reissued from a known-clean device - [ ] MFA re-enrolled fresh, not restored from a backup - [ ] Recovery channels re-verified (mail rules, recovery codes, phone, OAuth grants) -- [ ] Signing and deployer roles restored last, after [X days, 30 suggested] +- [ ] Signing and deployer roles restored last, after [monitoring period, 30 days suggested] - [ ] Approver: [Name] ### Prevention -Device tiers belong to [Endpoint Security](/opsec/endpoint/overview): signers and production admins on -managed hardware, contractors on VDI, everyone else behind a managed browser. - +- [ ] Device tier matched to role risk, per [Endpoint Security](/opsec/endpoint/overview) - [ ] Signing devices separate from daily-driver workstations -- [ ] Device tier matched to role risk, EDR and disk encryption on P1 holders +- [ ] EDR and disk encryption on the devices of P1 access holders - [ ] Extensions allowlisted, repo-recommended ones declined by default - [ ] Untrusted code run only in a sandbox or disposable VM - [ ] Phishing-resistant MFA on high-value accounts - [ ] Shortened session lifetimes for privileged tooling +- [ ] A known-clean device identified and kept ready - [ ] Contributors know where to report, including out of hours -### Closing the Incident +### Closing the incident The rebuild is the midpoint. [NIST SP 800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats improvement as a standing function rather than a meeting held once the noise dies down, so the work below @@ -319,38 +392,13 @@ the blameless format. - [ ] Timeline written up in the [Incident Log](/incident-management/incident-response-template/templates/incident-log-template) - [ ] Return-to-service gate signed off by the named approver -- [ ] Post-mortem scheduled if the compromise reached P1 or P2 - ([template](/incident-management/incident-response-template/templates/post-mortem-template)) +- [ ] Post-mortem scheduled per the + [Post-Mortem Template](/incident-management/incident-response-template/templates/post-mortem-template) - [ ] Indicators shared with the team, and with partners if the lure was reused against them - [ ] Prevention items above converted into owned work, not left in this checklist - [ ] Detection gap reviewed: what would have caught this a day earlier -## Escalation - - -- [ ] [Decision Makers](/incident-management/incident-response-template/contacts#decision-makers) - - immediately for any P1 access tier -- [ ] [Security Partners](/incident-management/incident-response-template/contacts#security-partners) - for - forensics, or whenever lateral movement is suspected -- [ ] [Legal](/incident-management/incident-response-template/contacts#legal--communications) - if funds - were stolen, data was exposed, or insider involvement is suspected -- [ ] [SEAL 911](https://t.me/seal_911_bot) - for crypto-specific containment help beyond your team - - -Legal owns external notification, but the clock starts at discovery and runs while you are still -containing. Raise it early, while the timing is still yours to pick. For wording, cadence, and who speaks, -use [Communications](/incident-management/incident-response-template/communications) and -[Communication Strategies](/incident-management/communication-strategies). - - -- [ ] Disclosure obligations checked against your jurisdictions and [X hours] reporting windows -- [ ] Affected users or customers notified if their data or funds were reachable -- [ ] Law enforcement report filed if funds were stolen (see - [Malware Infection](/incident-management/playbooks/malware) for the reporting venues) -- [ ] Partners and exchanges told which addresses and accounts to watch - - ## Further reading - [Runbooks overview](/incident-management/incident-response-template/runbooks/overview): how the runbooks in this section fit together diff --git a/docs/pages/incident-management/incident-response-template/runbooks/index.mdx b/docs/pages/incident-management/incident-response-template/runbooks/index.mdx index 8cbd801e5..f18bf84d4 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/index.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/index.mdx @@ -14,7 +14,7 @@ title: "Runbooks" - [Runbooks](/incident-management/incident-response-template/runbooks/overview) - [Runbook: Smart Contract Exploit](/incident-management/incident-response-template/runbooks/smart-contract-exploit) - [Runbook: Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise) -- [Runbook: Endpoint Compromise](/incident-management/incident-response-template/runbooks/endpoint-compromise) +- [Runbook: Endpoint compromise](/incident-management/incident-response-template/runbooks/endpoint-compromise) - [Runbook: Frontend Compromise](/incident-management/incident-response-template/runbooks/frontend-compromise) - [Runbook: DNS Hijack](/incident-management/incident-response-template/runbooks/dns-hijack) - [Runbook: CDN/Hosting Compromise](/incident-management/incident-response-template/runbooks/cdn-hosting-compromise) diff --git a/wordlist.txt b/wordlist.txt index 74bad3de3..d23a025a2 100644 --- a/wordlist.txt +++ b/wordlist.txt @@ -152,6 +152,7 @@ halmos Hdiv Helius Hetrix +hiberfil hevm HIDS HIPAA From ed831d1e8f0c5796cb9159d5dae61539c03bd6ab Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Wed, 7 Oct 2026 21:20:21 +0900 Subject: [PATCH 09/11] docs(incident-management): match checklist voice to sibling runbooks The runbooks state actions in the imperative and keep items short: "Remove compromised signer", "Rotate all secrets and tokens", and a Prevention section of bare noun phrases. This runbook wrote its action checklists as past participles instead, which read as verification rather than instruction and ran twice as long. Converted the action checklists to the imperative and Prevention to noun phrases, leaving the two that really are verification gates, the revocation confirmation and the return-to-service gate, as participles. Average checklist item is down from 8.8 words to 7.9; siblings sit at 4.1 to 4.4. Also cut eight blocks of prose that justified a decision rather than changing one: the section preamble to Confirming compromise, the restatement of why memory matters, the RFC 3227 note the numbered table already carries, the remark about unassigned owners, and the NIST citation behind Closing the incident, whose point is made by the sentence before it. Prose is 109 lines against 25 to 27 in the siblings. That gap is the revocation reasoning, the containment fork, and the step that confirms compromise; the siblings are shorter partly because they are thinner. --- .../runbooks/endpoint-compromise.mdx | 127 ++++++++---------- 1 file changed, 59 insertions(+), 68 deletions(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index 77967ca13..bcd6ffebc 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -48,10 +48,9 @@ the same incident. Do not hand it over while you still need their attention for | **Last updated** | [Date] | | **Owner** | [Name] | -A **known-clean device** means one that was outside the compromise window and is not the machine under -investigation: a freshly imaged spare, an administrator workstation, or a phone that was never paired to the -suspect machine. Identify one in advance. Sourcing it mid-incident costs hours, and four steps below depend -on having it. +A **known-clean device** means one outside the compromise window and not the machine under investigation: +a freshly imaged spare, an admin workstation, or a phone never paired to it. Four steps below need one, so +identify it in advance. Severity follows what the device could reach. Access is observable within the first twenty minutes and impact is not, so this table maps access onto the impact bands in the @@ -62,9 +61,8 @@ impact is not, so this table maps access onto the impact bands in the | Signing, deployer, or production admin, including any repository or pipeline path that reaches production without a human approval gate | P1 | Multisig signer key, contract upgrade or owner key, deployer EOA, hot wallet, cloud root, domain registrar, exchange withdrawal credentials, CI/CD secrets that deploy unreviewed | Move funds or land code in production, meeting the policy's fund-loss condition on the spot | | Everything else | P2 | Repository write behind review, cloud IAM write, internal admin panels, identity provider admin, official announcement channels, email, chat, shared documents | Impersonate a trusted colleague against the row above, publish to users, or pivot toward a P1 holder | -A multisig signer below threshold is still P1: one more compromised signer executes. A pipeline with no -approval gate is P1 because it deploys on the attacker's behalf, as in -[Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise). +A multisig signer below threshold is still P1, as is any pipeline that deploys without review (see +[Build Pipeline Compromise](/incident-management/incident-response-template/runbooks/build-pipeline-compromise)). **If you cannot tell which row applies, use P1.** Severity is set before anyone has finished the access inventory in [Step 3](#step-3-scope-what-the-device-could-reach), and in an organization of any size you will @@ -87,14 +85,12 @@ wait for a scheduled response. - [ ] Third-party notification (exchange, partner, another protocol's security team) -Extension and plugin installs deserve their own check: a malicious one reads every session the browser -holds, and a repository can pressure the install through its own recommendations. See +A malicious extension reads every session the browser holds, and a repository can pressure the install +through its own recommendations. See [Integrated Development Environments](/devsecops/integrated-development-environments). ### Confirming compromise -One symptom is enough to start. This section is about whether you can stop. - **With EDR**, work the console: process ancestry for anything spawned from a browser, archive, or downloads folder; credential-store access; and the agent's own health. A single quarantined detection still runs this runbook, since quarantine proves the product caught one payload. @@ -113,12 +109,12 @@ persistence lands and the two the user controls, from the device's own interface or from the disk image afterwards. -- [ ] Login items and launch agents or daemons +- [ ] Login items, launch agents, and daemons - [ ] Scheduled tasks and cron entries -- [ ] Shell profile files and PATH overrides -- [ ] Browser and editor extensions, against what the user remembers installing -- [ ] Identity provider sign-in history for logins the user does not recognize -- [ ] What the user actually ran, in their words, with the file or link if they still have it +- [ ] Shell profiles and PATH overrides +- [ ] Browser and editor extensions, against what the user remembers +- [ ] Identity provider sign-in history for logins they do not recognize +- [ ] What they ran, in their words, with the file or link Absent evidence is not absence. If the user executed something and you cannot rule out what it did, treat @@ -159,18 +155,17 @@ attacker's hands. Two tracks, two people, the same minute. The org-side track needs nothing from the user. -- [ ] Identity provider sign-in and session logs exported, before anything is revoked -- [ ] Sessions killed at identity provider, chat, email, and code host -- [ ] Cloud and production access suspended -- [ ] [Decision Makers](/incident-management/incident-response-template/contacts#decision-makers) paged if - the access is or may be P1 +- [ ] Export identity provider sign-in and session logs +- [ ] Kill sessions at identity provider, chat, email, and code host +- [ ] Suspend cloud and production access +- [ ] Page [Decision Makers](/incident-management/incident-response-template/contacts#decision-makers) at P1 - [ ] Call on a channel not tied to the suspect device -- [ ] Confirm identity by voice, never by text alone -- [ ] Tell them to stop using the device and leave it powered as it is -- [ ] Tell them to unplug hardware wallets and security keys now, not later +- [ ] Confirm identity by voice, never by text +- [ ] Tell them to stop using it, powered as it is +- [ ] Tell them to unplug hardware wallets and security keys - [ ] Say plainly that blame is not the point @@ -198,10 +193,9 @@ Find the row that matches your situation: Disconnect means unplug the cable and turn off the wireless radio from a hardware switch where one exists. Toggling Wi-Fi from inside the operating system means working on a live compromised machine. -Powering off ends access to running processes, injected code, and secrets decrypted in memory, which is the -evidence that tells you what actually ran and therefore what must be rotated. It also stops an active thief. -Choose speed when nobody is positioned to capture memory: a machine left running on a desk "for forensics" -that nobody ever images is the worst of both options. +Memory holds what actually ran, and therefore what must be rotated. But a machine left running on a desk +"for forensics" that nobody ever images is the worst of both options, so choose speed when nobody can +capture it. Expect the first row often, since [Malware Infection](/incident-management/playbooks/malware) tells the affected person to power off immediately. Booting it runs the persistence again, so that case proceeds from @@ -218,13 +212,13 @@ mid-incident is how memory gets lost. | 3 | Disk image | The device, before rebuild | [tool and command] | Most volatile first, per [RFC 3227](https://www.rfc-editor.org/rfc/rfc3227). Memory leads because a full -image also captures connections, login sessions, and the process table. +image also captures connections and sessions. -- [ ] Network access cut, by console containment or physical disconnect -- [ ] Device not wiped, not "cleaned", and not casually restarted -- [ ] Hardware wallets, security keys, and external drives unplugged and set aside -- [ ] Time and method of containment recorded +- [ ] Cut network access, at the console or physically +- [ ] Do not wipe, "clean", or restart the device +- [ ] Unplug hardware wallets, security keys, and external drives +- [ ] Record the time and method ### Step 3: Scope what the device could reach @@ -233,11 +227,11 @@ image also captures connections, login sessions, and the process table. reliably under pressure. -- [ ] Accounts held (identity provider, chat, email, code, cloud, treasury) -- [ ] Keys held, and signer roles the user can approve -- [ ] Paired devices (2FA apps, security keys, hardware wallets) -- [ ] Plaintext credentials on disk (.env files, cloud and SSH config, notes apps) -- [ ] Severity set from the table above, P1 wherever the access is unclear +- [ ] List accounts (identity provider, chat, email, code, cloud, treasury) +- [ ] List keys held and signer roles they can approve +- [ ] List paired devices (2FA apps, security keys, hardware wallets) +- [ ] Find plaintext credentials (.env, cloud and SSH config, notes apps) +- [ ] Set severity, P1 if the access is unclear **If you have endpoint management:** pull the installed application list, the extension inventory, and @@ -255,8 +249,7 @@ leaves the stolen cookie valid on any identity provider that does not revoke ses which is most of them. Re-enrolling 2FA from a cloud backup restores the same seed, which the attacker may already hold. -Rows 1 and 2 run together at P1. Otherwise work top to bottom. Assign every `[Name]` before an incident: -an unassigned row is where revocation stalls at 3am. +Rows 1 and 2 run together at P1. Otherwise work top to bottom. Assign every `[Name]` in advance. | # | Credential class | Where | Owner | Notes | | --- | ------------------ | ------- | ------- | ------- | @@ -331,12 +324,12 @@ Raise it early, while the timing is still yours to pick. For wording, cadence, a [Communication Strategies](/incident-management/communication-strategies). -- [ ] Disclosure obligations checked against your jurisdictions and +- [ ] Check disclosure obligations against your jurisdictions and [reporting window, 72 hours suggested] -- [ ] Affected users or customers notified if their data or funds were reachable -- [ ] Law enforcement report filed if funds were stolen (see - [Malware Infection](/incident-management/playbooks/malware) for the reporting venues) -- [ ] Partners and exchanges told which addresses and accounts to watch +- [ ] Notify affected users if their data or funds were reachable +- [ ] File a law enforcement report if funds were stolen (see + [Malware Infection](/incident-management/playbooks/malware) for venues) +- [ ] Tell partners and exchanges which addresses to watch ## Recovery and prevention @@ -348,10 +341,10 @@ compromise window opened, and do not accept "the antivirus removed it" as an out infostealer is gone, and the cost of being wrong is handing back signing authority. -- [ ] Device wiped and reinstalled from trusted media, or replaced -- [ ] Data moved file by file, never a whole user folder -- [ ] Browser profiles, extensions, and shell config rebuilt, not migrated -- [ ] Device retained if forensic or regulatory obligations require it +- [ ] Wipe and reinstall from trusted media, or replace the device +- [ ] Move data file by file, never a whole user folder +- [ ] Rebuild browser profiles, extensions, and shell config +- [ ] Retain the device if forensics or regulation require it ### Return-to-service gate @@ -370,33 +363,31 @@ Privileged access returns only when a named approver confirms all of the followi ### Prevention -- [ ] Device tier matched to role risk, per [Endpoint Security](/opsec/endpoint/overview) -- [ ] Signing devices separate from daily-driver workstations -- [ ] EDR and disk encryption on the devices of P1 access holders -- [ ] Extensions allowlisted, repo-recommended ones declined by default -- [ ] Untrusted code run only in a sandbox or disposable VM -- [ ] Phishing-resistant MFA on high-value accounts -- [ ] Shortened session lifetimes for privileged tooling -- [ ] A known-clean device identified and kept ready -- [ ] Contributors know where to report, including out of hours +- [ ] Device tiers, per [Endpoint Security](/opsec/endpoint/overview) +- [ ] Separate signing devices +- [ ] EDR and disk encryption for P1 access holders +- [ ] Extension allowlisting +- [ ] Sandboxing for untrusted code +- [ ] Phishing-resistant MFA +- [ ] Short session lifetimes for privileged tooling +- [ ] A known-clean device kept ready +- [ ] An out-of-hours reporting path ### Closing the incident -The rebuild is the midpoint. [NIST SP 800-61r3](https://csrc.nist.gov/pubs/sp/800/61/r3/final) treats -improvement as a standing function rather than a meeting held once the noise dies down, so the work below -closes the runbook. [Lessons Learned](/incident-management/lessons-learned) carries the review questions and -the blameless format. +The rebuild is the midpoint. [Lessons Learned](/incident-management/lessons-learned) carries +the review questions and the blameless format. -- [ ] Timeline written up in the +- [ ] Write up the timeline in the [Incident Log](/incident-management/incident-response-template/templates/incident-log-template) -- [ ] Return-to-service gate signed off by the named approver -- [ ] Post-mortem scheduled per the +- [ ] Get the return-to-service gate signed off +- [ ] Schedule a post-mortem per the [Post-Mortem Template](/incident-management/incident-response-template/templates/post-mortem-template) -- [ ] Indicators shared with the team, and with partners if the lure was reused against them -- [ ] Prevention items above converted into owned work, not left in this checklist -- [ ] Detection gap reviewed: what would have caught this a day earlier +- [ ] Share indicators with the team, and with partners if the lure was reused +- [ ] Assign owners to the prevention items above +- [ ] Review the detection gap: what would have caught this a day earlier ## Further reading From 0fc81c71db2fb6ce1de4f40df07f4c56206f5dbd Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Wed, 7 Oct 2026 21:26:09 +0900 Subject: [PATCH 10/11] docs(incident-management): re-check sessions after password rotation Review raised that a service with no second factor lets a stolen password open a new session right after the old one is killed. The revocation table only covered OAuth tokens, which session revocation also misses. Row 1 now says to re-check after row 8. --- .../incident-response-template/runbooks/endpoint-compromise.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index bcd6ffebc..4246629d5 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -253,7 +253,7 @@ Rows 1 and 2 run together at P1. Otherwise work top to bottom. Assign every `[Na | # | Credential class | Where | Owner | Notes | | --- | ------------------ | ------- | ------- | ------- | -| 1 | Active sessions | Identity provider, chat, email, code host | [Name] | Started in Step 1; confirm every system is covered. Session revocation does not reach OAuth tokens, so revoke those separately | +| 1 | Active sessions | Identity provider, chat, email, code host | [Name] | Started in Step 1; confirm every system is covered. Revoke OAuth tokens separately, since session revocation does not reach them. Re-check after row 8: on a service with no second factor, a stolen password buys a fresh session | | 2 | Signing and deployer keys | Multisig, contracts, treasury | [Name] | Run alongside row 1 at P1, per [Key Compromise](/incident-management/incident-response-template/runbooks/key-compromise). For a hot wallet, revoking the session is not enough: move the assets to a wallet generated on a known-clean device | | 3 | Password manager | Vault account | [Name] | Assume vault contents exposed; rotate master password and re-enroll 2FA from a known-clean device | | 4 | MFA enrollments | Identity provider, exchanges | [Name] | Re-enroll on a new device; never restore an authenticator from cloud backup | From b9664f43ffa4d9666cddcf878ea43b29d2fab10c Mon Sep 17 00:00:00 2001 From: s1ns3nz0 Date: Wed, 7 Oct 2026 21:32:03 +0900 Subject: [PATCH 11/11] docs(incident-management): say what removes a host foothold Review asked what happens when the attacker has a backdoor on the host. The foothold gating sentence only covered account-side persistence, so nothing stated that network containment and the rebuild are what remove a host foothold, and that rotating credentials does not. --- .../runbooks/endpoint-compromise.mdx | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx index 4246629d5..9a7fb1693 100644 --- a/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx +++ b/docs/pages/incident-management/incident-response-template/runbooks/endpoint-compromise.mdx @@ -241,7 +241,9 @@ recent process history instead of relying on the user's memory. Finish the access inventory in [Step 3](#step-3-scope-what-the-device-could-reach) before removing footholds. Revoking what you have found while an OAuth grant or mail rule you have not found survives moves -the attacker out of the device and leaves them in the accounts. +the attacker out of the device and leaves them in the accounts. A foothold on the host itself is handled by +Step 2 cutting the network and by rebuilding rather than cleaning; do not treat credential rotation as +having removed it. **Why order matters:** the usual failures are sequencing errors. Rotating an SSH key before killing the attacker's live session lets them add the new key themselves. Changing a password before revoking sessions