War stories: real incidents
What a real investigation looks like, from first complaint to root cause. Details changed, logic real.
Replication died on Friday evening
HIGHThe complaint
Branch users report a new password "does not work" — but only at the branch. Head office is fine.
Investigation logic
- 1repadmin /replsummary — one branch DC failing for 14 days.
- 2Directory Service log: event 1988 — lingering object.
- 3Turns out that DC was powered off for two months during an office refit.
- 4Monitoring check: nobody was alerting on replication at all.
Root cause
A DC offline past the tombstone lifetime still held deleted objects, so AD blocked replication with it.
What we did
We demoted the DC, cleaned metadata (ntdsutil metadata cleanup) and rebuilt it clean. Password changes flowed within the hour.
Takeaway
Monitoring replication matters as much as monitoring uptime. A DC can answer ping and still be dead to AD.
The same manager gets locked out every morning
MEDIUMThe complaint
One manager is locked out every morning at 08:10. Unlocking helps for two hours.
Investigation logic
- 1On the PDC Emulator, 4740 pointed at a terminal server via Caller Computer Name.
- 2On that server: 4625 events with Logon Type 3 every few minutes.
- 3Session check: a disconnected RDP session three weeks old under that user.
Root cause
The stale session cached the old password and kept resending it after the user changed it.
What we did
We killed the session and set a policy to end disconnected sessions after 8 hours. The lockouts stopped.
Takeaway
A repeating lockout is a symptom. Do not unlock in a loop — find the Caller Computer Name.
"No internet" that turned out to be a domain outage
MEDIUMThe complaint
After a router swap the branch has internet, but domain logon is odd and GPOs are stale.
Investigation logic
- 1ipconfig /all on a client: DNS = 192.168.1.1 (the router).
- 2nslookup -type=SRV _ldap._tcp.dc._msdcs.corp.local — no answer.
- 3System log: event 1058 — SYSVOL unreadable.
Root cause
The new router handed out its own DHCP with its own DNS, so clients stopped finding the DC.
What we did
We disabled DHCP on the router, restored the Windows scope with DNS = DC, then ran ipconfig /renew.
Takeaway
In a domain, client DNS always points at a DC. "Internet works" says nothing about the domain.
A nine-year-old service account brought down the domain
HIGHThe complaint
The SOC alerted: a normal user requested 40 service tickets in 3 minutes, all RC4.
Investigation logic
- 14769 events with Encryption Type 0x17 from one account — the classic Kerberoasting signature.
- 2We listed SPN accounts: svc_backup with PasswordLastSet in 2017.
- 3Event 4672 showed svc_backup logging into a file server with high privileges an hour later.
Root cause
A short, ancient password on an SPN service account — cracked offline in minutes.
What we did
We rotated to a 30-character password, moved the service to a gMSA, and phased out RC4.
Takeaway
Every SPN service account is an offline-crackable password. gMSA solves it permanently.
Someone restored a DC from a snapshot
HIGHThe complaint
After a "fix" in virtualization some users cannot log on and several servers lost domain trust.
Investigation logic
- 15722/5723 events on the DC — machine account passwords out of sync.
- 2repadmin showed a USN rollback on one DC.
- 3The virtualization team confirmed a two-week-old snapshot restore.
Root cause
A snapshot restore rewinds the AD database and causes a USN rollback: new changes never replicate again.
What we did
We demoted the damaged DC, cleaned metadata and rebuilt it. Clients were fixed with Test-ComputerSecureChannel -Repair.
Takeaway
Never restore a DC from a snapshot. Use Windows Server Backup / VM-Generation ID, or just build a new DC.
The security log that vanished on Saturday night
HIGHThe complaint
The SIEM went quiet for one server; next morning its log held a single event: 1102.
Investigation logic
- 1We pulled logs from the forwarder — history survived there.
- 2Half an hour earlier: 4624 Type 3 from an unknown host, then 7045 — a new randomly named service.
- 3Then 4728 — an addition to Domain Admins.
Root cause
An attacker entered via a weak-password host, moved to the server, gained admin and cleared the log.
What we did
Isolation, double krbtgt reset, privileged password rotation, and enforced real-time log forwarding.
Takeaway
A log kept only on the host is not evidence. Forwarding logs off-box is the cheapest defense there is.