Skip to main content
AD Academy

War stories: real incidents

What a real investigation looks like, from first complaint to root cause. Details changed, logic real.

Replication died on Friday evening

HIGH

The complaint

Branch users report a new password "does not work" — but only at the branch. Head office is fine.

Investigation logic

  1. 1repadmin /replsummary — one branch DC failing for 14 days.
  2. 2Directory Service log: event 1988 — lingering object.
  3. 3Turns out that DC was powered off for two months during an office refit.
  4. 4Monitoring check: nobody was alerting on replication at all.

Root cause

A DC offline past the tombstone lifetime still held deleted objects, so AD blocked replication with it.

What we did

We demoted the DC, cleaned metadata (ntdsutil metadata cleanup) and rebuilt it clean. Password changes flowed within the hour.

Takeaway

Monitoring replication matters as much as monitoring uptime. A DC can answer ping and still be dead to AD.

Related events:19882042

The same manager gets locked out every morning

MEDIUM

The complaint

One manager is locked out every morning at 08:10. Unlocking helps for two hours.

Investigation logic

  1. 1On the PDC Emulator, 4740 pointed at a terminal server via Caller Computer Name.
  2. 2On that server: 4625 events with Logon Type 3 every few minutes.
  3. 3Session check: a disconnected RDP session three weeks old under that user.

Root cause

The stale session cached the old password and kept resending it after the user changed it.

What we did

We killed the session and set a policy to end disconnected sessions after 8 hours. The lockouts stopped.

Takeaway

A repeating lockout is a symptom. Do not unlock in a loop — find the Caller Computer Name.

Related events:474046254771

"No internet" that turned out to be a domain outage

MEDIUM

The complaint

After a router swap the branch has internet, but domain logon is odd and GPOs are stale.

Investigation logic

  1. 1ipconfig /all on a client: DNS = 192.168.1.1 (the router).
  2. 2nslookup -type=SRV _ldap._tcp.dc._msdcs.corp.local — no answer.
  3. 3System log: event 1058 — SYSVOL unreadable.

Root cause

The new router handed out its own DHCP with its own DNS, so clients stopped finding the DC.

What we did

We disabled DHCP on the router, restored the Windows scope with DNS = DC, then ran ipconfig /renew.

Takeaway

In a domain, client DNS always points at a DC. "Internet works" says nothing about the domain.

Related events:10581085

A nine-year-old service account brought down the domain

HIGH

The complaint

The SOC alerted: a normal user requested 40 service tickets in 3 minutes, all RC4.

Investigation logic

  1. 14769 events with Encryption Type 0x17 from one account — the classic Kerberoasting signature.
  2. 2We listed SPN accounts: svc_backup with PasswordLastSet in 2017.
  3. 3Event 4672 showed svc_backup logging into a file server with high privileges an hour later.

Root cause

A short, ancient password on an SPN service account — cracked offline in minutes.

What we did

We rotated to a 30-character password, moved the service to a gMSA, and phased out RC4.

Takeaway

Every SPN service account is an offline-crackable password. gMSA solves it permanently.

Related events:476946724738

Someone restored a DC from a snapshot

HIGH

The complaint

After a "fix" in virtualization some users cannot log on and several servers lost domain trust.

Investigation logic

  1. 15722/5723 events on the DC — machine account passwords out of sync.
  2. 2repadmin showed a USN rollback on one DC.
  3. 3The virtualization team confirmed a two-week-old snapshot restore.

Root cause

A snapshot restore rewinds the AD database and causes a USN rollback: new changes never replicate again.

What we did

We demoted the damaged DC, cleaned metadata and rebuilt it. Clients were fixed with Test-ComputerSecureChannel -Repair.

Takeaway

Never restore a DC from a snapshot. Use Windows Server Backup / VM-Generation ID, or just build a new DC.

Related events:572257232042

The security log that vanished on Saturday night

HIGH

The complaint

The SIEM went quiet for one server; next morning its log held a single event: 1102.

Investigation logic

  1. 1We pulled logs from the forwarder — history survived there.
  2. 2Half an hour earlier: 4624 Type 3 from an unknown host, then 7045 — a new randomly named service.
  3. 3Then 4728 — an addition to Domain Admins.

Root cause

An attacker entered via a weak-password host, moved to the server, gained admin and cleared the log.

What we did

Isolation, double krbtgt reset, privileged password rotation, and enforced real-time log forwarding.

Takeaway

A log kept only on the host is not evidence. Forwarding logs off-box is the cheapest defense there is.

Related events:1102704547284624