1 What happens inside
src_ip, user, _timeEvery step can go wrong in a way that matters for evidence. A parser that reads a local time as UTC moves every
event by an hour. A field called user can mean the account in one source and the display name in
another. An event that failed to parse may be missing from searches entirely. Before you rely on a SIEM search,
look at the raw lines behind a few results, and keep the original exports.
2 Search, pivot, timeline
- Start from the trigger. Read the alert's raw events. Write down every entity it names: addresses, hosts, accounts, domains, times.
- Pivot on each entity. Search each one across all sources, with a time window around the event. Every new entity you find (a MAC address, a second account, a domain) becomes the next pivot.
- Widen the time. What did the same entities do in the days before and after? Preparation and follow-up are often more telling than the incident itself.
- Build the timeline. One list, one time base, every entry with its source and how reliable it is.
- Test hypotheses. For each explanation ask which events it predicts, then search for them. Look as hard for evidence against your favourite explanation as for evidence for it.
- Document the searches you ran, with time range and result counts, so another examiner can repeat them.
3 A small SIEM to practise on
One invented day at an invented company: Windows logon events, a web proxy, a VPN gateway and a GeoIP table. The
language is the same as in the case workbench: field=value, then commands after |, for
example | stats count by user, | table _time user src_ip or
| lookup geoip network AS src_ip. Times are UTC.
4 Alerts are hypotheses
An alert says that a rule matched, not that something bad happened. Triage means deciding quickly whether it is a true positive (the rule caught what it was meant to catch), a false positive (it matched something harmless), or needs investigation. Two habits help:
- Know the base rate. A rule that fires for every large upload will mostly catch backups and video calls. Check what the rule sees on normal days before you trust a single hit.
- Check the inputs. Many rules depend on enrichment data. "Impossible travel" compares GeoIP locations of two logins. But GeoIP databases record who registered an address range, not where the user stands. VPN exits, mobile carriers (CGNAT), satellite and corporate networks often geolocate hundreds of kilometres away; the accuracy radius and the organisation name tell you how far to trust a location.
5 Try it yourself
Answer with searches in the small SIEM above.
How many failed logons (event 4625) are there for the account bob?
sourcetype=WinEventLog:Security EventID=4625 user=bob, or | stats count by user.
6, between 03:12:05 and 03:13:30 UTC, all from the same address, all with SubStatus
0xC000006A (wrong password), followed by a successful logon at 03:15:02. A burst at 3 a.m. from outside is
password guessing that eventually worked, or the owner trying passwords; the next events decide.
From which address did the failed logons come?
Add | stats count by src_ip, then look the address up with | lookup geoip network AS src_ip.
203.0.113.50, a cheap hosting range according to the GeoIP table, not a home or mobile network. The same address starts the VPN session at 03:15:03.
How many MiB did bob upload through the proxy during that VPN session?
sourcetype=squid:access user=bob | eval mib = bytes_from_client / 1048576 | table _time url mib
671,088,640 bytes = 640 MiB to files.example.net, logged at 03:27:00, after a tunnel
of 410 s that started at 03:20:10. Failed logons, a login from a hosting address, then a large upload: a classic
account-takeover pattern, to be confirmed by asking bob.
carol logs in to the VPN from Leeds at 08:00 and from Lisbon at 09:10 UTC. What should you check first?
Run the "VPN with GeoIP" example.
Check the GeoIP records. The Lisbon range belongs to a satellite internet provider, whose addresses map to its ground station, not to the user. Carol may well have been on a train or a plane with satellite Wi-Fi. Ask her before you call it an attack.
Alert rules often compare volumes. These are the packet and byte counters of a flow record. Edit them and watch the decoded record:
- Set the byte counter to
00 01 00 00(64 KiB). Would a "more than 50 MB" rule still fire? - If the router sampled 1 in 100 packets, what real volume would the original counter stand for?