Parsing logs before you can model them
Raw log messages are free text with numbers, IPs and usernames mixed in — before any anomaly detection can happen, similar messages have to be grouped into templates so rare, new patterns stand out.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Before a model can spot an unusual log message, similar messages must first be grouped, so "unusual" has meaning.
Think about sorting a shoebox full of receipts. Some printed, some handwritten, some faded, from different shops. You cannot add up your spending by reading them one at a time, in whatever order they come out. First you sort them into piles: groceries, fuel, food delivery. Only once they are grouped does "I spent an unusual amount on fuel" mean anything at all.
Raw security logs need this exact same sorting step, before anything useful can be built on top of them.
Why it exists
A log line like "User alice logged in from 10.0.0.5" and "User bob logged in from 10.0.0.9" describe the same kind of event: a login. As raw text, they are completely different strings. A username and an IP change every time, even though the event type is identical.
Suppose you tried to spot "unusual" messages by comparing raw text directly. Every single login would look different from every other. Nothing would ever repeat. Nothing would ever look "usual," against which something else could look "unusual." The fix is turning each message into a template. This is the fixed, repeating shape of the message, with the changing parts replaced by placeholders.
Once messages are reduced to templates, counting becomes meaningful. A login template seen four times today is normal. A template appearing for the first time is worth a second look. Not because of what it says. Because of how rare its shape is.
How it works
"User alice logged in from 10.0.0.5" -> "User <USER> logged in from <IP>"
"User bob logged in from 10.0.0.9" -> "User <USER> logged in from <IP>"
"User dave logged in from 10.0.0.7" -> "User <USER> logged in from <IP>"
"Disk usage on /var at 92 percent" -> "Disk usage on /var at <NUM> percent"
|
v
Login template seen 4 times today: normal
Disk-usage template seen for the first time: worth a lookThe rarity of the shape, not the specific words, is the signal.
A real example you have seen
Server dashboards that show "unusual error pattern detected" are almost always built on this idea. Millions of log lines get grouped into a much smaller number of templates. Then the system watches for a template appearing far more, or far less, than its usual pattern.
Remember this
- Raw log text has to be grouped into templates, with variable parts replaced, before "unusual" can mean anything.
- A rare or brand-new template is a stronger signal than any single log line's exact wording.
- This grouping step is unglamorous. It is tedious to build well, and usually more work than the anomaly detection on top of it.
- A flagged anomaly is a lead for a human analyst to check, not an automatic verdict on its own.
What to learn next
- Data poisoning and backdoors — what happens if an attacker can influence the training data this whole pipeline is built on.
- Isolation forest — a standard tool for the anomaly-detection step once counts like these exist.
- Regex for text — the tool underneath the template-extraction code in this lesson.
Developer — Code and libraries.
Setup
python3 --versionNo installs needed — standard library only, continuing directly from security telemetry.
Minimal runnable code
We collapse seven raw log lines into templates by replacing the variable parts, then flag any template seen only once as worth a closer look.
import re
from collections import Counter
logs = [
"User alice logged in from 10.0.0.5",
"User bob logged in from 10.0.0.9",
"User carol logged in from 10.0.0.2",
"Failed to connect to database after 3 retries",
"Failed to connect to database after 5 retries",
"User dave logged in from 10.0.0.7",
"Disk usage on /var at 92 percent",
]
def to_template(message):
message = re.sub(r"User \w+", "User <USER>", message)
message = re.sub(r"\b\d+\.\d+\.\d+\.\d+\b", "<IP>", message)
message = re.sub(r"\b\d+\b", "<NUM>", message)
return message
templates = [to_template(m) for m in logs]
counts = Counter(templates)
for msg, tmpl in zip(logs, templates):
freq = counts[tmpl]
tag = " <- RARE, worth investigating" if freq == 1 else ""
print(msg)
print(f" template: {tmpl} (seen {freq}x so far in this batch){tag}")
print()
print(f"{len(logs)} raw log lines collapsed into {len(counts)} distinct templates")User alice logged in from 10.0.0.5 template: User <USER> logged in from <IP> (seen 4x so far in this batch) User bob logged in from 10.0.0.9 template: User <USER> logged in from <IP> (seen 4x so far in this batch) User carol logged in from 10.0.0.2 template: User <USER> logged in from <IP> (seen 4x so far in this batch) Failed to connect to database after 3 retries template: Failed to connect to database after <NUM> retries (seen 2x so far in this batch) Failed to connect to database after 5 retries template: Failed to connect to database after <NUM> retries (seen 2x so far in this batch) User dave logged in from 10.0.0.7 template: User <USER> logged in from <IP> (seen 4x so far in this batch) Disk usage on /var at 92 percent template: Disk usage on /var at <NUM> percent (seen 1x so far in this batch) <- RARE, worth investigating 7 raw log lines collapsed into 3 distinct templates
What actually happened
Seven completely different-looking log lines collapsed into only three templates. Four different logins — different users, different IPs — all correctly mapped to the exact same template, because the shape of a login message never changes, only the specific values inside it.
The disk-usage message is the interesting one. It appeared only once, and got correctly flagged, purely because its shape has not been seen before in this batch — not because "92 percent" is an inherently scary number. A model built on top of these template counts would learn to watch for new or rarely-seen templates, which is a far more general signal than trying to hardcode "alert if disk usage over 90%" by hand.
re.sub(r"User \w+", "User <USER>", message)replaces the pattern before the IP and number substitutions run — order matters here, since a username could theoretically contain digits that the later\d+pattern would otherwise touch.Counter(templates)is doing the real counting work; the loop afterward only exists to print results readably.- This is a genuinely simplified version of real log parsing algorithms. Production systems (Drain, Spell, and similar) handle far messier, more varied log formats automatically, without hand-written regex per log type.
Common mistakes
Writing one regex per log source by hand, forever. This does not scale past a handful of log formats. Real log parsers use general clustering algorithms (like Drain) that learn templates automatically from examples, rather than requiring a new regex every time a new application starts logging.
Treating "rare template" as automatically dangerous. A brand-new but completely legitimate application update can introduce a genuinely new, harmless log template. Rarity is a reason to look, not a verdict.
Ignoring multi-line log messages. Stack traces and multi-line error dumps break a simple "one line, one event" assumption immediately. Real log parsers need explicit logic to detect where one logical event starts and ends.
Forgetting that "seen so far" depends on the window. This demo's counts reset with each run, over a tiny fixed batch. A real system needs to decide: rare compared to the last hour? The last month? Same time last week (accounting for daily patterns)? That choice changes what counts as rare substantially.
Try it yourself
Add a log line describing a specific rare event, "Firewall rule DENY-9981 triggered from 203.0.113.4", and extend to_template with one more re.sub call to normalise the rule ID (something like \bDENY-\d+\b). Confirm it now collapses correctly if you add a second, similar line with a different rule number.
What to learn next
- Regex for text — the general technique this lesson's parsing is built on.
- Isolation forest — turning template counts into a proper anomaly score.
- Data poisoning and backdoors — the next lesson, on attacks against the training data itself.
Researcher — Mathematics and papers.
Log parsing as unsupervised template mining
Formally, log parsing seeks a function that maps each raw log message m to a template t such that messages sharing an underlying event type map to the same t, while the message-specific variable tokens (usernames, IPs, session IDs, byte counts) are abstracted away. This is unsupervised: no labelled dataset of "correct" templates typically exists for a new log source, so parsers must infer template boundaries from message structure and frequency alone.
Drain (He et al., 2017) is the most widely used and benchmarked algorithm for this. It builds a fixed-depth parse tree keyed first by message length, then by the first few tokens, then applies a similarity threshold to decide whether a new message matches an existing template's token pattern (using a <*> wildcard for positions where token content varies across matched messages) or should start a new template. Its key practical property is that it runs online, incrementally, and in roughly constant time per message, which is what makes it usable on live production log streams rather than only as an offline batch process.
Log-based anomaly detection, once parsed
With messages reduced to a sequence of template IDs, standard anomaly-detection framings become tractable:
- Frequency-based — flag templates whose occurrence count deviates significantly from a learned baseline (as in this lesson's simplified "seen once" heuristic), typically formalised via a control-chart or Poisson-rate model per template.
- Sequence-based — model the order of template IDs, since some anomalies show up only as an unusual sequence of otherwise-normal individual events (DeepLog, Du et al., 2017, uses an LSTM to predict the next expected template ID given recent history, flagging sequences the model finds surprising).
- Session/workflow-based — group template sequences into logical sessions (e.g. all logs for one user request or one system session) and detect anomalous sessions as a whole, closer to how a human investigator actually reasons about an incident.
The LogPAI benchmark ecosystem
He et al. (2020, LogPAI) established a standard benchmark suite of sixteen real-world log datasets (HDFS, Hadoop, Spark, and others) with human-annotated ground-truth templates, used to fairly compare parsing algorithms on accuracy (does the inferred template match the human-annotated one) and efficiency (throughput on large log volumes). This benchmark revealed that reported accuracy across popular parsers varies substantially by log source — a parser tuned well on one system's logs does not automatically generalise to another's, an important caveat when adopting any published log-parsing accuracy figure.
Cost
Template-count throughput matters directly at production scale: Drain and similar streaming parsers are designed to process on the order of thousands of log lines per second per core, because a parser that cannot keep pace with live log ingestion becomes the bottleneck for every downstream detection system built on top of it, regardless of how good those downstream models are.
Key references
- He, P., Zhu, J., Zheng, Z., & Lyu, M. R. (2017). Drain: An Online Log Parsing Approach with Fixed Depth Tree. IEEE ICWS.
- Du, M., Li, F., Zheng, G., & Srikumar, V. (2017). DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning. ACM CCS.
- He, S., et al. (2020). Loghub: A Large Collection of System Log Datasets for AI-Driven Log Analytics. IEEE ISSRE — the LogPAI benchmark suite.
- Zhu, J., et al. (2019). Tools and Benchmarks for Automated Log Parsing. IEEE/ACM ICSE-SEIP — a systematic comparison of parsing algorithm accuracy across datasets.
Current state
LLM-based log parsing (prompting a language model to extract templates directly, or to explain an anomalous log sequence in natural language for an analyst) is an active area of both research and commercial tooling as of the last few years, but throughput and cost per log line remain substantially worse than Drain-family algorithms, which is why most high-volume production pipelines still use classical parsing for the bulk ingestion path and reserve LLM-based analysis for the much smaller volume of already-flagged, human-reviewed events.
What to learn next
- Isolation forest — a general anomaly-detection technique applicable to parsed log features.
- Time-series features — building richer rolling-window features from template counts over time.
- Data poisoning and backdoors — the next lesson in this section.