Malware classification
A malware classifier looks for measurable telltale signs in a file, the same way a bank teller spots a fake note by feel, without reading every byte the way a human would read a document.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Malware classification decides whether a file is malicious by measuring specific, checkable properties. Not by "reading" it like a person would.
Think about a bank teller checking a currency note. They do not read every printed word on the note. They check specific things: the watermark, the security thread, how the paper feels, whether the number matches under UV light. A handful of checkable signals is enough to decide fast. Signals learned from handling thousands of real and fake notes.
A malware classifier checks a file the same way. Measurable properties, not a full read.
Why it exists
Antivirus software used to work almost entirely by signatures — an exact fingerprint of a known bad file. This works perfectly for malware seen before. It fails completely against anything new. A signature only matches an exact known file.
Malware authors know this. They change their file slightly with every release, specifically to dodge signature matching. This technique is called polymorphism. A machine learning classifier looks at properties of a file instead of an exact fingerprint. How random do the file's bytes look? Does it ask the operating system to do suspicious things? Is it digitally signed by a known publisher? A slightly modified file usually still shows the same suspicious properties, even with a completely different signature.
How it works
A file arrives
|
v
[ measure properties, WITHOUT running it ]
- how random do the bytes look? (packed/encrypted files look very random)
- is it digitally signed by a known publisher?
- what system functions does it ask to use?
- how big is it, compared to what it claims to be?
|
v
[ classifier ]
|
v
"malicious, 94% confident" or "benign, 98% confident"This is called static analysis: examining the file without executing it. It is fast and safe. The opposite approach, dynamic analysis, actually runs the file in an isolated, contained sandbox, and watches what it does. It is slower, but harder for malware to fool. The next lesson covers this further.
A real example you have seen
Modern antivirus tools (Windows Defender, and most enterprise security products) combine exactly this. A fast, static ML check runs on every new file. Known threats get caught by signatures. Anything still uncertain gets sandboxed.
Remember this
- A classifier judges a file by measurable properties, not by reading it like text. Randomness of its bytes, whether it is signed, what it asks the system to do.
- This approach catches new, slightly modified malware that exact signature matching misses.
- Static analysis (checking properties without running the file) is fast and safe. Dynamic analysis (running it in a sandbox) is slower but harder to fool.
- A "malicious" verdict should trigger quarantine and review, not silent auto-deletion, until a security analyst confirms it.
What to learn next
- Your opponent updates too — what happens once malware authors learn what your classifier looks for.
- Random forest — the model family used in this lesson's code.
- Feature importance done right — reading which properties a model actually relied on.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyMinimal runnable code
A small synthetic dataset built from well-documented, published malware heuristics — high byte entropy from packing, unsigned binaries, and suspicious system calls. No real malware files are used or needed.
import numpy as np
from sklearn.ensemble import RandomForestClassifier
# Each row: [byte_entropy(0-8), suspicious_api_calls, is_digitally_signed(0/1), file_size_kb]
files = np.array([
[3.1, 0, 1, 850], # ordinary signed installer
[2.8, 1, 1, 420],
[7.9, 12, 0, 38], # packed, unsigned, tiny -- classic packed-malware shape
[3.5, 0, 1, 1200],
[7.6, 9, 0, 55],
[2.9, 1, 1, 300],
[7.8, 15, 0, 22],
[3.2, 0, 1, 900],
[7.4, 10, 0, 60],
[3.0, 2, 1, 500],
])
is_malicious = np.array([0,0,1,0,1,0,1,0,1,0])
model = RandomForestClassifier(n_estimators=100, random_state=0)
model.fit(files, is_malicious)
test_files = np.array([
[3.3, 1, 1, 700], # looks ordinary
[7.7, 11, 0, 41], # looks packed and suspicious
])
predictions = model.predict(test_files)
probabilities = model.predict_proba(test_files)[:, 1]
for i, (pred, prob) in enumerate(zip(predictions, probabilities)):
label = "MALICIOUS" if pred == 1 else "benign"
print(f"file {i}: predicted={label} malicious_probability={prob:.2f}")
print()
importances = dict(zip(
["entropy", "suspicious_api_calls", "is_signed", "file_size_kb"],
model.feature_importances_.round(2),
))
print("which features the model leaned on:", importances)file 0: predicted=benign malicious_probability=0.00
file 1: predicted=MALICIOUS malicious_probability=0.99
which features the model leaned on: {'entropy': 0.21, 'suspicious_api_calls': 0.26, 'is_signed': 0.28, 'file_size_kb': 0.24}What actually happened
The two test files were built to resemble two clear archetypes from the training data — file 0 looks like the ordinary signed installers, file 1 looks like the packed, unsigned samples. The model separated them cleanly, and feature_importances_ shows all four properties contributed meaningfully, with none dominating completely — a sign the model is not relying on a single brittle rule.
Byte entropy deserves a plain explanation: it measures how random a sequence of bytes looks, on a scale where ordinary program code and text sit around 2 to 5, and encrypted or compressed data sits close to the maximum of 8. Malware is very often packed — compressed or encrypted to hide its real contents from static tools — which pushes entropy up sharply. That single number, computed without running the file at all, is one of the most useful cheap signals in this entire field.
RandomForestClassifierwas used, not logistic regression, because the true relationship here is not smoothly linear — entropy above roughly 7 matters a great deal, entropy between 3 and 5 barely matters at all, a "threshold-like" pattern trees capture naturally.feature_importances_for a random forest reflects how much each feature reduced impurity across all trees — a useful first read, though it can be misleading with correlated features, exactly as flagged for logistic regression coefficients in the finance section.
Common mistakes
Trusting static features alone against a well-resourced attacker. Entropy and signing status are cheap to fake or dodge — an attacker can sign a malicious file with a stolen certificate, or use techniques that keep entropy artificially low. Static analysis is a fast first filter, not a complete defence.
Training on too small or unrepresentative a sample. Real malware families vary enormously. Ten synthetic rows exist to show the mechanics; a real classifier trains on labelled datasets with many thousands of diverse samples, both malicious and benign, collected under careful legal and safety controls.
Running suspicious files to generate features. Even opening or running a file "to take a look" outside a proper sandbox can infect the machine doing the checking. Static features are computed by parsing the file's structure, never by executing it.
Ignoring class imbalance. In production, the overwhelming majority of files scanned are benign. This toy example is roughly balanced for clarity — see imbalanced data for how a real deployment has to handle the realistic ratio.
Try it yourself
Add a tricky file: [7.5, 1, 1, 900] — high entropy (looks packed) but signed and only one suspicious API call. This resembles a legitimate, commercially packed and signed application. Check what the model predicts, and consider why entropy alone is not a safe rule to hard-code.
What to learn next
- Random forest — the model type used here, explained from first principles.
- Feature importance done right — a deeper, more careful treatment of reading importances like the ones printed above.
- Your opponent updates too — why static features alone do not stay effective forever.
Researcher — Mathematics and papers.
Feature families in malware classification
Production systems typically draw on several distinct feature families, often combined:
- Static structural features — PE (Portable Executable, the Windows binary format) header fields, section table entropy per-section rather than whole-file, imported/exported function tables, digital signature validity chain.
- Byte n-gram / string features — frequency of byte sequences or extracted printable strings, treated as a bag-of-features exactly analogous to bag-of-words text classification; see text classification for the shared underlying technique.
- Dynamic behavioural features — API call sequences observed during sandboxed execution (process creation, registry writes, network connections attempted), capturing intent that static analysis of an obfuscated file can miss entirely.
- Raw-byte deep learning — treating the executable's raw bytes as a 1D sequence (Raff et al., 2018, MalConv) or reshaping them into a 2D "image" (Nataraj et al., 2011) and applying CNN architectures directly, avoiding hand-engineered features at the cost of interpretability.
Static analysis is fundamentally limited by obfuscation
Packing, encryption, and code virtualization can, in principle, make a file's static byte-level properties resemble anything the attacker chooses, including deliberately mimicking benign software's entropy profile. This is not a hypothetical: publicly documented "packer as a service" tooling exists specifically to defeat static ML and signature detection. This is the direct motivation for combining static classification with dynamic sandbox analysis and endpoint behavioural monitoring (EDR) rather than relying on any single layer.
Sandbox evasion
Malware increasingly checks whether it is running inside a sandbox before revealing malicious behaviour — detecting virtualization artifacts, unrealistic system uptime, absence of normal user activity, or known analysis tool processes — and behaves benignly if so, defeating dynamic analysis in turn. This is a direct, ongoing arms race between defenders and malware authors, covered structurally in the next lesson.
Evaluation and public benchmarks
The EMBER dataset (Anderson & Roth, 2018, Endgame/Elastic) is the most widely used open benchmark for static PE malware classification — roughly 1.1 million labelled samples with pre-extracted features, released specifically to give the research community a common, legally shareable benchmark without distributing live malware binaries. Reported accuracy figures on EMBER (typically well above 95% AUC for gradient-boosted models) are a reasonable reference point for the static classification sub-problem, but should not be read as a measure of real-world detection rate against a live, adaptive adversary — see the next lesson for why.
Cost
Dynamic sandbox analysis costs real compute time per file (often tens of seconds to minutes of actual execution, plus environment provisioning), which is why production pipelines apply it selectively — to files a fast static model flags as uncertain — rather than to every file scanned, mirroring the cascade pattern seen in the fraud-detection section (fast filter first, expensive check second).
Key references
- Anderson, H. S., & Roth, P. (2018). EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models. arXiv:1804.04637.
- Raff, E., et al. (2018). Malware Detection by Eating a Whole EXE (MalConv). AAAI Workshop — raw-byte CNN classification.
- Nataraj, L., et al. (2011). Malware Images: Visualization and Automatic Classification. VizSec — the byte-to-image representation.
- Rieck, K., et al. (2011). Automatic Analysis of Malware Behavior Using Machine Learning. Journal of Computer Security — a foundational dynamic-analysis ML treatment.
Current state
Gradient-boosted trees on hand-engineered static features remain competitive with, and often preferred over, raw deep learning approaches in production, largely because they are cheaper to run at endpoint scale and easier to audit when a detection needs to be explained to a security team — an interpretability pressure structurally similar to the one credit scoring faces from regulators.
What to learn next
- Text classification — the shared bag-of-features technique behind byte n-gram approaches.
- Feature importance done right — reading tree-ensemble importances correctly.
- Your opponent updates too — the adversarial dynamics that make static classification alone insufficient.