AI in Cybersecurity

Malware classification

A malware classifier looks for measurable telltale signs in a file, the same way a bank teller spots a fake note by feel, without reading every byte the way a human would read a document.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Malware classification decides whether a file is malicious by measuring specific, checkable properties. Not by "reading" it like a person would.

Think about a bank teller checking a currency note. They do not read every printed word on the note. They check specific things: the watermark, the security thread, how the paper feels, whether the number matches under UV light. A handful of checkable signals is enough to decide fast. Signals learned from handling thousands of real and fake notes.

A malware classifier checks a file the same way. Measurable properties, not a full read.

Why it exists

Antivirus software used to work almost entirely by signatures — an exact fingerprint of a known bad file. This works perfectly for malware seen before. It fails completely against anything new. A signature only matches an exact known file.

Malware authors know this. They change their file slightly with every release, specifically to dodge signature matching. This technique is called polymorphism. A machine learning classifier looks at properties of a file instead of an exact fingerprint. How random do the file's bytes look? Does it ask the operating system to do suspicious things? Is it digitally signed by a known publisher? A slightly modified file usually still shows the same suspicious properties, even with a completely different signature.

How it works

A file arrives
        |
        v
  [ measure properties, WITHOUT running it ]
    - how random do the bytes look? (packed/encrypted files look very random)
    - is it digitally signed by a known publisher?
    - what system functions does it ask to use?
    - how big is it, compared to what it claims to be?
        |
        v
  [ classifier ]
        |
        v
  "malicious, 94% confident"  or  "benign, 98% confident"

This is called static analysis: examining the file without executing it. It is fast and safe. The opposite approach, dynamic analysis, actually runs the file in an isolated, contained sandbox, and watches what it does. It is slower, but harder for malware to fool. The next lesson covers this further.

A real example you have seen

Modern antivirus tools (Windows Defender, and most enterprise security products) combine exactly this. A fast, static ML check runs on every new file. Known threats get caught by signatures. Anything still uncertain gets sandboxed.

Remember this

  • A classifier judges a file by measurable properties, not by reading it like text. Randomness of its bytes, whether it is signed, what it asks the system to do.
  • This approach catches new, slightly modified malware that exact signature matching misses.
  • Static analysis (checking properties without running the file) is fast and safe. Dynamic analysis (running it in a sandbox) is slower but harder to fool.
  • A "malicious" verdict should trigger quarantine and review, not silent auto-deletion, until a security analyst confirms it.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Minimal runnable code

A small synthetic dataset built from well-documented, published malware heuristics — high byte entropy from packing, unsigned binaries, and suspicious system calls. No real malware files are used or needed.

malware_classifier.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier

# Each row: [byte_entropy(0-8), suspicious_api_calls, is_digitally_signed(0/1), file_size_kb]
files = np.array([
    [3.1, 0, 1, 850],   # ordinary signed installer
    [2.8, 1, 1, 420],
    [7.9, 12, 0, 38],   # packed, unsigned, tiny -- classic packed-malware shape
    [3.5, 0, 1, 1200],
    [7.6, 9, 0, 55],
    [2.9, 1, 1, 300],
    [7.8, 15, 0, 22],
    [3.2, 0, 1, 900],
    [7.4, 10, 0, 60],
    [3.0, 2, 1, 500],
])
is_malicious = np.array([0,0,1,0,1,0,1,0,1,0])

model = RandomForestClassifier(n_estimators=100, random_state=0)
model.fit(files, is_malicious)

test_files = np.array([
    [3.3, 1, 1, 700],   # looks ordinary
    [7.7, 11, 0, 41],   # looks packed and suspicious
])
predictions = model.predict(test_files)
probabilities = model.predict_proba(test_files)[:, 1]

for i, (pred, prob) in enumerate(zip(predictions, probabilities)):
    label = "MALICIOUS" if pred == 1 else "benign"
    print(f"file {i}: predicted={label}  malicious_probability={prob:.2f}")

print()
importances = dict(zip(
    ["entropy", "suspicious_api_calls", "is_signed", "file_size_kb"],
    model.feature_importances_.round(2),
))
print("which features the model leaned on:", importances)
Output
file 0: predicted=benign  malicious_probability=0.00
file 1: predicted=MALICIOUS  malicious_probability=0.99

which features the model leaned on: {'entropy': 0.21, 'suspicious_api_calls': 0.26, 'is_signed': 0.28, 'file_size_kb': 0.24}

What actually happened

The two test files were built to resemble two clear archetypes from the training data — file 0 looks like the ordinary signed installers, file 1 looks like the packed, unsigned samples. The model separated them cleanly, and feature_importances_ shows all four properties contributed meaningfully, with none dominating completely — a sign the model is not relying on a single brittle rule.

Byte entropy deserves a plain explanation: it measures how random a sequence of bytes looks, on a scale where ordinary program code and text sit around 2 to 5, and encrypted or compressed data sits close to the maximum of 8. Malware is very often packed — compressed or encrypted to hide its real contents from static tools — which pushes entropy up sharply. That single number, computed without running the file at all, is one of the most useful cheap signals in this entire field.

  • RandomForestClassifier was used, not logistic regression, because the true relationship here is not smoothly linear — entropy above roughly 7 matters a great deal, entropy between 3 and 5 barely matters at all, a "threshold-like" pattern trees capture naturally.
  • feature_importances_ for a random forest reflects how much each feature reduced impurity across all trees — a useful first read, though it can be misleading with correlated features, exactly as flagged for logistic regression coefficients in the finance section.

Common mistakes

Trusting static features alone against a well-resourced attacker. Entropy and signing status are cheap to fake or dodge — an attacker can sign a malicious file with a stolen certificate, or use techniques that keep entropy artificially low. Static analysis is a fast first filter, not a complete defence.

Training on too small or unrepresentative a sample. Real malware families vary enormously. Ten synthetic rows exist to show the mechanics; a real classifier trains on labelled datasets with many thousands of diverse samples, both malicious and benign, collected under careful legal and safety controls.

Running suspicious files to generate features. Even opening or running a file "to take a look" outside a proper sandbox can infect the machine doing the checking. Static features are computed by parsing the file's structure, never by executing it.

Ignoring class imbalance. In production, the overwhelming majority of files scanned are benign. This toy example is roughly balanced for clarity — see imbalanced data for how a real deployment has to handle the realistic ratio.

Try it yourself

Add a tricky file: [7.5, 1, 1, 900] — high entropy (looks packed) but signed and only one suspicious API call. This resembles a legitimate, commercially packed and signed application. Check what the model predicts, and consider why entropy alone is not a safe rule to hard-code.

What to learn next

Researcher — Mathematics and papers.

Feature families in malware classification

Production systems typically draw on several distinct feature families, often combined:

  • Static structural features — PE (Portable Executable, the Windows binary format) header fields, section table entropy per-section rather than whole-file, imported/exported function tables, digital signature validity chain.
  • Byte n-gram / string features — frequency of byte sequences or extracted printable strings, treated as a bag-of-features exactly analogous to bag-of-words text classification; see text classification for the shared underlying technique.
  • Dynamic behavioural features — API call sequences observed during sandboxed execution (process creation, registry writes, network connections attempted), capturing intent that static analysis of an obfuscated file can miss entirely.
  • Raw-byte deep learning — treating the executable's raw bytes as a 1D sequence (Raff et al., 2018, MalConv) or reshaping them into a 2D "image" (Nataraj et al., 2011) and applying CNN architectures directly, avoiding hand-engineered features at the cost of interpretability.

Static analysis is fundamentally limited by obfuscation

Packing, encryption, and code virtualization can, in principle, make a file's static byte-level properties resemble anything the attacker chooses, including deliberately mimicking benign software's entropy profile. This is not a hypothetical: publicly documented "packer as a service" tooling exists specifically to defeat static ML and signature detection. This is the direct motivation for combining static classification with dynamic sandbox analysis and endpoint behavioural monitoring (EDR) rather than relying on any single layer.

Sandbox evasion

Malware increasingly checks whether it is running inside a sandbox before revealing malicious behaviour — detecting virtualization artifacts, unrealistic system uptime, absence of normal user activity, or known analysis tool processes — and behaves benignly if so, defeating dynamic analysis in turn. This is a direct, ongoing arms race between defenders and malware authors, covered structurally in the next lesson.

Evaluation and public benchmarks

The EMBER dataset (Anderson & Roth, 2018, Endgame/Elastic) is the most widely used open benchmark for static PE malware classification — roughly 1.1 million labelled samples with pre-extracted features, released specifically to give the research community a common, legally shareable benchmark without distributing live malware binaries. Reported accuracy figures on EMBER (typically well above 95% AUC for gradient-boosted models) are a reasonable reference point for the static classification sub-problem, but should not be read as a measure of real-world detection rate against a live, adaptive adversary — see the next lesson for why.

Cost

Dynamic sandbox analysis costs real compute time per file (often tens of seconds to minutes of actual execution, plus environment provisioning), which is why production pipelines apply it selectively — to files a fast static model flags as uncertain — rather than to every file scanned, mirroring the cascade pattern seen in the fraud-detection section (fast filter first, expensive check second).

Key references

  • Anderson, H. S., & Roth, P. (2018). EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models. arXiv:1804.04637.
  • Raff, E., et al. (2018). Malware Detection by Eating a Whole EXE (MalConv). AAAI Workshop — raw-byte CNN classification.
  • Nataraj, L., et al. (2011). Malware Images: Visualization and Automatic Classification. VizSec — the byte-to-image representation.
  • Rieck, K., et al. (2011). Automatic Analysis of Malware Behavior Using Machine Learning. Journal of Computer Security — a foundational dynamic-analysis ML treatment.

Current state

Gradient-boosted trees on hand-engineered static features remain competitive with, and often preferred over, raw deep learning approaches in production, largely because they are cheaper to run at endpoint scale and easier to audit when a detection needs to be explained to a security team — an interpretability pressure structurally similar to the one credit scoring faces from regulators.

What to learn next