Home › Blog

What if you could fingerprint an attacker by their behaviour, not their IP address?

What if you could fingerprint an attacker by their behaviour, not their IP address?

Today I deployed something I've been thinking about since last summer: Adversary Behavioral DNA.

Most threat intelligence is IOC-based — IP blocklists, file hashes, domain feeds. The problem? Attackers rotate IPs. They use VPNs. They spin up cloud VMs. The IOC is dead before it hits your SIEM.

So I built something different. Every HTTP request that hits my infrastructure gets parsed, and every IP gets profiled across 18 behavioral dimensions:

 - Temporal entropy (how random is their timing?)
 - Request velocity and acceleration
 - Path vocabulary richness
 - Method and status code distributions
 - Inter-request timing statistics
 - Domain spread across the ecosystem
 - Auth attempt ratio
 - Sensitive path targeting ratio
 - User-agent consistency
 - Sequential pattern repetition (bigram repeat rate)

This creates a unique behavioral vector — a DNA fingerprint — for every attacker. From there:

Archetype classification scores each IP as a scanner, brute-forcer, researcher, bot/crawler, or targeted operator.

Kill chain mapping places observed behaviour on the Lockheed Martin framework — so you know if they're in reconnaissance or actively attempting exploitation.

Markov chain prediction learns transition probabilities between request categories and predicts their next move.

And the part I'm most proud of — same-operator detection. Using cosine similarity with a union-find clustering algorithm, the system identifies when the same human is operating from different IPs.

First real result? Two Microsoft Azure IPs from completely different subnets, both running PHP webshell scanners against my infrastructure. Traditional tools see two separate IPs. Adversary DNA returned 0.9975 cosine similarity — virtually identical behavioural fingerprints. Same toolkit, same path vocabulary, same timing patterns. One operator, two VMs. Confirmed.

No external ML libraries. No training data. No GPU. Pure statistical inference from raw access logs. ~800 lines of Python.

5 API endpoints, live in production, profiling every visitor in real time across 10 domains.

The industry spends billions on IOC feeds that expire in hours. Behaviour doesn't expire. An attacker can change their IP — they can't change how they think.

This is part of the ninja.ing intelligence platform — 14 platforms, 160K+ graph nodes, built solo from a bedroom in Glasgow. The full capability map is at https://lnkd.in/enc-YQD8

If you're working in threat intelligence, detection engineering, or behavioural analytics — I'd love to hear your thoughts. This is just the beginning.

#cybersecurity#threatintelligence#datascience#behaviouralanalytics#infosec#machinelearning#python#startups#buildinpublic
The Probably Fine Daily

Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.

Subscribe to the Daily →

View the original on LinkedIn ↗

← All writing