We're splitting our production infrastructure to two servers this month — and there may be
We're splitting our production infrastructure to two servers this month — and there may be a brief outage during the migration.
Here's why.
Since launching ninja.ing, we've been running 11 production domains, 15+ containers, and multiple Neo4j graph databases on a single 64GB server. Signal's threat graph holds 160,000+ nodes. The ML workloads push CPU to 99%.
It worked. Until it didn't.
THE ATTACKS
Running a cybersecurity platform publicly means you become a target.
55,000+ automated scanner attempts blocked across our 11 domains. Within 60 seconds of deploying detection jails, 4 IPs were already caught.
We built an Adversary Behavioral DNA engine — 18-dimensional fingerprinting for every IP that touches our infrastructure. Temporal entropy, request velocity, path vocabulary, auth endpoint targeting. Every attacker classified into archetypes: scanners, brute forcers, bot crawlers, researchers, and targeted operators.
The system maps behavior to kill chain stages and uses Markov chains to predict an attacker's next move. IPs with cosine similarity above 0.85 get clustered as the same operator across different source addresses.
Real examples: Swedish IP trying /wp-admin/ on our SIEM domain. Credential stuffing across all 11 domains simultaneously. .git/config probes, .env extraction attempts, phpMyAdmin scans — every hour.
Three fail2ban jails run 24/7. Scanner detection bans for 24 hours. Auth failures ban for 1 hour. Aggressive 404 patterns ban for 12 hours.
THE SPLIT
Box 1 (NEW) — Ryzen 9 7950X3D, 128GB DDR5, 3.84TB NVMe. Dedicated to Signal and Fusion — the two apps consuming 80% of resources. The 3D V-Cache is chosen specifically for Neo4j graph traversal workloads.
Box 2 (CURRENT) — Ryzen 5 3600, 64GB. The other 13 apps that collectively use less than Signal alone.
Connected via WireGuard tunnel. Caddy on Box 1 as single front door, reverse-proxying lighter apps to Box 2. Same data centre.
3-LAYER BACKUP
Layer 1: Local — 18 repos, 11 Neo4j databases pulled from prod, all secrets. Daily.
Layer 2: Server self-backup — each box backs up its own volumes locally. 7-day retention.
Layer 3: Cross-box sync over WireGuard. If either box dies, the other has a copy.
THE RESULT
Signal's threat diff drops from 139 seconds to under 30. ML workloads get 128GB instead of fighting 13 apps for 64GB. Proper redundancy — if either server goes down, the other still serves.
There may be a brief outage — minutes, not hours. We'd rather be transparent about that.
This is what running production cybersecurity infrastructure actually looks like. Not just dashboards — the 3AM scanner blocks, the behavioral fingerprinting, and the honest acknowledgment that there might be 5 minutes of downtime while we make it better.
Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.
Subscribe to the Daily →
Scott Gardner ·