Home › Blog

A day in the life of VibeOps...

Today I migrated a production security platform to a new server, ran a data science audit on a million-node graph, fixed 19 data quality issues, and tuned Linux kernel parameters. Before dinner.

Let me walk you through it.

06:00 — The Great Migration

We went live with a 2-box architecture. Box 1: a Ryzen 9 7950X3D with 128GB DDR5 and 3.84TB NVMe. Box 2: the old faithful Ryzen 5 3600 with 64GB. Connected by a WireGuard tunnel running at sub-millisecond latency.

Signal (our threat intelligence platform) and Fusion got promoted to the big box. Everything else stays on Box 2. Caddy on Box 1 is now the front door for all 11 domains, reverse-proxying to Box 2 over WireGuard for the rest.

Neo4j volumes copied. User data migrated. 5 TLS SNI proxy snippets configured. DNS cutover. 3-layer backup strategy across both boxes with cross-replication over WireGuard at 05:00 UTC daily.

Zero downtime. Zero data loss. Zero drama. Suspicious.

10:00 — The Data Science Audit

With the new hardware humming, we unleashed 4 parallel data science agents on the Neo4j graph. 1.07 million nodes. 12.8 million edges. The findings were... educational.

54% of the graph was Sighting nodes. Five hundred and seventy-four thousand OpenCTI provenance artifacts that no human has ever queried. Over half a million nodes just... existing. Consuming resources. Contributing nothing. We've all worked with someone like that.

46,000 Indicator nodes had zero relationships. Orphans. Floating in the void. Connected to nothing. Predicting nothing. Indicating nothing.

39,000 Indicators had their type set to "Indicator." That's like filling in your job title as "Employee." Technically correct. Completely useless.

3 threat actors had identity crises: "MuddyWater" and "muddywater" were living separate lives in the graph. Same actor. Different nodes. Different relationships. Nobody introduced them.

Out of 20 data ingesters, only 4 had created Source nodes. The other 16 were dumping intelligence into the graph anonymously. Like leaving Post-it notes on someone's desk with no signature.

And my personal favourite: every ingester was faithfully calculating confidence scores, passing them through the entire pipeline, and then... not writing them to the database. The digital equivalent of doing your homework and leaving it on the kitchen table.

14:00 — The Purge

Round 1: 10 cleanup operations. Campaign deduplication. Software nodes named "2024-03-15" (that's a date, not malware). Fancy Bear living separately from APT28. Fixed.

Round 2: 10 more. The Sighting massacre. The orphan cull. Case-insensitive actor merges. Indicator type detection by regex pattern matching. Vulnerability nodes with NULL names getting them from their own cve_id field. The creation of 16 missing Source nodes.

Total damage: approximately 620,000 nodes deleted. The graph lost half its weight and gained twice the intelligence.

16:00 — "Why Is Neo4j at 303% CPU?"

Ah. Yes. That.

The Ryzen 9 7950X3D — 32 cores, 128GB RAM — was running Neo4j with a 2GB heap. Two gigabytes. For a million-node graph with Graph Data Science and APOC loaded.

The JVM was spending more time garbage collecting than actually processing queries. It was like giving a Formula 1 car a 2-litre fuel tank and wondering why it keeps stopping.

Neo4j's own memory-recommendation tool said "4GB minimum please." We were giving it half that.

Fixed: heap to 8GB. Page cache to 8GB. Container limit to 28GB. Added -XX:+ExitOnOutOfMemoryError because a half-crashed database is worse than a fully crashed one.

CPU dropped from 303% to 1.4%.

18:00 — "Go Deeper"

Then we went full Linux kernel tuning.

Disabled Transparent Huge Pages. THP is the silent killer of JVM workloads — the kernel helpfully defragments memory into 2MB pages while your database stutters for 200ms wondering what happened. Set to never. Created a systemd service to keep it dead across reboots.

Set noatime,nodiratime on the root filesystem. Every single file read was updating metadata timestamps. On a database that reads millions of files. ext4 was keeping a diary nobody asked for.

Cranked readahead from 256 to 4096. Graph traversals aren't random 4K reads — they're sequential scans across relationship chains. Let the kernel prefetch.

Dropped vm.dirty_ratio from 40 to 10 and vm.dirty_background_ratio from 10 to 5. The kernel was hoarding 40% of RAM in dirty pages before flushing. For a database. That needs predictable I/O latency. Flushing 50GB of dirty pages in one burst is not "predictable."

Bumped vm.max_map_count to 524K because Neo4j memory-maps everything and the default was getting tight.

Tripled the TCP buffer sizes. The Bolt protocol was pushing graph query results through 208KB buffers. That's a garden hose for a fire hydrant.

All persisted. sysctl.d config. fstab entries. systemd oneshot service. This box will boot tuned.

20:00 — The Numbers

Box 1 final state:

- 32 cores, 128GB RAM, 3.84TB NVMe

- Neo4j: 8GB heap + 8GB page cache + 12GB native headroom = 28GB container

- 78GB RAM still free. The server is basically on holiday.

- CPU: 1.4% at idle (was 303%)

- All 11 domains serving through Caddy with TLS

- 3-layer backup strategy across both boxes

- Linux kernel tuned to vendor spec and beyond

620,000 useless nodes eliminated. 19 data quality issues fixed. Confidence scores actually persisting. Source provenance for all 20 ingesters. Tactics extracted from STIX kill chain phases.

One day. Two boxes. Zero downtime.

Moo Moo still must die.

#CyberSecurity #ThreatIntelligence #Neo4j #GraphDatabase #Linux #PerformanceTuning #Infrastructure #DataEngineering #NinjaEngineering

The Probably Fine Daily

Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.

Subscribe to the Daily →

Originally published on LinkedIn ↗

← All writing