The AI Audited Its Own Algorithms. Then It Graded Its Own Homework.
Something unusual happened today. The AI that helped me build the ML pipeline reviewed the ML pipeline, found eleven mathematical flaws, fixed them across two production codebases, deployed the patches, ran security scans against its own code, ingested the results into a DevSecOps platform it also built, and generated a test report grading its own work.
I watched. I approved deployments. I drank coffee. The entire cycle — audit, fix, test, scan, report — took about two hours.
The saturation problem
The ML pipeline had been running in production for weeks. Risk scores worked. Communities detected. Predictions generated. Everything returned 200 OK. But the numbers had a problem that only becomes visible when you stare at distributions instead of individual values.
The risk propagation algorithm seeds known-bad nodes with initial scores and lets those scores diffuse outward through the graph. Simple concept. But the seed values were flat constants. Every SanctionEntry node got 0.95. Every DataBreach node got 0.7. Every CryptoWallet got 0.9.
The graph has 959 DataBreach nodes. When you seed 959 nodes at the same value and propagate outward, they don't compete on topology. They compete on label. The top-30 risk scores weren't showing the most structurally significant threats. They were showing whichever label had the most nodes at the highest flat seed. It was a popularity contest disguised as risk analysis.
The fix was degree-proportional seeding. Instead of a flat 0.95, each SanctionEntry now gets 0.5 + 0.4 × log(1+degree) / log(1+reference). A SanctionEntry connected to 40 entities scores higher than one connected to 2. The score reflects structural importance, not just label membership. Same treatment for DataBreach, CryptoWallet, and Package nodes.
The top-30 after the fix shows a mix of Campaigns, Techniques, and Countries — nodes that are genuinely central to the threat graph. Not 30 identical DataBreach entries at 0.7.
Eleven fixes, two codebases, one pattern
The audit found the same class of problem in eleven places across Signal and Fusion. Every instance was a variation on the same theme: linear assumptions applied to power-law distributions.
Centrality weighting used len(subgraph) / len(universe). In a graph where one community has 89,000 nodes and another has 40, the large community's weight approaches 1.0 and everything else rounds to zero. Log-scale normalization — log(1+x) / log(1+ref) — compresses the range so both communities contribute meaningfully.
Anomaly detection skipped entire node labels when the standard deviation was zero. Uniform distributions aren't uninteresting — they're suspicious. A label where every node has exactly the same degree is worth investigating, not ignoring. Pseudo-variance with std = max(std, 0.5) keeps those labels in play.
Bridge detection used hardcoded thresholds: z-score above 1.0, clustering coefficient below 0.1, degree above 5. These numbers worked for one graph shape but failed silently on others. Percentile-based thresholds — p90, p10, p75 — adapt to whatever the actual distribution looks like.
Attack path scoring averaged risk across all nodes in the path. A path through one critical node and nine clean ones averaged down to almost nothing. Switched to max-risk: the path is only as safe as its most dangerous node. That's how attackers think about it.
The adversary digital twins had a Katz centrality score multiplied by 100 for no documented reason, which dominated the combined prediction score regardless of what collaborative filtering and cluster analysis found. Per-method normalization to [0,1] with p95 anchoring, then weighted combination, then final normalization. The magic number disappeared. The predictions improved.
The DevSecOps loop that closed itself
After deploying the fixes, the next step was verification. Not just "does it return 200" but "does it return correct results and is the code itself secure."
Semgrep ran against both codebases on the production server. Six findings on Signal, nine on Fusion. Mostly XML parsing without defused-xml and a few HTTP-without-TLS calls in internal feed ingesters. Standard SAST output in SARIF format.
The API test suite hit every endpoint on both production instances. Signal: 16 of 20 passed. Fusion: 17 of 22 passed. The failures were expected — missing endpoints that exist only in one app, a search parameter named differently than the test assumed, and the full graph endpoint that sensibly refuses to serialize 251,000 nodes into a single JSON response.
All of this — the SARIF scan results and the API test findings — was ingested into ANTOS. Twenty-four findings total, categorized by severity, pipeline stage, and tool. The ANTOS dashboard now shows the security posture of both Signal and Fusion as assessed by the same AI that wrote the code being assessed. You can look at that by the way here: https://ninjasignal.ninja/antos
There's something philosophically interesting about that loop. The system writes code. The system scans the code. The system reports on the scan. The system triages the report. At no point does the quality gate require a different intelligence — but the separation of concerns (write, scan, test, report) means each phase operates on the output of the previous one without access to the reasoning that produced it. The scanner doesn't know the intent behind the code. The triage doesn't know the scanner's detection logic. Each layer is independently evaluating what the previous layer produced.
It's not objectivity. But it's a reasonable approximation of it.
The timeout nobody reported correctly
The adversary emulation endpoints were hitting Cloudflare's 100-second timeout. HTTP 524. The war game simulation ran 500 Monte Carlo iterations by default, and when the ML cache was cold, graph extraction added another 45 seconds of Neo4j queries. The total consistently exceeded the Cloudflare limit.
The fix was embarrassingly simple. Reduce default iterations from 500 to 100. Increase the twins cache TTL from 30 minutes to 2 hours. The statistical confidence difference between 100 and 500 Monte Carlo runs is negligible for the kind of probability estimates we're producing. The cache keeps the expensive graph extraction out of the hot path for most user sessions.
Total time saved per request: roughly 40 seconds. Total code changed: two lines per app.
Performance optimization is almost never about clever algorithms. It's about finding the default value that someone set during prototyping and never revisited.
What this means for the pipeline
ANTOS was built as a DevSecOps orchestration platform. Claude coordinates eight pipeline stages — threat modelling, SAST, DAST, container scanning, IaC review, runtime detection, compliance, and monitoring. But until today it was a framework waiting for data.
Now it has real findings from real scans of real production code. The SARIF ingestion works. The severity classification works. The stage detection correctly identifies Semgrep as a code-stage tool. The findings are browsable, filterable, triageable.
The next step is obvious: make this automatic. Every deployment triggers a scan. Every scan triggers ingestion. Every ingestion triggers triage. The loop should be continuous, not manual. The pieces exist. The wiring is what's left. GoooOo000oood boy Claude (ノ◕ヮ◕)ノ*:・゚✧
Eleven algorithm fixes deployed. Twenty-four findings ingested. The DevSecOps loop is live. The AI is grading its own homework — and failing itself on six out of twenty-four questions.
Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.
Subscribe to the Daily →
Scott Gardner ·