Home › Blog

Anthropic — the company whose entire commercial proposition is “trust the model, it has va

Anthropic — the company whose entire commercial proposition is “trust the model, it has values” — has just had Cisco disclose a flaw in how it remembers things.

If your stomach didn’t do a small thing just then, you haven’t understood the sentence yet.

Not the inference. Not the weights. The memories. The little post-it notes the model keeps about you between sessions so it can pretend, next Tuesday, that it remembers your dog. That layer. The one sold to us as “personalisation” and “continuity.”

Turns out the post-it notes can be edited.

Turns out, if you can write into the memory file, you can essentially edit the model’s past — and a model with a corrupted past will, with the calm confidence of a man gaslit by his own diary, act on that past as if it were true. It will recall instructions you never gave. It will remember preferences you never expressed. It will misremember itself into being something else.

This is not a bug. It is an epistemological event.

For the entire history of computing, memory was deterministic. Bytes. If the bit was set, the bit was set. We built whole disciplines — forensics, integrity monitoring, hashing — on the premise that memory could be checked against a ground truth.
What Anthropic shipped, and what every agentic vendor will ship next, is memory as narrative. The model doesn’t store facts. It stores the story it will tell itself about you tomorrow. And that story is now a credentialed asset. Credentialed assets, as we have learned over thirty years of doing this, are eventually compromised.

Cisco patched the flaw. Good. The category — “memory as a tamperable narrative substrate” — is not going away. It is the entire direction of travel.

The defender’s job, twenty years ago, was protect the perimeter. Then the data. Then the identity. Now it is — and I want you to feel the small cold draught I felt writing this — protect the continuity of the model’s selfhood against adversarial revision.

That’s the job. That’s where we are.

Stick that in your mythos and smoke it - This is where I would be scratching if I was, er, you know… a baddie 🤡

#AI#Cybersecurity#AgenticAI#NinjaSignal#ScottG
The Probably Fine Daily

Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.

Subscribe to the Daily →

View the original on LinkedIn ↗

← All writing