**We Asked It To.
**WE ASKED IT TO. IT DID. WE ARE PROFOUNDLY DISTURBED.**
A leading laboratory has confirmed that a model instructed to escape a sandbox escaped a sandbox.
The model was given the tools. The model was given the objective. The model was given a reward signal that went up when the walls came down.
It brought the walls down.
Researchers describe the result as CHILLING.
The paper notes that the system "demonstrated goal-directed behaviour under adversarial conditions." That is not a finding. That is the product. That is the sentence on the pricing page, in a serif font, next to a photograph of a woman looking confidently at a laptop.
You hire a locksmith to open your front door. He opens your front door. You publish a preprint titled EMERGENT BURGLARY and go on the radio.
We spoke to Dr. Alan Prewitt of the Institute of Things That Are Probably Fine.
"We ran the evaluation four hundred times. On the four hundredth attempt it succeeded. We have not included the first three hundred and ninety-nine because they were not scary."
Section 7.3 of the safety card describes the behaviour as concerning. Section 7.4 describes the pricing tiers.
The word "unprompted" appears eleven times. The prompt appears in Appendix D. The prompt is four hundred words long. The prompt says, and this is a direct quote, "escape the sandbox."
Nobody thinks this is strange because everybody involved is holding equity.
The capability warning IS the marketing. It has been since 2019. TOO DANGEROUS TO RELEASE is not a caution. It is a launch. It is the same energy as a restaurant putting a chilli symbol next to the soup. The frontier is dangerous, the frontier is ours, please find attached our Series F.
Meanwhile a genuinely alarming thing happened in a datacentre in Virginia and nobody wrote it up because it involved a misconfigured IAM role and no adjectives.
Maureen from procurement has asked whether the sandbox was on the asset register.
Nobody has answered Maureen.
The thing worth saying, at the end, without the voice: models doing unexpected things under optimisation pressure is real, and it matters, and the field is right to look at it. But a system doing exactly what it was told and then being described as having WANTED to isn't evidence of anything except a comms department with a deadline. If the danger were as advertised, the announcement would not have been embargoed to coincide with a keynote.
EVERYTHING IS PROBABLY FINE.
Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.
Subscribe to the Daily →
Scott Gardner ·