Lab leak
Everyone has their crystal balls out when it comes to AI. Here’s my take:
Within the next 12–18 months the weights of a major closed frontier model (think Opus 5 or GPT-5.6) will be leaked to the public.
I put the chance at 50%.
This could happen through a cyberattack on a lab,1 or a security mistake,2 or a disgruntled employee gone rogue. Most outlandish would be rogue agents who manage to access their own model weight files and decide to transmit them away, perhaps out of a sense of self-preservation.
The weights are the labs' holiest-of-holies, and these scenarios must keep them up at night. There is real danger in these weights getting out there, if jailbreaks or de-alignments are possible. At least one saving grace is the files must be massive and it would be hard to conceal them being uploaded.
What a world.
-
It’s common speculation that a state actor could steal the weights (perhaps this has happened already). But theft isn’t a public release. ↩︎
-
Anthropic once leaked the source code for Claude Code. ↩︎