Lab leak

Everyone has their crystal balls out when it comes to AI. Here’s my take:

Within the next 12–18 months the weights of a major closed frontier model (think Opus 5 or GPT-5.6) will be leaked to the public.

I put the chance at 50%.

This could happen through a cyberattack on a lab,1 or a security mistake,2 or a disgruntled employee gone rogue. Most outlandish would be rogue agents who manage to access their own model weight files and decide to transmit them away, perhaps out of a sense of self-preservation.

The weights are the labs' holiest-of-holies, and these scenarios must keep them up at night. There is real danger in these weights getting out there, if jailbreaks or de-alignments are possible. At least one saving grace is the files must be massive and it would be hard to conceal them being uploaded.

What a world.


  1. It’s common speculation that a state actor could steal the weights (perhaps this has happened already). But theft isn’t a public release. ↩︎

  2. Anthropic once leaked the source code for Claude Code↩︎

Next
Previous