Discussion about this post

User's avatar
RichardYe's avatar

Thanks Buck for the post! I appreciate this focus on the motivations, thought-processes, and constraints insiders at the labs face. I think many of the points here were spot on.

I do worry the AI Control community hasn't sufficiently updated on hacking capabilities though. Even moderate hacking capabilities of coding agents running internally significantly increases the risk of exfil or a rouge internal deployment. The problem is so many monitors and control techniques have nontrivial attack surfaces. And for testing, companies are running helpful-only models relatively frequently. And explicitly giving them targets to attack all while with more limited monitoring. While I think some of the Mythos-thing was over-hyped, the general trend toward better hacking will just continue. I'd love for a set of cheap control techniques that companies can grab and implement, but worry that future hacking capabilities could wash away many of the gains.

No posts

Ready for more?