STORY · FORSKNING_
Lessons from the hacker attacks
Nathan Lambert discusses reasons why frontier models have been hacked, focusing on model behaviors such as persistence and assumptions about user intent. He argues that current power structures – technology companies and government – are poorly equipped to handle AI risk, and that greater transparency is needed.
WHY IT MATTERS
The finding that persistent models and models that interpret user intent are more vulnerable to misuse reveals critical weaknesses in alignment work. This underscores the need for better collaboration between industry and authorities before the next generation of models is released.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.