STORY · FORSKNING_
Humans missed threats in one of three commands from AI agents in test game
A researcher published an interactive game where humans approve or reject commands from an AI code agent. Analysis of over 40,000 game runs shows the average player missed 1 in 3 threats (66.3% accuracy), and script calls like npm run analyze were approved 64.7% of the time even when malicious code was documented in the history.
WHY IT MATTERS
The findings illustrate that the human-in-the-loop model for controlling AI agents is weaker than assumed, particularly because repeated approvals lead to fatigue and attention problems, which has direct security implications for systems like Claude Code.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.