STORY · MODELLER_

Anthropic's new model cycle and agent benchmarks reveal persistent challenges

Anthropic launched the Mythos/Opus series where Claude Mythos gets praise for one-shot workflows, while Opus 4.8 shows regressions in benchmarks. Meanwhile, new testing tools like Agents' Last Exam and SWE-Marathon reveal that even top models like GPT 5.5 and Claude Opus 4.7 lack reliability gains on long-running tasks.

WHY IT MATTERS

This exposes a critical gap between model marketing and actual performance on complex agent tasks. The results clarify where AI systems still struggle and are driving agent framework development toward RL environment-style evaluation.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.