STORY · FORSKNING_
MUD games as LLM evaluation reveal problems with AI judging
Researchers used classic text-based games (MUD) from the 1970s to evaluate language models with just $99 in API costs. The experiment showed that when they removed two evaluation criteria based on LLM classifiers, a frontier model dropped six places, and the methods showed low reliability between different AI judges.
WHY IT MATTERS
The finding underscores a critical problem with benchmark-based LLM evaluation: using AI as a judge introduces systematic bias and noise that make it difficult to reliably compare models. This has implications for how the industry validates and ranks new models.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.