STORY · FORSKNING_

MUD games as LLM evaluation reveal problems with AI judging

Researchers used classic text-based games (MUD) from the 1970s to evaluate language models with just $99 in API costs. The experiment showed that when they removed two evaluation criteria based on LLM classifiers, a frontier model dropped six places, and the methods showed low reliability between different AI judges.

WHY IT MATTERS

The finding underscores a critical problem with benchmark-based LLM evaluation: using AI as a judge introduces systematic bias and noise that make it difficult to reliably compare models. This has implications for how the industry validates and ranks new models.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.