STORY · VERKTOY_
smevals – tool for evaluating AI models and prompts
Simon Willison has worked with Jesse Vincent's Prime Radiant lab to develop smevals, an open tool for running small evaluation suites across different model configurations. The tool lets users define evaluations as YAML files, run them against models like GPT-5.5 or Claude Opus 4.6, and visualize results through a web server or static HTML report.
WHY IT MATTERS
smevals addresses a long-standing need in AI development to systematically evaluate models' capabilities on metrics defined by the user. The tool lowers the barrier for people who want to benchmark models and prompts.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.