STORY · FORSKNING_
Test-time compute changes how AI models are measured, according to OpenAI researcher
OpenAI researcher Noam Brown argues that the industry's traditional benchmark grids don't account for how much time models are given to think. With large amounts of test-time compute, today's models can reason for weeks or months on complex tasks, making old evaluations meaningless.
WHY IT MATTERS
This shifts the foundation for how we measure AI models' capability and safety. If models can actually solve problems that previously seemed unsolvable given enough time, we need to reconsider both benchmarking practices and safety testing when capability scales with computational budget.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.