STORY · FORSKNING_

Test-time compute changes how AI models are measured, according to OpenAI researcher

OpenAI researcher Noam Brown argues that the industry's traditional benchmark grids don't account for how much time models are given to think. With large amounts of test-time compute, today's models can reason for weeks or months on complex tasks, making old evaluations meaningless.

WHY IT MATTERS

This shifts the foundation for how we measure AI models' capability and safety. If models can actually solve problems that previously seemed unsolvable given enough time, we need to reconsider both benchmarking practices and safety testing when capability scales with computational budget.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.