STORY · MODELLER_

Qwen3-Max Thinking beats Gemini 3 Pro and GPT-5.2 on Humanity's Last Exam

Alibaba's Qwen3-Max Thinking model achieved the highest score on the Humanity's Last Exam benchmark test compared to Google's Gemini 3 Pro and OpenAI's GPT-5.2. The model was tested with search functionality enabled.

WHY IT MATTERS

It shows that Qwen now competes with the leading labs on advanced reasoning tasks, and suggests that competition is opening up at the frontier of model development outside the US.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.