STORY · MODELLER_
Qwen3-Max Thinking beats Gemini 3 Pro and GPT-5.2 on Humanity's Last Exam
Alibaba's Qwen3-Max Thinking model achieved the highest score on the Humanity's Last Exam benchmark test compared to Google's Gemini 3 Pro and OpenAI's GPT-5.2. The model was tested with search functionality enabled.
WHY IT MATTERS
It shows that Qwen now competes with the leading labs on advanced reasoning tasks, and suggests that competition is opening up at the frontier of model development outside the US.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.