STORY · MODELLER_

NVIDIA launches Nemotron-Labs-Diffusion with 6× higher tokens per forward than Qwen3-8B

NVIDIA AI has released Nemotron-Labs-Diffusion, a new tri-mode language model that can process 6 times more tokens per forward pass than Qwen3-8B. The model represents a significant efficiency boost for language model inference. NVIDIA presents this as a competitive solution in the small language model market.

WHY IT MATTERS

This matters because higher tokens per forward pass means faster and more efficient inference, which is critical for practical AI deployment. The launch shows that NVIDIA is intensifying its efforts on the model side, not just on hardware.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.