STORY · MODELLER_

NVIDIA launches Nemotron 3 Ultra, 550B MoE model optimized for agent workflows

NVIDIA released the fully open Nemotron 3 Ultra, a 550B MoE model with 55B active parameters and 1M context window, promising 5x faster throughput and 30% cost reduction for long-running agent tasks. The model uses a hybrid Mamba/attention architecture, LatentMoE and native MTP, and was pretrained on 20T tokens. In parallel, Anthropic shared observations on recursive self-improvement in Claude models, where the models themselves wrote 80%+ of merged code patches and enabled trainers to ship 8x more code.

WHY IT MATTERS

Nemotron 3 Ultra opens significantly faster and cheaper agent development for the developer community, while Anthropic's findings on recursive self-improvement point toward a possible qualitative shift in how AI assistants can accelerate AI development itself.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.