Nemotron Diffusion tests the one-token-at-a-time bottleneck
NVIDIA released tri-mode diffusion LLMs that switch between AR, diffusion, and self-speculation generation in one checkpoint.
NVIDIA released tri-mode diffusion LLMs that switch between AR, diffusion, and self-speculation generation in one checkpoint.
NVIDIA Verified Agent Skills treats agent skills as scanned, carded, and signed artifacts, pointing to a new supply-chain checkpoint for AI agents.
NVIDIA Vera CPU deliveries show that AI agent competition is shifting beyond GPU inference toward CPU-bound sandboxes, tool calls, and RL evaluation.
NVIDIA AI-Q agent skill lets Claude Code, Codex, and other harnesses delegate enterprise research to a local AI-Q server.
NVIDIA SANA-WM claims 720p, 60-second world modeling from a 2.6B backbone. The real story is not video polish but the cost structure of open models.
NVIDIA is positioning Nous Research Hermes Agent on RTX and DGX Spark as a local, always-on self-improving agent runtime.
NVIDIA and Ineffable Intelligence are pointing the model race toward RL infrastructure for agents that learn from experience.