Skip to content

Nemotron 3 Ultra 550B

Also accepted:nvidia/nemotron-3-ultranemotron-3-ultranvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron-3-Ultra-550B-A55B is a 550B parameter (55B active) frontier model built on a LatentMoE hybrid architecture combining Mamba-2, MoE, and Attention with Multi-Token Prediction. Features a 1M token context window, configurable reasoning mode (enable_thinking), and strong multilingual support across English, French, Spanish, Italian, German, Japanese, Korean, Hindi, Brazilian Portuguese, and Chinese. Best suited for complex agentic workflows, long-context analysis, tool use, and high-stakes RAG.

Providers
Capabilities
AIHubMix
aihubmix-byok
$0
$0
Unavailable
NVIDIA
nvidia-byok
$0
$0
Unavailable
OpenRouter
openrouter-byok
$0
$0
Unavailable
Ollama Cloud
ollama-byok
$0
$0
Unavailable
Baseten
baseten-byok
$0
$0
Unavailable
OpenCode Zen
opencode-zen-byok
$0
$0
Unavailable
CommandCode
commandcode-byok
$0
$0
Unavailable
NVIDIA
nvidia
$0.50
$2.50
Usage analytics

Loading usage…

API & code
Uptime & Health
No uptime data yet

These providers haven't been health-probed for this model yet. The router still routes around upstreams that fail live requests — uptime fills in once probe history accrues.

Share cards
Nemotron 3 Ultra 550B share card
Nemotron 3 Ultra 550B
AIHubMix upstream share card
AIHubMix upstream
NVIDIA upstream share card
NVIDIA upstream
NVIDIA upstream share card
NVIDIA upstream
Hue upstream share card
Hue upstream
OpenRouter upstream share card
OpenRouter upstream
Ollama Cloud upstream share card
Ollama Cloud upstream
Baseten upstream share card
Baseten upstream
OpenCode Zen upstream share card
OpenCode Zen upstream
CommandCode upstream share card
CommandCode upstream
Credits
Use your own key

Run Nemotron 3 Ultra 550B on your own key — your requests are billed by the provider. Pool callers pay AnyRouter credits.

No BYOK keys configured for this model yet.

Share a key with the pool to earn credits for every request it serves, covering your plan cost.

Text generation
Context length1,000,000 tokens
Max output65,536 tokens
ArchitectureTransformer
Categorytext
ReleasedJun 4, 2026
Modalities
Capabilities