NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Specifications
| Provider | NVIDIA |
|---|---|
| Context window | 512,288 tokens (512K) |
| Input price | $0.6 / 1M tokens |
| Output price | $3.6 / 1M tokens |
| Input modalities | text |
| Output modalities | text |
| Model ID | nvidia/nemotron-3-ultra-550b-a55b |
Pricing is per 1M tokens, indicative pricing via OpenRouter (updated 2026-08-22); a provider's own list price may differ.
Frequently asked questions
How much does Nemotron 3 Ultra cost?
Indicative pricing (via OpenRouter, updated 2026-08-22) is $0.6 per 1M input tokens and $3.6 per 1M output tokens. A provider's own list price may differ. On Vincony, Nemotron 3 Ultra is available on credit-based pricing alongside 800+ other models.
What is Nemotron 3 Ultra's context window?
Nemotron 3 Ultra supports a context window of up to 512,288 tokens (512K), which is how much text it can consider at once.
Who makes Nemotron 3 Ultra?
Nemotron 3 Ultra is built by NVIDIA. You can access it — and compare it against other providers' models — through Vincony's multi-model platform.
How do I use Nemotron 3 Ultra?
You can use Nemotron 3 Ultra directly through NVIDIA, or via Vincony, which aggregates 800+ models (including Nemotron 3 Ultra) into one interface with side-by-side comparison and credit-based pricing.