Llama 3.3 Nemotron Super 49B V1.5
nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA's Llama 3.3 Nemotron Super 49B V1.5 is a text‑only model with a context length of 131,072 tokens. It supports tool usage and reasoning, costs $0.40 per million input tokens and $0.40 per million output tokens, and its weights are not open. The model was released on 2025‑03‑16.
- tool use
- structured outputs
- reasoning
At a glance
Specifications
What the API gives you
- Context window
- 131K tokens
- Max output
- 16K tokens
- Input price
- $0.40 / 1M tokens
- Output price
- $0.40 / 1M tokens
- Cache read
- —
- Modalities
- text → text
- Knowledge cutoff
- Mar 31, 2024
- Released
- Mar 16, 2025
Recommendation coverage
Where this model appears
Not currently included in a published recommendation. This means the available task-specific evidence did not support a ranking yet—not that the model was reviewed and rejected.
Benchmarks
No independent benchmark results are linked to this model yet. We keep it visible while withholding quality claims until usable results arrive.
Open-weight model on Hugging Face.
Catalog data via OpenRouter and models.dev. Last checked within the past 6 hours.