Llama 3.3 Nemotron Super 49B V1.5

nvidia/llama-3.3-nemotron-super-49b-v1.5

NVIDIA's Llama 3.3 Nemotron Super 49B V1.5 is a text‑only model with a context length of 131,072 tokens. It supports tool usage and reasoning, costs $0.40 per million input tokens and $0.40 per million output tokens, and its weights are not open. The model was released on 2025‑03‑16.

  • tool use
  • structured outputs
  • reasoning

At a glance

providernvidia
releasedMar 16, 2025
context131K
input / M$0.40

Specifications

What the API gives you

Context window
131K tokens
Max output
16K tokens
Input price
$0.40 / 1M tokens
Output price
$0.40 / 1M tokens
Cache read
Modalities
text → text
Knowledge cutoff
Mar 31, 2024
Released
Mar 16, 2025

Recommendation coverage

Where this model appears

Not currently included in a published recommendation. This means the available task-specific evidence did not support a ranking yet—not that the model was reviewed and rejected.

Benchmarks

No independent benchmark results are linked to this model yet. We keep it visible while withholding quality claims until usable results arrive.

Open-weight model on Hugging Face.

Catalog data via OpenRouter and models.dev. Last checked within the past 6 hours.