Llama 3.2 11B Vision Instruct

meta-llama/llama-3.2-11b-vision-instruct

Meta's Llama 3.2 11B Vision Instruct is an 11‑billion‑parameter model that accepts text and image inputs and generates text output, with a context window of 131 072 tokens. It is released under open weights, costs $0.345 per 1 M input and output tokens, does not support tools or reasoning, and was released on 2024‑09‑25.

  • structured outputs
  • open-weight

At a glance

providermeta-llama
releasedSep 25, 2024
context131K
input / M$0.34

Specifications

What the API gives you

Context window
131K tokens
Max output
16K tokens
Input price
$0.34 / 1M tokens
Output price
$0.34 / 1M tokens
Cache read
Modalities
text, image → text
Knowledge cutoff
Dec 31, 2023
Released
Sep 25, 2024

Recommendation coverage

Where this model appears

Not currently included in a published recommendation. This means the available task-specific evidence did not support a ranking yet—not that the model was reviewed and rejected.

Benchmarks

No independent benchmark results are linked to this model yet. We keep it visible while withholding quality claims until usable results arrive.

Open-weight model on Hugging Face.

Catalog data via OpenRouter and models.dev. Last checked within the past 6 hours.