GPT Audio
openai/gpt-audio
OpenAI's GPT Audio, released on January 20 2026, accepts text and audio inputs and can generate both text and audio outputs. It supports tool usage, does not support reasoning, and has a context length of 128 000 tokens. The model is not open weights, with pricing of $2.5 per million input tokens and $10 per million output tokens.
- tool use
- structured outputs
At a glance
Specifications
What the API gives you
- Context window
- 128K tokens
- Max output
- 16K tokens
- Input price
- $2.50 / 1M tokens
- Output price
- $10 / 1M tokens
- Cache read
- —
- Modalities
- text, audio → text, audio
- Knowledge cutoff
- —
- Released
- Jan 20, 2026
Recommendation coverage
Where this model appears
Not currently included in a published recommendation. This means the available task-specific evidence did not support a ranking yet—not that the model was reviewed and rejected.
Benchmarks
No independent benchmark results are linked to this model yet. We keep it visible while withholding quality claims until usable results arrive.
Catalog data via OpenRouter and models.dev. Last checked within the past 6 hours.