OpenAI: GPT Audio✓ Catalog verified
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
Overview
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
Access: available through the official OpenAI API. Context window: 128K.
Specifications
| API identifier | openai/gpt-audio |
| Provider | OpenAI |
| Model type | Multimodal LLM |
| Context window | 128K catalog |
| Input modalities | text · audio |
| Output modalities | text · audio |
| Released | Jan 19, 2026 |
| Tokenizer | GPT |
| Moderated | Yes |
| Architecture modality | text+audio->text+audio |
API defaults
{
"top_k": null,
"top_p": null,
"temperature": null,
"presence_penalty": null,
"frequency_penalty": null,
"repetition_penalty": null
}Pricing
Live pricing components for openai/gpt-audio as published in the catalog.
| Pricing component | Raw unit price | Normalized | Unit |
|---|---|---|---|
| Prompt tokens | $0.0000025 | $2.50 / 1M | Per input token |
| Completion tokens | $0.00001 | $10.00 / 1M | Per output token |
| Request fee | N/A | N/A | Per request |
| Image fee | N/A | N/A | Per image unit |
| Web search fee | N/A | N/A | Per search request |
Cost calculator
Estimate monthly spend from your own token volumes.
Estimated Total Cost
$3.25Calculated from current list pricing in the ModelsAtlas catalog.Capabilities
| Capability | Status | What it means |
|---|---|---|
| Visual Understanding | Not advertised | Image and document analysis support |
| Audio Processing | ✓ Supported | Speech and voice aligned flows |
| Tool Calling | ✓ Supported | Supports tools / function calling |
| Self-Hosting | Not advertised | Deploy outside managed APIs |
Capability signals
Directional signals derived from model metadata and capability tags — not official benchmark submissions.
Quick start
Call OpenAI: GPT Audio through an OpenAI-compatible client.
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key="YOUR_API_KEY"
)
response = client.chat.completions.create(
model="openai/gpt-audio",
messages=[{"role": "user", "content": "Explain quantum physics."}]
)
print(response.choices[0].message.content)Alternatives
Closest models by context window from a different provider.
| Model | Provider | Context | Input /1M | Actions |
|---|---|---|---|---|
| AllenAI: Olmo 2 32B Instruct | Allenai | 128K | $0.05 | View → |
| Amazon: Nova Micro 1.0 | Amazon | 128K | $0.04 | View → |
Sources & attribution
Pricing and metadata are maintained in the ModelsAtlas catalog.
Frequently asked questions
How much does OpenAI: GPT Audio cost?
$2.50 per 1M input tokens and $10.00 per 1M output tokens on the official OpenAI API.
What is the context window of OpenAI: GPT Audio?
OpenAI: GPT Audio supports up to 128K tokens of context.
Does OpenAI: GPT Audio support tool / function calling?
Yes, OpenAI: GPT Audio supports tool / function calling — you can register tools and the model will emit structured tool calls.
How do I access OpenAI: GPT Audio?
Use the official OpenAI API with the model id `openai/gpt-audio`. See the Quick start section above for code examples.