OpenAI: GPT-4o Audio✓ Catalog verified
The gpt-4o-audio-preview model adds support for audio inputs as prompts. This enhancement allows the model to detect nuances within audio recordings and add depth to generated user experiences. Audio outputs are currently not supported. Audio tokens are priced at $40 per million input and $80 per million output audio tokens.
Overview
The gpt-4o-audio-preview model adds support for audio inputs as prompts. This enhancement allows the model to detect nuances within audio recordings and add depth to generated user experiences. Audio outputs are currently not supported. Audio tokens are priced at $40 per million input and $80 per million output audio tokens.
Access: available through the official OpenAI API. Context window: 128K.
Specifications
| API identifier | openai/gpt-4o-audio-preview |
| Provider | OpenAI |
| Model type | Multimodal LLM |
| Context window | 128K catalog |
| Input modalities | audio · text |
| Output modalities | text · audio |
| Released | Aug 15, 2025 |
| Tokenizer | GPT |
| Moderated | Yes |
| Architecture modality | text+audio->text+audio |
API defaults
{
"top_p": null,
"temperature": null,
"frequency_penalty": null
}Pricing
Live pricing components for openai/gpt-4o-audio-preview as published in the catalog.
| Pricing component | Raw unit price | Normalized | Unit |
|---|---|---|---|
| Prompt tokens | $0.0000025 | $2.50 / 1M | Per input token |
| Completion tokens | $0.00001 | $10.00 / 1M | Per output token |
| Request fee | N/A | N/A | Per request |
| Image fee | N/A | N/A | Per image unit |
| Web search fee | N/A | N/A | Per search request |
Cost calculator
Estimate monthly spend from your own token volumes.
Estimated Total Cost
$3.25Calculated from current list pricing in the ModelsAtlas catalog.Capabilities
| Capability | Status | What it means |
|---|---|---|
| Visual Understanding | Not advertised | Image and document analysis support |
| Audio Processing | ✓ Supported | Speech and voice aligned flows |
| Tool Calling | ✓ Supported | Supports tools / function calling |
| Self-Hosting | Not advertised | Deploy outside managed APIs |
Capability signals
Directional signals derived from model metadata and capability tags — not official benchmark submissions.
Quick start
Call OpenAI: GPT-4o Audio through an OpenAI-compatible client.
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key="YOUR_API_KEY"
)
response = client.chat.completions.create(
model="openai/gpt-4o-audio-preview",
messages=[{"role": "user", "content": "Explain quantum physics."}]
)
print(response.choices[0].message.content)Alternatives
Closest models by context window from a different provider.
| Model | Provider | Context | Input /1M | Actions |
|---|---|---|---|---|
| AllenAI: Olmo 2 32B Instruct | Allenai | 128K | $0.05 | View → |
| Amazon: Nova Micro 1.0 | Amazon | 128K | $0.04 | View → |
Sources & attribution
Pricing and metadata are maintained in the ModelsAtlas catalog.
Frequently asked questions
How much does OpenAI: GPT-4o Audio cost?
$2.50 per 1M input tokens and $10.00 per 1M output tokens on the official OpenAI API.
What is the context window of OpenAI: GPT-4o Audio?
OpenAI: GPT-4o Audio supports up to 128K tokens of context.
Does OpenAI: GPT-4o Audio support tool / function calling?
Yes, OpenAI: GPT-4o Audio supports tool / function calling — you can register tools and the model will emit structured tool calls.
How do I access OpenAI: GPT-4o Audio?
Use the official OpenAI API with the model id `openai/gpt-4o-audio-preview`. See the Quick start section above for code examples.