Use case

Best AI voice / audio models

Audio models handle either transcription, audio reasoning, or speech synthesis. Some support real-time streaming.

Top 25 models for speech-to-text

#ModelProviderContextInput price / 1MTier
1Google: Gemini 2.0 FlashGoogle1.0M
Ultra context (1M+)
$0.10Budget
2Google: Gemini 2.0 Flash LiteGoogle1.0M
Ultra context (1M+)
$0.07Budget
3Google: Gemini 2.5 FlashGoogle1.0M
Ultra context (1M+)
$0.30Budget
4Google: Gemini 2.5 Flash (batch)Google1.0M
Ultra context (1M+)
$0.15Budget
5Google: Gemini 2.5 Flash LiteGoogle1.0M
Ultra context (1M+)
$0.10Budget
6Google: Gemini 2.5 Flash Lite (batch)Google1.0M
Ultra context (1M+)
$0.05Budget
7Google: Gemini 2.5 Flash Lite Preview 09-2025Google1.0M
Ultra context (1M+)
$0.10Budget
8Google: Gemini 2.5 ProGoogle1.0M
Ultra context (1M+)
$1.25Standard
9Google: Gemini 2.5 Pro (batch)Google1.0M
Ultra context (1M+)
$0.63Budget
10Google: Gemini 2.5 Pro Preview 05-06Google1.0M
Ultra context (1M+)
$1.25Standard
11Google: Gemini 2.5 Pro Preview 06-05Google1.0M
Ultra context (1M+)
$1.25Standard
12Google: Gemini 3 Flash PreviewGoogle1.0M
Ultra context (1M+)
$0.50Budget
13Google: Gemini 3 Flash Preview (batch)Google1.0M
Ultra context (1M+)
$0.25Budget
14Google: Gemini 3 Pro PreviewGoogle1.0M
Ultra context (1M+)
$2.00Standard
15Google: Gemini 3.1 Flash LiteGoogle1.0M
Ultra context (1M+)
$0.25Budget
16Google: Gemini 3.1 Flash Lite (batch)Google1.0M
Ultra context (1M+)
$0.13Budget
17Google: Gemini 3.1 Flash Lite PreviewGoogle1.0M
Ultra context (1M+)
$0.25Budget
18Google: Gemini 3.1 Pro PreviewGoogle1.0M
Ultra context (1M+)
$2.00Standard
19Google: Gemini 3.1 Pro Preview (batch)Google1.0M
Ultra context (1M+)
$1.00Standard
20Google: Gemini 3.1 Pro Preview Custom ToolsGoogle1.0M
Ultra context (1M+)
$2.00Standard
21Google: Gemini 3.5 FlashGoogle1.0M
Ultra context (1M+)
$1.50Standard
22Google: Gemini 3.5 Flash (batch)Google1.0M
Ultra context (1M+)
$0.75Budget
23Google: Gemini 3.5 Flash LiteGoogle1.0M
Ultra context (1M+)
$0.30Budget
24Google: Gemini 3.5 Flash Lite (batch)Google1.0M
Ultra context (1M+)
$0.15Budget
25Google: Gemini 3.6 FlashGoogle1.0M
Ultra context (1M+)
$0.75Budget

Frequently asked questions

Which model is best for transcription?

GPT-4o audio variants, Gemini audio, and Whisper-class models lead. Whisper is open-weight and self-hostable.

Related use cases