RAG-ready models
Good RAG models have long context (to fit retrieved chunks), low hallucination, and tool use for query refinement.
50 models in this category. See also the “Best AI models for RAG” ranking with editorial picks and FAQ.
| Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|
| DeepSeek V4 Flash Latest ~deepseek/deepseek-v4-flash-latest | ~deepseek | 1.0M Ultra context (1M+) | $0.08 | Budget |
| DeepSeek: DeepSeek V4 Flash 0731 deepseek/deepseek-v4-flash-0731 | DeepSeek | 1.0M Ultra context (1M+) | $0.09 | Budget |
| Google: Gemini 2.0 Flash Lite google/gemini-2.0-flash-lite-001 | 1.0M Ultra context (1M+) | $0.07 | Budget | |
| Google: Gemini 2.5 Flash Lite (batch) google/gemini-2.5-flash-lite:batch | 1.0M Ultra context (1M+) | $0.05 | Budget | |
| OpenAI: GPT-4.1 Nano (batch) openai/gpt-4.1-nano:batch | OpenAI | 1.0M Ultra context (1M+) | $0.05 | Budget |
| Auto Router openrouter/auto | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| Auto Router (Beta) openrouter/auto-beta | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| Qwen: Qwen3.5-Flash qwen/qwen3.5-flash-02-23 | Qwen | 1M Ultra context (1M+) | $0.07 | Budget |
| Qwen: Qwen3.6 Plus Preview (free) qwen/qwen3.6-plus-preview:free | Qwen | 1M Ultra context (1M+) | $0.00 | Budget |
| Qwen: Qwen3.7 Flash qwen/qwen3.7-flash | Qwen | 1M Ultra context (1M+) | $0.03 | Budget |
| Z.ai: GLM 5.2 z-ai/glm-5.2 | Z Ai | 1.0M Ultra context (1M+) | $0.10 | Budget |
| ByteDance Seed: Seed 1.6 Flash bytedance-seed/seed-1.6-flash | Bytedance Seed | 262K Long context (128K+) | $0.07 | Budget |
| DeepSeek: DeepSeek V4 Flash 0423 deepseek/deepseek-v4-flash | DeepSeek | 1.0M Ultra context (1M+) | $0.14 | Budget |
| DeepSeek: DeepSeek V4 Pro deepseek/deepseek-v4-pro | DeepSeek | 1.0M Ultra context (1M+) | $0.43 | Budget |
| Google: Gemini 2.0 Flash google/gemini-2.0-flash-001 | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| Google: Gemini 2.5 Flash google/gemini-2.5-flash | 1.0M Ultra context (1M+) | $0.30 | Budget | |
| Google: Gemini 2.5 Flash (batch) google/gemini-2.5-flash:batch | 1.0M Ultra context (1M+) | $0.15 | Budget | |
| Google: Gemini 2.5 Flash Lite google/gemini-2.5-flash-lite | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| Google: Gemini 2.5 Flash Lite Preview 09-2025 google/gemini-2.5-flash-lite-preview-09-2025 | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| Google: Gemini 2.5 Pro (batch) google/gemini-2.5-pro:batch | 1.0M Ultra context (1M+) | $0.63 | Budget | |
| Google: Gemini 3 Flash Preview google/gemini-3-flash-preview | 1.0M Ultra context (1M+) | $0.50 | Budget | |
| Google: Gemini 3 Flash Preview (batch) google/gemini-3-flash-preview:batch | 1.0M Ultra context (1M+) | $0.25 | Budget | |
| Google: Gemini 3.1 Flash Lite google/gemini-3.1-flash-lite | 1.0M Ultra context (1M+) | $0.25 | Budget | |
| Google: Gemini 3.1 Flash Lite (batch) google/gemini-3.1-flash-lite:batch | 1.0M Ultra context (1M+) | $0.13 | Budget | |
| Google: Gemini 3.1 Flash Lite Preview google/gemini-3.1-flash-lite-preview | 1.0M Ultra context (1M+) | $0.25 | Budget | |
| Google: Gemini 3.5 Flash (batch) google/gemini-3.5-flash:batch | 1.0M Ultra context (1M+) | $0.75 | Budget | |
| Google: Gemini 3.5 Flash Lite google/gemini-3.5-flash-lite | 1.0M Ultra context (1M+) | $0.30 | Budget | |
| Google: Gemini 3.5 Flash Lite (batch) google/gemini-3.5-flash-lite:batch | 1.0M Ultra context (1M+) | $0.15 | Budget | |
| Google: Gemini 3.6 Flash (batch) google/gemini-3.6-flash:batch | 1.0M Ultra context (1M+) | $0.75 | Budget | |
| Google: Gemma 3 27B google/gemma-3-27b-it | 262K Long context (128K+) | $0.08 | Budget | |
| Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 262K Long context (128K+) | $0.07 | Budget | |
| Google: Gemma 4 26B A4B (free) google/gemma-4-26b-a4b-it:free | 262K Long context (128K+) | $0.00 | Budget | |
| Google: Gemma 4 31B (free) google/gemma-4-31b-it:free | 262K Long context (128K+) | $0.00 | Budget | |
| Qwen: Qwen3 30B A3B Instruct 2507 qwen/qwen3-30b-a3b-instruct-2507 | Qwen | 262K Long context (128K+) | $0.05 | Budget |
| inclusionAI: Ling-2.6-1T inclusionai/ling-2.6-1t | Inclusionai | 262K Long context (128K+) | $0.07 | Budget |
| inclusionAI: Ling-2.6-flash inclusionai/ling-2.6-flash | Inclusionai | 262K Long context (128K+) | $0.01 | Budget |
| inclusionAI: Ring-2.6-1T inclusionai/ring-2.6-1t | Inclusionai | 262K Long context (128K+) | $0.07 | Budget |
| Meta: Llama 4 Maverick meta-llama/llama-4-maverick | Meta | 1.0M Ultra context (1M+) | $0.20 | Budget |
| Meta: Llama 4 Scout meta-llama/llama-4-scout | Meta | 1.3M Ultra context (1M+) | $0.10 | Budget |
| MiniMax: MiniMax M3 minimax/minimax-m3 | Minimax | 1.0M Ultra context (1M+) | $0.30 | Budget |
| Mistral: Mistral Small 3.2 24B mistralai/mistral-small-3.2-24b-instruct | Mistral AI | 256K Long context (128K+) | $0.09 | Budget |
| Nex AGI: Nex-N2-Mini nex-agi/nex-n2-mini | Nex Agi | 262K Long context (128K+) | $0.03 | Budget |
| NVIDIA: Nemotron 3 Nano 30B A3B nvidia/nemotron-3-nano-30b-a3b | Nvidia | 262K Long context (128K+) | $0.05 | Budget |
| NVIDIA: Nemotron 3 Super nvidia/nemotron-3-super-120b-a12b | Nvidia | 1M Ultra context (1M+) | $0.30 | Budget |
| NVIDIA: Nemotron 3 Super (free) nvidia/nemotron-3-super-120b-a12b:free | Nvidia | 262K Long context (128K+) | $0.00 | Budget |
| OpenAI: GPT-4.1 Mini openai/gpt-4.1-mini | OpenAI | 1.0M Ultra context (1M+) | $0.40 | Budget |
| OpenAI: GPT-4.1 Mini (batch) openai/gpt-4.1-mini:batch | OpenAI | 1.0M Ultra context (1M+) | $0.20 | Budget |
| OpenAI: GPT-4.1 Nano openai/gpt-4.1-nano | OpenAI | 1.0M Ultra context (1M+) | $0.10 | Budget |
| OpenAI: GPT-5 Nano openai/gpt-5-nano | OpenAI | 400K Long context (128K+) | $0.05 | Budget |
| OpenAI: GPT-5 Nano (batch) openai/gpt-5-nano:batch | OpenAI | 400K Long context (128K+) | $0.03 | Budget |