This model always redirects to the latest model in the DeepSeek V4 Flash family.
Model details →Best AI models for RAG
Good RAG models have long context (to fit retrieved chunks), low hallucination, and tool use for query refinement.
81Fit score
81Fit score
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Model details →81Fit score
Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/g…
Model details →Top 25 models for retrieval-augmented generation
| # | Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash Latest | ~deepseek | 1.0M Ultra context (1M+) | $0.08 | Budget |
| 2 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 3 | Google: Gemini 2.0 Flash Lite | 1.0M Ultra context (1M+) | $0.07 | Budget | |
| 4 | Google: Gemini 2.5 Flash Lite (batch) | 1.0M Ultra context (1M+) | $0.05 | Budget | |
| 5 | OpenAI: GPT-4.1 Nano (batch) | OpenAI | 1.0M Ultra context (1M+) | $0.05 | Budget |
| 6 | Auto Router | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 7 | Auto Router (Beta) | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 8 | Qwen: Qwen3.5-Flash | Qwen | 1M Ultra context (1M+) | $0.07 | Budget |
| 9 | Qwen: Qwen3.6 Plus Preview (free) | Qwen | 1M Ultra context (1M+) | $0.00 | Budget |
| 10 | Qwen: Qwen3.7 Flash | Qwen | 1M Ultra context (1M+) | $0.03 | Budget |
| 11 | Z.ai: GLM 5.2 | Z Ai | 1.0M Ultra context (1M+) | $0.10 | Budget |
| 12 | ByteDance Seed: Seed 1.6 Flash | Bytedance Seed | 262K Long context (128K+) | $0.07 | Budget |
| 13 | DeepSeek: DeepSeek V4 Flash 0423 | DeepSeek | 1.0M Ultra context (1M+) | $0.14 | Budget |
| 14 | DeepSeek: DeepSeek V4 Pro | DeepSeek | 1.0M Ultra context (1M+) | $0.43 | Budget |
| 15 | Google: Gemini 2.0 Flash | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| 16 | Google: Gemini 2.5 Flash | 1.0M Ultra context (1M+) | $0.30 | Budget | |
| 17 | Google: Gemini 2.5 Flash (batch) | 1.0M Ultra context (1M+) | $0.15 | Budget | |
| 18 | Google: Gemini 2.5 Flash Lite | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| 19 | Google: Gemini 2.5 Flash Lite Preview 09-2025 | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| 20 | Google: Gemini 2.5 Pro (batch) | 1.0M Ultra context (1M+) | $0.63 | Budget | |
| 21 | Google: Gemini 3 Flash Preview | 1.0M Ultra context (1M+) | $0.50 | Budget | |
| 22 | Google: Gemini 3 Flash Preview (batch) | 1.0M Ultra context (1M+) | $0.25 | Budget | |
| 23 | Google: Gemini 3.1 Flash Lite | 1.0M Ultra context (1M+) | $0.25 | Budget | |
| 24 | Google: Gemini 3.1 Flash Lite (batch) | 1.0M Ultra context (1M+) | $0.13 | Budget | |
| 25 | Google: Gemini 3.1 Flash Lite Preview | 1.0M Ultra context (1M+) | $0.25 | Budget |
Frequently asked questions
How big should the context window be for RAG?
32K is a comfortable minimum for typical knowledge-base RAG. 128K+ lets you include large source documents without summarization. 1M+ enables 'cache the whole codebase' patterns.
Related use cases
AI models for codingAI reasoning modelsAI vision modelslong-context AI models (128K+)AI models for function callingAI models for structured / JSON outputsAI models for agentsfree AI modelsCheapest AI models per tokenmultilingual AI modelsembedding modelssmall / on-device AI modelsAI image generation modelsAI voice / audio modelsAI models for mathAI models for writingopen-source AI modelsAI models for summarizationAI models for translationAI models for enterprise