DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Model details →Best AI models for RAG
Good RAG models have long context (to fit retrieved chunks), low hallucination, and tool use for query refinement.
81Fit score
81Fit score
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Model details →81Fit score
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Model details →Top 25 models for retrieval-augmented generation
| # | Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|---|
| 1 | DeepSeek: DeepSeek V4 Flash 0423 | DeepSeek | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 2 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 1.3M Ultra context (1M+) | $0.04 | Budget |
| 3 | DeepSeek: DeepSeek V4 Flash 0731 (free) | DeepSeek | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 4 | Google: Gemini 2.0 Flash Lite | 1.0M Ultra context (1M+) | $0.07 | Budget | |
| 5 | Google: Gemini 2.5 Flash Lite (batch) | 1.0M Ultra context (1M+) | $0.05 | Budget | |
| 6 | MiniMax: MiniMax M3 (free) | Minimax | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 7 | OpenAI: GPT-4.1 Nano (batch) | OpenAI | 1.0M Ultra context (1M+) | $0.05 | Budget |
| 8 | OpenAI: GPT-6 Luna (batch) | OpenAI | 1.1M Ultra context (1M+) | $0.05 | Budget |
| 9 | OpenAI: GPT-6 Luna Pro (batch) | OpenAI | 1.1M Ultra context (1M+) | $0.05 | Budget |
| 10 | Auto Router | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 11 | Auto Router (Beta) | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 12 | Qwen: Qwen3.5-Flash | Qwen | 1M Ultra context (1M+) | $0.07 | Budget |
| 13 | Qwen: Qwen3.6 Plus Preview (free) | Qwen | 1M Ultra context (1M+) | $0.00 | Budget |
| 14 | Qwen: Qwen3.7 Flash | Qwen | 1M Ultra context (1M+) | $0.03 | Budget |
| 15 | Ox Alpha | Stealth | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 16 | Z.ai: GLM 5.3 Flash (batch) | Z Ai | 1.0M Ultra context (1M+) | $0.06 | Budget |
| 17 | ByteDance Seed: Seed 1.6 Flash | Bytedance Seed | 262K Long context (128K+) | $0.07 | Budget |
| 18 | DeepSeek: DeepSeek V4 Flash 0731 (batch) | DeepSeek | 1.0M Ultra context (1M+) | $0.11 | Budget |
| 19 | DeepSeek: DeepSeek V4 Flash Vision Exp | DeepSeek | 1.0M Ultra context (1M+) | $0.22 | Budget |
| 20 | DeepSeek: DeepSeek V4 Flash Vision Exp (batch) | DeepSeek | 1.0M Ultra context (1M+) | $0.11 | Budget |
| 21 | DeepSeek: DeepSeek V4 Pro 0423 | DeepSeek | 1.0M Ultra context (1M+) | $0.96 | Budget |
| 22 | DeepSeek: DeepSeek V4 Pro 0813 (batch) | DeepSeek | 1.0M Ultra context (1M+) | $0.66 | Budget |
| 23 | DeepSeek: DeepSeek V4.1 Flash | DeepSeek | 1.0M Ultra context (1M+) | $0.10 | Budget |
| 24 | DeepSeek: DeepSeek V4.1 Flash (batch) | DeepSeek | 1.0M Ultra context (1M+) | $0.11 | Budget |
| 25 | Dots Studio: Dots3-Note Preview (free) | Dots Studio | 512K Long context (128K+) | $0.00 | Budget |
Frequently asked questions
How big should the context window be for RAG?
32K is a comfortable minimum for typical knowledge-base RAG. 128K+ lets you include large source documents without summarization. 1M+ enables 'cache the whole codebase' patterns.
Related use cases
AI models for codingAI reasoning modelsAI vision modelslong-context AI models (128K+)AI models for function callingAI models for structured / JSON outputsAI models for agentsfree AI modelsCheapest AI models per tokenmultilingual AI modelsembedding modelssmall / on-device AI modelsAI image generation modelsAI voice / audio modelsAI models for mathAI models for writingopen-source AI modelsAI models for summarizationAI models for translationAI models for enterprise