This model always redirects to the latest model in the DeepSeek V4 Flash family.
Model details →Best AI models for summarization
Summarization is mostly a long-context + instruction-following problem. The winners overlap with the long-context list.
69Fit score
69Fit score
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Model details →69Fit score
Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/g…
Model details →Top 25 models for document summarization
| # | Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash Latest | ~deepseek | 1.0M Ultra context (1M+) | $0.08 | Budget |
| 2 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 3 | Google: Gemini 2.0 Flash Lite | 1.0M Ultra context (1M+) | $0.07 | Budget | |
| 4 | Google: Gemini 2.5 Flash Lite (batch) | 1.0M Ultra context (1M+) | $0.05 | Budget | |
| 5 | Google: Lyria 3 Clip Preview | 1.0M Ultra context (1M+) | $0.00 | Budget | |
| 6 | Google: Lyria 3 Pro Preview | 1.0M Ultra context (1M+) | $0.00 | Budget | |
| 7 | NVIDIA: Nemotron 3 Ultra (free) | Nvidia | 1M Ultra context (1M+) | $0.00 | Budget |
| 8 | OpenAI: GPT-4.1 Nano (batch) | OpenAI | 1.0M Ultra context (1M+) | $0.05 | Budget |
| 9 | Auto Router | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 10 | Auto Router (Beta) | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 11 | OpenRouter: Fusion | Openrouter | 1M Ultra context (1M+) | $-1000000.00 | Budget |
| 12 | Pareto Code Router | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 13 | Poolside: Laguna S 2.1 | Poolside | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 14 | Qwen: Qwen3.5-Flash | Qwen | 1M Ultra context (1M+) | $0.07 | Budget |
| 15 | Qwen: Qwen3.6 Plus Preview (free) | Qwen | 1M Ultra context (1M+) | $0.00 | Budget |
| 16 | Qwen: Qwen3.7 Flash | Qwen | 1M Ultra context (1M+) | $0.03 | Budget |
| 17 | Z.ai: GLM 5.2 | Z Ai | 1.0M Ultra context (1M+) | $0.10 | Budget |
| 18 | Amazon: Nova 2 Lite | Amazon | 1M Ultra context (1M+) | $0.30 | Budget |
| 19 | Amazon: Nova Lite 1.0 | Amazon | 300K Long context (128K+) | $0.06 | Budget |
| 20 | ByteDance Seed: Seed 1.6 Flash | Bytedance Seed | 262K Long context (128K+) | $0.07 | Budget |
| 21 | Cohere: North Mini Code (free) | Cohere | 256K Long context (128K+) | $0.00 | Budget |
| 22 | DeepSeek: DeepSeek V4 Flash 0423 | DeepSeek | 1.0M Ultra context (1M+) | $0.14 | Budget |
| 23 | DeepSeek: DeepSeek V4 Pro | DeepSeek | 1.0M Ultra context (1M+) | $0.43 | Budget |
| 24 | Google: Gemini 2.0 Flash | 1.0M Ultra context (1M+) | $0.10 | Budget | |
| 25 | Google: Gemini 2.5 Flash | 1.0M Ultra context (1M+) | $0.30 | Budget |
Frequently asked questions
Should I use a small or large model for summarization?
Smaller models are fine for simple TL;DR. For domain-specific or accuracy-sensitive summaries (legal, medical, technical), use frontier models.
Related use cases
AI models for codingAI reasoning modelsAI vision modelslong-context AI models (128K+)AI models for function callingAI models for structured / JSON outputsAI models for agentsfree AI modelsCheapest AI models per tokenmultilingual AI modelsembedding modelssmall / on-device AI modelsAI image generation modelsAI voice / audio modelsAI models for mathAI models for writingopen-source AI modelsAI models for RAGAI models for translationAI models for enterprise