DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Model details →Best AI models for summarization
Summarization is mostly a long-context + instruction-following problem. The winners overlap with the long-context list.
69Fit score
69Fit score
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Model details →69Fit score
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Model details →Top 25 models for document summarization
| # | Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|---|
| 1 | DeepSeek: DeepSeek V4 Flash 0423 | DeepSeek | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 2 | DeepSeek: DeepSeek V4 Flash 0731 | DeepSeek | 1.3M Ultra context (1M+) | $0.04 | Budget |
| 3 | DeepSeek: DeepSeek V4 Flash 0731 (free) | DeepSeek | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 4 | Google: Gemini 2.0 Flash Lite | 1.0M Ultra context (1M+) | $0.07 | Budget | |
| 5 | Google: Gemini 2.5 Flash Lite (batch) | 1.0M Ultra context (1M+) | $0.05 | Budget | |
| 6 | Google: Lyria 3 Clip Preview | 1.0M Ultra context (1M+) | $0.00 | Budget | |
| 7 | Google: Lyria 3 Pro Preview | 1.0M Ultra context (1M+) | $0.00 | Budget | |
| 8 | MiniMax: MiniMax M3 (free) | Minimax | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 9 | NVIDIA: Nemotron 3 Ultra (free) | Nvidia | 1M Ultra context (1M+) | $0.00 | Budget |
| 10 | NVIDIA: Nemotron 3.5 Lightning (free) | Nvidia | 1M Ultra context (1M+) | $0.00 | Budget |
| 11 | OpenAI: GPT-4.1 Nano (batch) | OpenAI | 1.0M Ultra context (1M+) | $0.05 | Budget |
| 12 | OpenAI: GPT-6 Luna (batch) | OpenAI | 1.1M Ultra context (1M+) | $0.05 | Budget |
| 13 | OpenAI: GPT-6 Luna Pro (batch) | OpenAI | 1.1M Ultra context (1M+) | $0.05 | Budget |
| 14 | Auto Router | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 15 | Auto Router (Beta) | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 16 | OpenRouter: Fusion | Openrouter | 1M Ultra context (1M+) | $-1000000.00 | Budget |
| 17 | Pareto Code Router | Openrouter | 2M Ultra context (1M+) | $-1000000.00 | Budget |
| 18 | Poolside: Laguna S 2.1 | Poolside | 1.0M Ultra context (1M+) | $0.09 | Budget |
| 19 | Qwen: Qwen3.5-Flash | Qwen | 1M Ultra context (1M+) | $0.07 | Budget |
| 20 | Qwen: Qwen3.6 Plus Preview (free) | Qwen | 1M Ultra context (1M+) | $0.00 | Budget |
| 21 | Qwen: Qwen3.7 Flash | Qwen | 1M Ultra context (1M+) | $0.03 | Budget |
| 22 | Ox Alpha | Stealth | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 23 | Thinking Machines: Inkling (free) | Thinkingmachines | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 24 | Thinking Machines: Inkling Small (free) | Thinkingmachines | 1.0M Ultra context (1M+) | $0.00 | Budget |
| 25 | Z.ai: GLM 5.3 Flash (batch) | Z Ai | 1.0M Ultra context (1M+) | $0.06 | Budget |
Frequently asked questions
Should I use a small or large model for summarization?
Smaller models are fine for simple TL;DR. For domain-specific or accuracy-sensitive summaries (legal, medical, technical), use frontier models.
Related use cases
AI models for codingAI reasoning modelsAI vision modelslong-context AI models (128K+)AI models for function callingAI models for structured / JSON outputsAI models for agentsfree AI modelsCheapest AI models per tokenmultilingual AI modelsembedding modelssmall / on-device AI modelsAI image generation modelsAI voice / audio modelsAI models for mathAI models for writingopen-source AI modelsAI models for RAGAI models for translationAI models for enterprise