Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning to reach state-of-the-art performance on m…
Model details →Deep Cogito: Cogito v2.1 671B vs NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 wins on 2 of 6 axes — pricing and capability skew in its favour for most workloads.
Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has bee…
Model details →Side-by-side comparison
| Capability | Deep Cogito: Cogito v2.1 671B | NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 | Winner |
|---|---|---|---|
| Context window Maximum number of input tokens the model can attend to in a single request. | 128K | 131K | 🏆 NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 |
| Input price (per 1M) Cost per million input tokens billed by the provider. | $1.25 | $0.60 | 🏆 NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 |
| Output price (per 1M) Cost per million output tokens billed by the provider. | $1.25 | $1.80 | 🏆 Deep Cogito: Cogito v2.1 671B |
| Tool / function calling First-class support for emitting structured tool calls. | — | — | Tie |
| Vision input Accepts image inputs alongside text. | — | — | Tie |
| Reasoning mode Internal chain-of-thought / extended-thinking support. | Yes | Yes | Tie |
Frequently asked questions
Is Deep Cogito: Cogito v2.1 671B better than NVIDIA: Llama 3.1 Nemotron Ultra 253B v1?
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 wins on 2 of 6 axes — pricing and capability skew in its favour for most workloads.
What's the price difference between Deep Cogito: Cogito v2.1 671B and NVIDIA: Llama 3.1 Nemotron Ultra 253B v1?
Input: $1.25 vs $0.60 per 1M tokens. Output: $1.25 vs $1.80 per 1M tokens.
What context windows do Deep Cogito: Cogito v2.1 671B and NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 support?
Deep Cogito: Cogito v2.1 671B supports up to 128K tokens. NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 supports up to 131K tokens.
Do both Deep Cogito: Cogito v2.1 671B and NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 support tool calling?
Deep Cogito: Cogito v2.1 671B: not advertised. NVIDIA: Llama 3.1 Nemotron Ultra 253B v1: not advertised.