Inception: Mercury✓ Catalog verified
Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots. Read more in the [blog post] (https://www.inceptionlabs.ai/blog/introducing-mercury) here.
Overview
Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots. Read more in the [blog post] (https://www.inceptionlabs.ai/blog/introducing-mercury) here.
Access: available through the official Inception API. Context window: 128K.
Specifications
| API identifier | inception/mercury |
| Provider | Inception |
| Model type | Text LLM |
| Context window | 128K catalog |
| Input modalities | text |
| Output modalities | text |
| Released | Jun 26, 2025 |
| Moderated | No |
| Architecture modality | text->text |
| Scheduled expiry | 2026-04-15 |
API defaults
{
"top_p": null,
"temperature": 0,
"frequency_penalty": null
}Pricing
Live pricing components for inception/mercury as published in the catalog.
| Pricing component | Raw unit price | Normalized | Unit |
|---|---|---|---|
| Prompt tokens | $2.5e-7 | $0.25 / 1M | Per input token |
| Completion tokens | $7.5e-7 | $0.75 / 1M | Per output token |
| Request fee | N/A | N/A | Per request |
| Image fee | N/A | N/A | Per image unit |
| Web search fee | N/A | N/A | Per search request |
Cost calculator
Estimate monthly spend from your own token volumes.
Estimated Total Cost
$0.28Calculated from current list pricing in the ModelsAtlas catalog.Capabilities
| Capability | Status | What it means |
|---|---|---|
| Visual Understanding | Not advertised | Image and document analysis support |
| Audio Processing | Not advertised | Speech and voice aligned flows |
| Tool Calling | ✓ Supported | Supports tools / function calling |
| Self-Hosting | Not advertised | Deploy outside managed APIs |
Capability signals
Directional signals derived from model metadata and capability tags — not official benchmark submissions.
Quick start
Call Inception: Mercury through an OpenAI-compatible client.
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key="YOUR_API_KEY"
)
response = client.chat.completions.create(
model="inception/mercury",
messages=[{"role": "user", "content": "Explain quantum physics."}]
)
print(response.choices[0].message.content)Alternatives
Closest models by context window from a different provider.
| Model | Provider | Context | Input /1M | Actions |
|---|---|---|---|---|
| AllenAI: Olmo 2 32B Instruct | Allenai | 128K | $0.05 | View → |
| Amazon: Nova Micro 1.0 | Amazon | 128K | $0.04 | View → |
Sources & attribution
Pricing and metadata are maintained in the ModelsAtlas catalog.
Frequently asked questions
How much does Inception: Mercury cost?
$0.25 per 1M input tokens and $0.75 per 1M output tokens on the official Inception API.
What is the context window of Inception: Mercury?
Inception: Mercury supports up to 128K tokens of context.
Does Inception: Mercury support tool / function calling?
Yes, Inception: Mercury supports tool / function calling — you can register tools and the model will emit structured tool calls.
How do I access Inception: Mercury?
Use the official Inception API with the model id `inception/mercury`. See the Quick start section above for code examples.