Atlas / Models / Nvidia / NVIDIA: Llama 3.1 Nemotron Ultra 253B v1

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1✓ Catalog verified

nvidia/llama-3.1-nemotron-ultra-253b-v1

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural Architecture Search (NAS), resulting in enhanced efficiency, reduced memory usage, and improved inference latency. The model supports a context length of up to 128K tokens and can operate efficiently on an 8x NVIDIA H100 node. Note: you must include `detailed thinking on` in the system prompt to enable reasoning. Please see Usage Recommendations for more.

Input price
$0.60 /1M
Output price
$1.80 /1M
Context
131K
Modalities
text
Released
Apr 8, 2025
Tool calling
Atlas signal
85/100
01

Overview

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural Architecture Search (NAS), resulting in enhanced efficiency, reduced memory usage, and improved inference latency. The model supports a context length of up to 128K tokens and can operate efficiently on an 8x NVIDIA H100 node. Note: you must include `detailed thinking on` in the system prompt to enable reasoning. Please see Usage Recommendations for more.

Access: available through the official Nvidia API. Context window: 131K.

Pricing and metadata from the ModelsAtlas catalog.Last refreshed Aug 5, 2026
02

Specifications

API identifiernvidia/llama-3.1-nemotron-ultra-253b-v1
ProviderNvidia
Model typeText LLM
Context window131K catalog
Input modalitiestext
Output modalitiestext
ReleasedApr 8, 2025
TokenizerLlama3
ModeratedNo
Architecture modalitytext->text
Supported parameters
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatstructured_outputstemperaturetop_ktop_p
Source: provider documentation + catalog feedMethodology →

API defaults

Default parameters
{
  "top_k": null,
  "top_p": null,
  "temperature": null,
  "presence_penalty": null,
  "frequency_penalty": null,
  "repetition_penalty": null
}
03

Pricing

Live pricing components for nvidia/llama-3.1-nemotron-ultra-253b-v1 as published in the catalog.

$0.60
Input /1M
$1.80
Output /1M
Per request
Per image
Pricing componentRaw unit priceNormalizedUnit
Prompt tokens$6e-7$0.60 / 1MPer input token
Completion tokens$0.0000018$1.80 / 1MPer output token
Request feeN/AN/APer request
Image feeN/AN/APer image unit
Web search feeN/AN/APer search request
Source: catalog pricing feedCompare all pricing →Cheapest models →
04

Cost calculator

Estimate monthly spend from your own token volumes.

Scale / Volume

Estimated Total Cost

$0.66Calculated from current list pricing in the ModelsAtlas catalog.
05

Capabilities

CapabilityStatusWhat it means
Visual UnderstandingNot advertisedImage and document analysis support
Audio ProcessingNot advertisedSpeech and voice aligned flows
Tool CallingNot advertisedSupports tools / function calling
Self-HostingNot advertisedDeploy outside managed APIs
Derived from catalog capability tags and supported parameters
06

Capability signals

Directional signals derived from model metadata and capability tags — not official benchmark submissions.

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
MMLU Signal (Reasoning)
97signal
Coding Signal (HumanEval proxy)
82signal
Math Signal (GSM8K proxy)
83signal
Science Signal (GPQA proxy)
79signal
Estimated from metadata — not an official benchmark runAll benchmarks →Methodology →
07

Quick start

Call NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 through an OpenAI-compatible client.

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_API_KEY"
)

response = client.chat.completions.create(
    model="nvidia/llama-3.1-nemotron-ultra-253b-v1",
    messages=[{"role": "user", "content": "Explain quantum physics."}]
)

print(response.choices[0].message.content)
08

Alternatives

Closest models by context window from a different provider.

09

Sources & attribution

Pricing and metadata are maintained in the ModelsAtlas catalog.

Capability bars use metadata tags — directional estimates onlyHow we source and verify data →
10

Frequently asked questions

How much does NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 cost?

$0.60 per 1M input tokens and $1.80 per 1M output tokens on the official Nvidia API.

What is the context window of NVIDIA: Llama 3.1 Nemotron Ultra 253B v1?

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 supports up to 131K tokens of context.

Does NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 support tool / function calling?

NVIDIA: Llama 3.1 Nemotron Ultra 253B v1 does not advertise first-class tool/function calling. Check vendor docs for the latest capabilities.

How do I access NVIDIA: Llama 3.1 Nemotron Ultra 253B v1?

Use the official Nvidia API with the model id `nvidia/llama-3.1-nemotron-ultra-253b-v1`. See the Quick start section above for code examples.