Atlas / Models / Nvidia / NVIDIA: Llama 3.1 Nemotron 70B Instruct

NVIDIA: Llama 3.1 Nemotron 70B Instruct✓ Catalog verified

nvidia/llama-3.1-nemotron-70b-instruct

NVIDIA's Llama 3.1 Nemotron 70B is a language model designed for generating precise and useful responses. Leveraging Llama 3.1 70B architecture and Reinforcement Learning from Human Feedback (RLHF), it excels in automatic alignment benchmarks. This model is tailored for applications requiring high accuracy in helpfulness and response generation, suitable for diverse user queries across multiple domains. Usage of this model is subject to Meta's Acceptable Use Policy.

Input price
$1.20 /1M
Output price
$1.20 /1M
Context
131K
Modalities
text
Released
Oct 15, 2024
Tool calling
✓ Yes
Atlas signal
81/100
01

Overview

NVIDIA's Llama 3.1 Nemotron 70B is a language model designed for generating precise and useful responses. Leveraging Llama 3.1 70B architecture and Reinforcement Learning from Human Feedback (RLHF), it excels in automatic alignment benchmarks. This model is tailored for applications requiring high accuracy in helpfulness and response generation, suitable for diverse user queries across multiple domains. Usage of this model is subject to Meta's Acceptable Use Policy.

Access: available through the official Nvidia API. Context window: 131K.

Pricing and metadata from the ModelsAtlas catalog.Last refreshed Aug 5, 2026
02

Specifications

API identifiernvidia/llama-3.1-nemotron-70b-instruct
ProviderNvidia
Model typeText LLM
Context window131K catalog
Input modalitiestext
Output modalitiestext
ReleasedOct 15, 2024
TokenizerLlama3
Instruction formatllama3
ModeratedNo
Architecture modalitytext->text
Supported parameters
frequency_penaltymax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_p
Source: provider documentation + catalog feedMethodology →
03

Pricing

Live pricing components for nvidia/llama-3.1-nemotron-70b-instruct as published in the catalog.

$1.20
Input /1M
$1.20
Output /1M
Per request
Per image
Pricing componentRaw unit priceNormalizedUnit
Prompt tokens$0.0000012$1.20 / 1MPer input token
Completion tokens$0.0000012$1.20 / 1MPer output token
Request feeN/AN/APer request
Image feeN/AN/APer image unit
Web search feeN/AN/APer search request
Source: catalog pricing feedCompare all pricing →Cheapest models →
04

Cost calculator

Estimate monthly spend from your own token volumes.

Scale / Volume

Estimated Total Cost

$0.84Calculated from current list pricing in the ModelsAtlas catalog.
05

Capabilities

CapabilityStatusWhat it means
Visual UnderstandingNot advertisedImage and document analysis support
Audio ProcessingNot advertisedSpeech and voice aligned flows
Tool Calling✓ SupportedSupports tools / function calling
Self-HostingNot advertisedDeploy outside managed APIs
Derived from catalog capability tags and supported parameters
06

Capability signals

Directional signals derived from model metadata and capability tags — not official benchmark submissions.

NVIDIA: Llama 3.1 Nemotron 70B Instruct
MMLU Signal (Reasoning)
79signal
Coding Signal (HumanEval proxy)
88signal
Math Signal (GSM8K proxy)
79signal
Science Signal (GPQA proxy)
79signal
Estimated from metadata — not an official benchmark runAll benchmarks →Methodology →
07

Quick start

Call NVIDIA: Llama 3.1 Nemotron 70B Instruct through an OpenAI-compatible client.

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_API_KEY"
)

response = client.chat.completions.create(
    model="nvidia/llama-3.1-nemotron-70b-instruct",
    messages=[{"role": "user", "content": "Explain quantum physics."}]
)

print(response.choices[0].message.content)
08

Alternatives

Closest models by context window from a different provider.

09

Sources & attribution

Pricing and metadata are maintained in the ModelsAtlas catalog.

Capability bars use metadata tags — directional estimates onlyHow we source and verify data →
10

Frequently asked questions

How much does NVIDIA: Llama 3.1 Nemotron 70B Instruct cost?

$1.20 per 1M input tokens and $1.20 per 1M output tokens on the official Nvidia API.

What is the context window of NVIDIA: Llama 3.1 Nemotron 70B Instruct?

NVIDIA: Llama 3.1 Nemotron 70B Instruct supports up to 131K tokens of context.

Does NVIDIA: Llama 3.1 Nemotron 70B Instruct support tool / function calling?

Yes, NVIDIA: Llama 3.1 Nemotron 70B Instruct supports tool / function calling — you can register tools and the model will emit structured tool calls.

How do I access NVIDIA: Llama 3.1 Nemotron 70B Instruct?

Use the official Nvidia API with the model id `nvidia/llama-3.1-nemotron-70b-instruct`. See the Quick start section above for code examples.