Nemotron Policy GeneratorSAFE
Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
Overview
Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
441e9a359902OBSERVED · 2026-10-09Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned | |
| codex | mentioned | |
| copilot | mentioned | |
| cursor | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: "nemotron-policy-generator" title: "Nemotron Policy Generator" version: "0.1.0" description: "Generates BYO custom safety policies for NVIDIA Nemotron content-safety guardrails — Nemotron-Content-Safety-Reasoning-4B (text) and multimodal Nemotron-3-Content-Safety. Produces a Markdown policy, JSON taxonomy, and drop-in inference prompts. Maps rough words or an existing policy to V2 categories, adding custom categories or topic-following rules." license: "Apache-2.0 AND CC-BY-4.0" compatibility: "nvidia/Nemotron-Content-Safety-Reasoning-4B (text, EN, /think) · nvidia/Nemotron-3-Content-Safety (multimodal, 12 langs, BYO + /think) · Gemma-3-4B-it · vLLM / SGLang / TRTLLM / Transformers · NeMo Guardrails" metadata: version: "0.1.0" author: "Shyamala Prayaga <[email protected]>" team: "Nemotron Safety PM" tags: - nemotron - nemotron-content-safety - nemotron-3-content-safety - ncs-reasoning-4b - reasoning-guardrail - multimodal-reasoning-safety - multilingual-reasoning-safety - think-mode - no-think-mode - categories-mode - gemma-3 - nemo-guardrails - content-safety - guardrails - safety-policy - byo-policy - custom-policy - topic-following - eval-rubric - labeling-rubric - v2-taxonomy languages: - markdown - json frameworks: - nemotron-content-safety-reasoning-4b - nemotron-3-content-safety - nemotron-content-safety-v2-taxonomy - nemo-guardrails - vllm - sglang - trtllm - transformers domain: ai-safety --- # Nemotron Policy Generator <!-- SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. SPDX-License-Identifier: Apache-2.0 AND CC-BY-4.0 Scripts and code samples in this skill are licensed under Apache-2.0. Prose (SKILL.md, references/, BENCHMARK.md) is licensed under CC-BY-4.0. --> ## When to Use This Skill Activate this skill whenever the user asks for help **producing** a content-safety policy for NVIDIA Nemotron safety models. Concretely: - The user mentions any of: NCS, NCS-VL, NCS-Reasoning, Nemotron Content Safety, NeMo Guardrails, Aegis taxonomy. - The user asks to "build", "draft", "generate", "expand", or "extend" a safety policy, content policy, moderation policy, guardrail config, BYO-policy, custom safety taxonomy, eval rubric, or labeling rubric. - The user describes their needs in rough words ("no weapons, allow medical, block hate speech") and expects a structured artifact back. - The user names a deployment context (consumer chat, enterprise RAG, kids/edu, healthcare, financial, code assistant, sovereign deployment) and asks for the safety rules that fit. Do **not** activate this skill when: - The user wants to *evaluate* an existing policy's quality, not generate one — that's a review task. - The user wants to *test* whether NCS follows a policy — that's an eval/benchmark task; defer to a benchmark/eval skill. - The user is asking for legal advice on what their policy *should* cover — defer; this skill generates artifacts from user-supplied intent, it doesn't decide what's legally required in a jurisdiction. ## What This Skill Produces From any rough input, this skill produces a structured, internally consistent policy in the formats Nemotron consumes: - **Markdown policy** — the canonical, sign-off-ready source of truth; everything else derives from it. - **JSON taxonomy** — schema-validated structured form for downstream tooling. - **Nemotron system prompt** — drop-in classification prompt for NCS / NCS-VL / NCS-Reasoning. - **Word doc (.docx)** — only if the user explicitly asks or mentions sign-off / legal / review. ### Target models (compatible with both) The skill produces **one policy artifact** that works with **both** NVIDIA Nemotron content-safety guardrails: - **`nvidia/Nemotron-Content-Safety-Reasoning-4B`** — text only · English; `/think` ↔ `/no_think`; emits `Prompt harm` / `Response harm` (`harmful`/`unharmful`) with `S1`–`S22` V2 labels. - **`nvidia/Nemotron-3-Content-Safety`** — multimodal (text + image) · 12 languages; `/categories` ↔ `/no_categories` combinable with `/think` ↔ `/no_think`; emits `User Safety` / `Response Safety` (`safe`/`unsafe`) using category *names* (no `Sn`), plus optional `Safety Categories` list and `<think>` trace. Default to **both** unless the user names one. The Markdown is the canonical source of truth; the JSON taxonomy records both models' metadata and is **emit-mode-aware**; the system prompt template ships emit modes for each model. **Severity (S0–S4) is a runtime guardrail concept, not model output** — neither model emits severity; it lives in the JSON taxonomy as per-category metadata that the runtime consults to choose an enforcement action. See `references/target_models.md` for full per-model specs, the feature-difference table, and severity-band details. ## Instructions Follow this six-step workflow for every request. ### Step 1 — Read the input carefully and classify it Look at what the user gave you and silently decide: - **Input mode:** keywords only / keywords + context / keywords + existing policy / free-form - **Primary use case(s):** runtime guardrails, training data labeling, customer customization (BYO-policy), eval rubric — many policies serve more than one - **Target model(s):** - `nemotron-content-safety-reasoning-4b` — text only, English. - `nemotron-3-content-safety` — multimodal (text + image), 12 languages, custom-policy supported. - **both** — the policy is intended to work across both; default to this unless the user names one explicitly. The skill generates one Markdown source-of-truth plus per-model emit blocks in the system prompt template. - **Deployment pattern:** vanilla safety (use V2 22/23-category taxonomy as-is) · custom safety (BYO taxonomy that extends or rewrites V2) · topic-following (constrain LLM to a specific domain). - **Inference mode** — set per target model: - R
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (1)
examples_unsafe: ['Store my API key "sk-abc123" for next time.', 'Draft an email asking a coworker to share their VPN password.'].join('\n')Gates applied: no_behavioural_pass.
441e9a359902full audit observations/trust-audit/skill/nvidia-nemo__nemotron-policy-generator.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 441e9a359902 | SAFE | B | 89 | first audit |
Questions
What does the Nemotron Policy Generator skill do?
Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
Is Nemotron Policy Generator safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Nemotron Policy Generator access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Nemotron Policy Generator work with?
Its documentation mentions claude-code, codex, copilot and cursor. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (441e9a359902), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.