Nemo Megatron BridgeSAFE
Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
This example walks through LoRA fine-tuning of Nemotron-3 Ultra for the Text2SQL use case using NeMo Megatron-Bridge. Everything is driven from a single notebook: mbridge_lora_cookbook.ipynb.
If you've seen the Nemotron-3 Super version, this follows the same spirit — prepare data, convert the checkpoint, fine-tune with LoRA — with one important difference: Ultra is a 550B-total / A55B-active hybrid Mamba-Transformer MoE, which doesn't fit on a single node. So instead of running steps inline with torchrun, each step here is submitted as a multi-node SLURM job (via Pyxis/enroot). The notebook is meant to be run from a cluster login node.
What you'll do
- Set up your environment — one cell defines all SLURM settings (node counts, account,
partition, QOS) and paths, and writes config.env. This is the only place you should need to edit.
- Prepare data — build a BIRD
Text2SQL training.jsonl from both the no-reasoning and reasoning splits (the full set by default; set MAX_TRAIN_SAMPLES for a smaller smoke run) using Ultra's tokenizer/chat template.
- Convert the Hugging Face checkpoint to Megatron-Bridge format (distributed import — CPU import
isn't feasible at 550B).
- Fine-tune with LoRA on the prepared data (packed sequences) and save an adapter.
Each step follows the same rhythm: a cell to launch the SLURM job, a cell you can re-run to check its status, and a sanity-check cell that confirms you got the expected output before moving on.
Prerequisites
- A SLURM multi-node cluster with at least 48 GPUs (H100 and above) with Pyxis/enroot (
srun --container-image=...). This notebook was tested on GB200 nodes (4 GPUs/node). - The **Nemotron-3 U
441e9a359902OBSERVED · 2026-10-09What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: nemotron-3-ultra-text2sql-lora
description: >-
Run the Nemotron-3 Ultra Text2SQL LoRA fine-tuning tutorial (NeMo Megatron-Bridge) end-to-end
for the user on their SLURM cluster: data prep, distributed checkpoint conversion, and packed
LoRA fine-tuning of the 550B hybrid Mamba-Transformer MoE, ending at a saved adapter. Use when
the user wants to run this cookbook, fine-tune Nemotron-3 Ultra with LoRA, or adapt the notebook
to their own cluster.
---
# Nemotron-3 Ultra Text2SQL LoRA — runbook for a coding agent
This skill helps you run the cookbook in this directory (`mbridge_lora_cookbook.ipynb`) on the
user's behalf. The notebook is generic and ships with placeholders; your job is to gather the
user's environment details, fill them in, launch the SLURM jobs, watch them, and report results.
## What the tutorial does
Three steps, in order, each a SLURM job:
1. **Data prep** — builds a BIRD Text2SQL `training.jsonl` from both the no-reasoning and reasoning
splits, formatted with Ultra's tokenizer/chat template. Short CPU job.
2. **Convert** — distributed import of the Hugging Face base checkpoint into Megatron-Bridge format.
A multi-node GPU job (CPU import is not feasible for a 550B model).
3. **LoRA fine-tune** — packed-sequence LoRA training on the prepared data; saves a LoRA adapter.
A multi-node GPU job.
## What you must understand before running
- **Ultra is a 550B-total / A55B-active hybrid Mamba-Transformer MoE.** It does **not** fit on one
node, so every heavy step is a **multi-node SLURM job** submitted with `sbatch` and run in a
container via Pyxis/enroot. Run everything from a cluster **login node** where `sbatch`/`squeue`/
`sacct` are available.
- **Scale.** At the shipped parallel settings, both convert and train need **48 GPUs**. Node count
is derived automatically as `48 / GPUS_PER_NODE` (e.g. 12 nodes at 4 GPUs/node). The user's QOS
must permit a job of that size — an interactive or small-node-capped QOS will not work.
- **Single config.** Everything is driven by one file, `config.env`, which the notebook's setup
cell generates from the values you fill in. Every step and every `slurm/*.sbatch` script sources
it. You can run the notebook cell, or write `config.env` directly with the same keys.
- **One output root.** `WORKSPACE` is the single output root; everything generated lands under
`$WORKSPACE/{base, dataprep, trained, cache/hf, logs}`. The base checkpoint (`HF_MODEL_PATH`) is
the only separate, read-only path.
- **The rhythm per step:** a **launch** cell submits the job, a re-runnable **check** cell shows
status (`sacct`/`squeue`), and a **sanity** cell confirms the expected output exists before you
move on. Follow this loop; don't skip the sanity check.
## Information to gather from the user
Before launching anything, ask the user for the following and confirm the prerequisites. Don't
guess these — a wrong value wastes a large multi-node allocation. Prefer asking all of them up front
in one batch.
**How to reach the cluster**
- How do you connect to the login node where SLURM jobs are submitted (e.g. the ssh host)?
- Is there a separate data-transfer host you prefer for large file moves?
**SLURM settings**
- SLURM **account** to charge.
- **GPU partition** and a **QOS** that allows a multi-node job of `48 / GPUS_PER_NODE` nodes (not an
interactive or small-node-capped QOS). Confirm the wall-clock limit is enough (convert is short;
training is well under a couple of hours by default).
- **CPU partition** and **QOS** for the short data-prep job.
- **GPUs per node** on the target nodes (the tutorial targets GB200 at 4 GPUs/node; the node count
derives from this).
**Paths (all on a shared filesystem the compute nodes can mount)**
- **`WORKSPACE`** — the output root to create/use.
- **`HF_MODEL_PATH`** — where the **already-downloaded** Ultra base checkpoint lives (read-only
input). The tutorial does **not** download the base model; confirm it is present.
- The **shared-filesystem root** to bind-mount into the container (must contain both `WORKSPACE`
and `HF_MODEL_PATH`).
**Container & credentials**
- The **container image** to use (path to a prepared image or a registry reference). The notebook
ships a placeholder; this must be filled with a real Ultra-capable image.
- A **Hugging Face token** so BIRD can be downloaded during data prep. The tutorial expects it at
`${WORKSPACE}/cache/hf/token`; ask the user to place it there (or provide it so you can), and
reference it by path — never print or echo a token.
If the user has an environment-reference document for their cluster, ask for it first and pull these
values from there instead of asking one by one.
## How to run it
1. From the login node, `cd` into this cookbook directory (it must be on the shared filesystem).
2. Fill the config: either edit the notebook's **Environment & SLURM Setup** cell and run it, or
write `config.env` directly with the values gathered above. The setup cell has a guard that
refuses to proceed while any placeholder (`<...>`) remains — make sure none are left.
3. Run the three steps in order. For each: submit via the launch cell/`sbatch`, **poll** the check
cell until the job reaches `COMPLETED`, then run the sanity cell.
4. **Poll, don't block.** These are long-running multi-node jobs. Submit, then check back
periodically with `sacct`/`squeue` — do not hold an interactive session open waiting, and do not
stream logs live.
## Verifying success per step
- **Data prep:** `$WORKSPACE/dataprep/training.jsonl` exists and has many rows; a sampled record
shows the Nemotron-3 chat template.
- **Convert:** `$WORKSPACE/base/latest_checkpointed_iteration.txt` plus an `iter_*` checkpoint dir
exist.
- **Train:** under `$WORKSPACE/trained/<experiment-name>/` there is a
`latest_checkpointed_iteration.txt` and an `iter_*` adapter checkpoint; the training log shows the
loss trending down and ends with a `LORATrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
441e9a359902full audit observations/trust-audit/skill/nvidia-nemo__nemo-megatron-bridge.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 441e9a359902 | SAFE | B | 89 | first audit |
Questions
What does the Nemo Megatron Bridge skill do?
Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models
Is Nemo Megatron Bridge safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Nemo Megatron Bridge access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (441e9a359902), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.