Atlas / Skills / nvidia-nemo / Nemo Megatron Bridge

Nemo Megatron BridgeSAFE

skills/nvidia-nemo/nemo-megatron-bridge

Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
Apache-2.0
Stars
2,139
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

This example walks through LoRA fine-tuning of Nemotron-3 Ultra for the Text2SQL use case using NeMo Megatron-Bridge. Everything is driven from a single notebook: mbridge_lora_cookbook.ipynb.

If you've seen the Nemotron-3 Super version, this follows the same spirit — prepare data, convert the checkpoint, fine-tune with LoRA — with one important difference: Ultra is a 550B-total / A55B-active hybrid Mamba-Transformer MoE, which doesn't fit on a single node. So instead of running steps inline with torchrun, each step here is submitted as a multi-node SLURM job (via Pyxis/enroot). The notebook is meant to be run from a cluster login node.

What you'll do

  1. Set up your environment — one cell defines all SLURM settings (node counts, account,

partition, QOS) and paths, and writes config.env. This is the only place you should need to edit.

  1. Prepare data — build a BIRD

Text2SQL training.jsonl from both the no-reasoning and reasoning splits (the full set by default; set MAX_TRAIN_SAMPLES for a smaller smoke run) using Ultra's tokenizer/chat template.

  1. Convert the Hugging Face checkpoint to Megatron-Bridge format (distributed import — CPU import

isn't feasible at 550B).

  1. Fine-tune with LoRA on the prepared data (packed sequences) and save an adapter.

Each step follows the same rhythm: a cell to launch the SLURM job, a cell you can re-run to check its status, and a sanity-check cell that confirms you got the expected output before moving on.

Prerequisites

  • A SLURM multi-node cluster with at least 48 GPUs (H100 and above) with Pyxis/enroot (srun --container-image=...). This notebook was tested on GB200 nodes (4 GPUs/node).
  • The **Nemotron-3 U
Read from source at commit 441e9a359902OBSERVED · 2026-10-09
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: nemotron-3-ultra-text2sql-lora
description: >-
  Run the Nemotron-3 Ultra Text2SQL LoRA fine-tuning tutorial (NeMo Megatron-Bridge) end-to-end
  for the user on their SLURM cluster: data prep, distributed checkpoint conversion, and packed
  LoRA fine-tuning of the 550B hybrid Mamba-Transformer MoE, ending at a saved adapter. Use when
  the user wants to run this cookbook, fine-tune Nemotron-3 Ultra with LoRA, or adapt the notebook
  to their own cluster.
---

# Nemotron-3 Ultra Text2SQL LoRA — runbook for a coding agent

This skill helps you run the cookbook in this directory (`mbridge_lora_cookbook.ipynb`) on the
user's behalf. The notebook is generic and ships with placeholders; your job is to gather the
user's environment details, fill them in, launch the SLURM jobs, watch them, and report results.

## What the tutorial does

Three steps, in order, each a SLURM job:

1. **Data prep** — builds a BIRD Text2SQL `training.jsonl` from both the no-reasoning and reasoning
   splits, formatted with Ultra's tokenizer/chat template. Short CPU job.
2. **Convert** — distributed import of the Hugging Face base checkpoint into Megatron-Bridge format.
   A multi-node GPU job (CPU import is not feasible for a 550B model).
3. **LoRA fine-tune** — packed-sequence LoRA training on the prepared data; saves a LoRA adapter.
   A multi-node GPU job.

## What you must understand before running

- **Ultra is a 550B-total / A55B-active hybrid Mamba-Transformer MoE.** It does **not** fit on one
  node, so every heavy step is a **multi-node SLURM job** submitted with `sbatch` and run in a
  container via Pyxis/enroot. Run everything from a cluster **login node** where `sbatch`/`squeue`/
  `sacct` are available.
- **Scale.** At the shipped parallel settings, both convert and train need **48 GPUs**. Node count
  is derived automatically as `48 / GPUS_PER_NODE` (e.g. 12 nodes at 4 GPUs/node). The user's QOS
  must permit a job of that size — an interactive or small-node-capped QOS will not work.
- **Single config.** Everything is driven by one file, `config.env`, which the notebook's setup
  cell generates from the values you fill in. Every step and every `slurm/*.sbatch` script sources
  it. You can run the notebook cell, or write `config.env` directly with the same keys.
- **One output root.** `WORKSPACE` is the single output root; everything generated lands under
  `$WORKSPACE/{base, dataprep, trained, cache/hf, logs}`. The base checkpoint (`HF_MODEL_PATH`) is
  the only separate, read-only path.
- **The rhythm per step:** a **launch** cell submits the job, a re-runnable **check** cell shows
  status (`sacct`/`squeue`), and a **sanity** cell confirms the expected output exists before you
  move on. Follow this loop; don't skip the sanity check.

## Information to gather from the user

Before launching anything, ask the user for the following and confirm the prerequisites. Don't
guess these — a wrong value wastes a large multi-node allocation. Prefer asking all of them up front
in one batch.

**How to reach the cluster**
- How do you connect to the login node where SLURM jobs are submitted (e.g. the ssh host)?
- Is there a separate data-transfer host you prefer for large file moves?

**SLURM settings**
- SLURM **account** to charge.
- **GPU partition** and a **QOS** that allows a multi-node job of `48 / GPUS_PER_NODE` nodes (not an
  interactive or small-node-capped QOS). Confirm the wall-clock limit is enough (convert is short;
  training is well under a couple of hours by default).
- **CPU partition** and **QOS** for the short data-prep job.
- **GPUs per node** on the target nodes (the tutorial targets GB200 at 4 GPUs/node; the node count
  derives from this).

**Paths (all on a shared filesystem the compute nodes can mount)**
- **`WORKSPACE`** — the output root to create/use.
- **`HF_MODEL_PATH`** — where the **already-downloaded** Ultra base checkpoint lives (read-only
  input). The tutorial does **not** download the base model; confirm it is present.
- The **shared-filesystem root** to bind-mount into the container (must contain both `WORKSPACE`
  and `HF_MODEL_PATH`).

**Container & credentials**
- The **container image** to use (path to a prepared image or a registry reference). The notebook
  ships a placeholder; this must be filled with a real Ultra-capable image.
- A **Hugging Face token** so BIRD can be downloaded during data prep. The tutorial expects it at
  `${WORKSPACE}/cache/hf/token`; ask the user to place it there (or provide it so you can), and
  reference it by path — never print or echo a token.

If the user has an environment-reference document for their cluster, ask for it first and pull these
values from there instead of asking one by one.

## How to run it

1. From the login node, `cd` into this cookbook directory (it must be on the shared filesystem).
2. Fill the config: either edit the notebook's **Environment & SLURM Setup** cell and run it, or
   write `config.env` directly with the values gathered above. The setup cell has a guard that
   refuses to proceed while any placeholder (`<...>`) remains — make sure none are left.
3. Run the three steps in order. For each: submit via the launch cell/`sbatch`, **poll** the check
   cell until the job reaches `COMPLETED`, then run the sanity cell.
4. **Poll, don't block.** These are long-running multi-node jobs. Submit, then check back
   periodically with `sacct`/`squeue` — do not hold an interactive session open waiting, and do not
   stream logs live.

## Verifying success per step

- **Data prep:** `$WORKSPACE/dataprep/training.jsonl` exists and has many rows; a sampled record
  shows the Nemotron-3 chat template.
- **Convert:** `$WORKSPACE/base/latest_checkpointed_iteration.txt` plus an `iter_*` checkpoint dir
  exist.
- **Train:** under `$WORKSPACE/trained/<experiment-name>/` there is a
  `latest_checkpointed_iteration.txt` and an `iter_*` adapter checkpoint; the training log shows the
  loss trending down and ends with a `LORA
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-09 · audit v0.4.1 · source sha 441e9a359902full audit observations/trust-audit/skill/nvidia-nemo__nemo-megatron-bridge.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-09441e9a359902SAFEB89first audit
05

Questions

What does the Nemo Megatron Bridge skill do?

Developer Asset Hub for NVIDIA Nemotron — A one-stop resource for training recipes, usage cookbooks, datasets, and full end-to-end reference examples to build with Nemotron models

Is Nemo Megatron Bridge safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Nemo Megatron Bridge access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (441e9a359902), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement