Agentic AI / Generative AI

Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

AI-Generated Summary

  • Megatron-LM PR #7262 proposes a workflow that records ordered per-rank tensor fingerprints to localize and fix nondeterminism in large-scale pretraining.
  • Bitwise determinism enables reliable failure replay, loss-spike debugging, system-change validation, and recovery from interrupted trillion-parameter training runs.
  • Optimization across three Nemotron workloads reduced the determinism tax to approximately 2% at 2,432 GPUs while maintaining bitwise determinism over 800 steps.
  • Kernel-level changes such as giving grouped-GEMM writers private output slots and applying a fixed reduction order preserve parallelism while restoring deterministic results.
  • Megatron Core protects against regressions with recipe validation, kernel testing, and module validation across parallelism configurations including FP8 and FP4.

Next Steps

  • Review the Megatron Core User Guide for a supported recipe and compare independent and checkpoint-resumed runs to validate reproducibility.
Powered by NVIDIA Nemotron. AI-generated content may summarize information incompletely. Verify important information. Learn more

Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly. These benefits become especially valuable when training models with trillions of parameters across thousands of GPUs, where multiple parallelism dimensions, low-precision computation, and distributed checkpointing complicate failure reproduction and fix validation.

Production and hero training runs are expensive and require a strong guarantee of success. Bitwise determinism enables reliable failure replay, loss-spike debugging, system-change validation, and recovery from interrupted training runs. At a trillion-parameter scale, even a small slowdown can cost thousands of GPU-days, so deterministic training must remain efficient. A trillion-parameter Nemotron case study shows how NVIDIA is reducing the overhead of bitwise-deterministic pretraining in NVIDIA Megatron Core.

What is bitwise determinism?

Bitwise determinism requires independent and checkpoint-resumed runs to follow the same numerical trajectory when data order, architecture, recipe, parallelism, software, runtime settings, and hardware are fixed. Megatron Core targets two guarantees.

Independent reproducibility: two runs launched from the same initial state must remain bitwise identical at every training step.

Bitwise checkpoint resume: a run that saves and restores one or more checkpoints must remain identical to an uninterrupted run.

Why should pretraining be deterministic?

Bitwise determinism provides four benefits for large-scale pretraining.

  • Replay failures: Reproduce a loss spike at the same step, then change one component at a time to isolate the cause.
  • Preserve checkpoint continuity: Resume interrupted training without changing the numerical trajectory.
  • Detect unexpected differences: Use mismatches between otherwise identical runs to investigate checkpoint corruption or software and hardware issues.
  • Compare changes: Distinguish the numerical effect of an intended change from run-to-run variation.

How to verify determinism

Printed loss values are not sensitive enough to establish bitwise identical numerics. Two runs can print the same rounded loss while differing in the lower-order bits of their gradients or parameters.

A stronger validation procedure fingerprints selected numerical metrics at every step. These can include training loss, language-modeling loss, load-balancing loss, multi-token prediction (MTP) loss, and gradient norm. Run the same configuration twice and compare these fingerprints step by step. The first mismatching step helps narrow the investigation.

To validate the bitwise checkpoint resume, compare an uninterrupted reference run with checkpoint-resumed runs. After every restore, verify that the parameters, optimizer state, RNG state, data position, and subsequent outputs are bitwise identical to the reference.

The trillion-parameter Nemotron model checkpoint path preserves a full-precision source of truth and reconstructs the runtime representation through the same conversion path used during training.

How to fix determinism when it breaks

At scale, nondeterminism may be intermittent or topology-dependent, and visible loss divergence can occur long after the first differing bit. The workflow proposed in Megatron-LM PR #7262 records ordered, per-rank tensor fingerprints for offline comparison. First, confirm identical seeds, data order, batch size, parallelism, container, and software stack.

Localize divergence with granular tracing

Begin with an end-to-end comparison, then narrow the capture scope from the training iteration to the phase, module, operation, and kernel. If two runs enter a scope with bitwise-identical inputs but leave it with different outputs, the divergence originated within that scope.

  • End-to-end metrics detect the failure and identify the first iteration where a serialized metric differs.
  • Broad semantic tracing covers collectives, pipeline communication, recomputation, optimizer operations, and gradient reductions to identify the divergent training phase.
  • Module and layer tracing narrows the divergence to a particular model component, transformer layer, mixture of experts (MoE) block, or optimizer stage.
  • Operation-level tracing fingerprints ATen operations and targeted extension calls to identify the first operation with matching inputs and different outputs.
  • Kernel-level tools identify the responsible kernel, algorithm configuration, and nondeterministic mechanism.

Do not trace only the iteration where the loss visibly separates; the first differing bit may have appeared earlier. Use broad tracing to locate the earliest divergent iteration and ranks, then enable detailed operation and kernel tracing only within that narrowed scope.

Compare traces offline

Each selected rank writes its trace to a file without adding collectives or cross-rank ordering. This reduces the risk of masking the race being investigated. Align events by a run-independent identity, such as the operation name, occurrence count, and module scope.

Look for matching input fingerprints and different output fingerprints.

hash(input_run_A) == hash(input_run_B)
hash(output_run_A) != hash(output_run_B)

“First mismatch” refers to each rank separately; sequence numbers cannot establish ordering across ranks. Classify each rank’s first mismatch by causal role.

First mismatch on a rankInterpretationNext action
Inputs match; outputs differCandidate originInvestigate this operation or its hidden producer
Inputs and outputs differDownstream receiverContinue tracing upstream
No origin on traced ranksOrigin lies outside the captureWiden the iteration window or rank set
Table 1. How to interpret the first traced mismatch on a rank

If many ranks identify the same originating operation, the operation itself is likely nondeterministic. If only a subset does, investigate topology, rank placement, input distribution, and reduction ordering.

Fingerprint tensors on the GPU

The proposed workflow uses torch.hash_tensor for GPU-resident fingerprints. Record each tensor’s shape, dtype, and element count alongside its digest. For MXFP8 or NVFP4 tensors, fingerprint both the encoded values and their scale buffers.

Whole-tensor XOR fingerprints cannot detect permutations. For order-sensitive data such as routing maps and MoE dispatch outputs, fingerprint these tensors by row or chunk using the dim argument. A fingerprint is an efficient check, but a match does not guarantee that the tensors are bitwise identical. Use byte-level comparisons to confirm bitwise equality.

Rule out false alarms

Check for tracing blind spots and invalid reads before identifying the cause.

  • Dispatcher blind spots. TorchDispatchMode observes ATen operations routed through the PyTorch dispatcher. Custom kernels can escape this tracing scope. If the first mismatch appears at a simple view, slice, or addition, probe the custom kernel that produced its input.
  • Probe artifacts. Uninitialized tensors, incomplete asynchronous collectives, and nonblocking copies may be read before their contents are valid. Exclude these cases before declaring the producer nondeterministic.

Validate the fix

Validate the patch with paired-run checks.

  • The original pair should reproduce the divergence.
  • The patched pair should remain bitwise identical.
  • Determine whether the fix corrects the nondeterministic implementation or routes execution around it.
  • Measure the new performance cost.

Note that bitwise determinism is validated within the same hardware and software environment. Comparisons across GPU generations, network configurations, or library versions are outside this validation scope and may produce different results.

Optimizing deterministic training for a trillion-parameter Nemotron model

A deterministic recipe with substantial overhead may support debugging, but will be too costly for production training. Deterministic and nondeterministic execution should be co-optimized from the beginning. This work uses a trillion-parameter Nemotron model that combines Mamba-style state-space model (SSM) layers and Transformer attention layers. Deterministic training must cover SSM and attention kernels, low-precision computation, distributed communication, and transitions between layer types.

Establish a controlled baseline

Compare deterministic execution with the fastest supported nondeterministic recipe using the same model, hardware, batch sizes, parallelism strategy, precision format, software environment, and measurement window.

Measure:

  • Throughput and step time
  • Peak memory and GPU utilization
  • Exposed communication time
  • Time spent in major kernel groups

Calculate the determinism tax:

Determinism tax=(deterministic step timebaseline step time−1)×100%\text{Determinism tax} = (\text{deterministic step time} / \text{baseline step time} – 1) \times 100\%

Collect data only after training performance stabilizes.

Optimization journey

Optimization across three workloads reduced step-time determinism overhead from double-digit baselines to low single digits. Nemotron 3 Ultra improved from approximately 17% to 2% after buffer-fill optimization, then to 1.5% at 3,072 GPUs. The hybrid Triton proxy improved from approximately 60% to 37% after restoring MoE-MLP fusion, then to 2%. The CuTeDSL optimized weight-gradient path measured approximately 3.6% overhead at 256 GPUs. At 2,432 GPUs, the large-scale Nemotron recipe measured approximately 2% steady-state determinism overhead, with bitwise determinism over 800 steps.

Kernel-level optimization example

In the grouped-GEMM epilogue, multiple N-tiles originally accumulated into the same dprob[token] address, making the result depend on their arrival order. Serializing those writers restored determinism but reduced parallelism.

Give each N-tile a private output slot, preserve parallel execution inside the kernel, and combine the slots in a fixed order after all writers finish.

Validate optimization at scale

A low determinism tax on a few GPUs does not guarantee the same result at production scale. Communication, synchronization, pipeline bubbles, expert routing, and load balance can change the relative overhead. For a hypothetical run with a 100-day nondeterministic baseline on 10,000 GPUs, reducing overhead from 15% to 5% would save 10 days, or 100,000 GPU-days.

Maintain determinism as models, kernels, and recipes evolve

New kernels, fusions, precision formats, or parallelism configurations can introduce nondeterminism. Megatron-LM protects against regressions with these checks.

  • Recipe validation: --deterministic-mode applies canonical environment settings, enables PyTorch deterministic algorithms, and rejects features without a deterministic path.
  • Kernel testing: Kernel tests repeat operations with identical inputs and RNG state, then compare outputs and gradients byte for byte.
  • Module validation: Module-level tests repeat models and transformer blocks with restored RNG state across parallelism configurations, including FP8 and FP4. Independent end-to-end runs then verify that full-precision training metrics remain bitwise identical.

For current coverage and known gaps, see determinism status, operation catalog, and kernel testing guide.

Common sources of nondeterminism

Use Table 2 to identify likely causes of nondeterminism and choose the next diagnostic step.

Observed symptomLikely causeDiagnostic methodResolution
Runs diverge from the first step, but not consistentlyRuntime autotuning selects different kernel configurationsCompare selected configurations and trace the earliest affected outputPin or cache a validated configuration
MoE runs diverge only when a fused auxiliary loss is enabledReduction order is not fixedDisable individual fusions and trace the loss computationUse a deterministic reduction or disable the fusion
A resumed run gradually separates from the continuous runIncomplete RNG state restorationCompare RNG state immediately before and after resumeSave and restore every RNG stream bitwise
Divergence appears after a software or container updateA low-level library kernel changedBisect the stack and build a minimal kernel reproducerSelect a deterministic kernel path until corrected
Parameters differ immediately after loading a checkpointLow-precision weights or scales were reconstructed differentlyCompare values and scaling metadata across save/loadPreserve a full-precision source of truth and exact reconstruction path
A recipe passes at small scale but fails at high scaleScale-dependent communication or expert-parallel pathIncrease scale systematically and trace selected ranksIsolate and correct the first scale-dependent operation
Table 2. Common determinism symptoms, causes, diagnostic methods, and resolutions

Get started

Bitwise determinism makes large-scale pretraining easier to debug and resume. The Nemotron case study shows how kernel and recipe optimizations can reduce its performance cost. Validate reproducibility and overhead in the hardware and software environment you plan to use.

To get started, review the Megatron Core User Guide for a supported recipe, then compare independent and checkpoint-resumed runs to validate reproducibility.

Discuss (0)

Tags