Quick Start
Overview
This section helps you quickly use dynamic graphs (PyNative) to complete an LLM training and gain an intuitive understanding of the overall process. The following uses DeepSeek-V3 as an example to perform the minimum-scale pre-training on two devices.
Dynamic graph training uses the unified script run_mindformer.py as the entry point and routes to the dynamic graph trainer mindformers.pynative.trainer.Trainer through --mode 1. The trainer builds the model, dataset, and optimizer and drives the training loop. A task can be summarized into three steps:
Preparing the configuration file: Compile a dynamic graph YAML file to connect the model, data, parallelism, and optimizer.
Starting training: Use
msrunto start multi-device training.Viewing the result: Check the logs to ensure that the loss/grad norm decreases properly in each step.
For details about the complete process and detailed configuration of each capability, see Function Overview. (The training guide and feature pages will be released later.)
Prerequisites
MindSpore and MindSpore Transformers have been installed. For details, see Installation Guide.
The Ascend hardware environment has been set up, and the CANN has been correctly configured.
A Megatron dataset (
.bin/.idx) has been prepared. The repository scriptpreprocess_indexed_dataset.pycan be used to create a dataset (JSON is transmitted to BIN/IDX). For details, see "Dataset."
Configuring the Data Path Connection
The preprocessing output is a pair of
xxx.binandxxx.idxfiles. The file name is in the format of<output-prefix>_text_document.bin(the suffix_text_documentis added by default in the script). In the YAML file on this page,train_dataset.dataloader.config.data_pathmust be set to the prefix containing the suffix. For example, if--output-prefix /path/megatron_datais used,/path/megatron_data_text_document.binwill be generated. In this case,/path/megatron_data_text_documentwill be filled in the configuration. In addition, record theeod(eos) token ID corresponding to the tokenizer used during preprocessing, which will be used in the followingconfig.eod.
Step 1: Preparing the Configuration File
The dynamic graph uses the YAML configuration in dataclass style. The top-level sections correspond to the weight, training, parallelism, optimizer, learning rate, data, and model. This page provides a complete example configuration for two devices pynative_ds3.yaml, which can be directly downloaded and used. The content of each section is as follows (for details about the complete field description, see "Configuration File Description"):
checkpoint:
enable_save: False # Do not save the weight for quick verification.
save_path: "./output/ds3"
training:
steps: 10 # Number of training steps.
local_batch_size: 1 # Batch size per device (per forward pass).
global_batch_size: 2 # Global batch size. For details, see the following description.
max_norm: 1.0 # Gradient clipping (valid when the value is greater than 0).
seed: 42
parallelism:
data_parallel_shard: -1 # FSDP. The value -1 indicates automatic sharding based on the number of available devices.
expert_parallel: 1
tensor_parallel: 1
context_parallel: 1
pipeline_parallel: 1
# SP is automatically enabled with TP. Currently, it cannot be disabled. (You do not need to configure this item. For details, see the document on distributed parallel training.)
optimizer:
type: AdamW
betas: [0.9, 0.95]
eps: 1.e-8
weight_decay: 0.01
lr_scheduler:
type: ConstantWarmUpLR
learning_rate: 1.e-5
warmup_ratio: 0
train_dataset:
dataloader:
type: BlendedMegatronDatasetDataLoader
datasets_type: "GPTDataset"
sizes: [1000, 0, 0] # Number of samples for [Training, Testing, Evaluation]. Currently, this parameter takes effect only for the training set.
column_names: ["input_ids", "labels", "loss_mask", "position_ids"]
shuffle: false
config:
seed: 1234
seq_length: 4096 # The value must be the same as that of model.seq_length.
split: "1, 0, 0" # Split based on [Training, Testing, Evaluation]. This parameter is required if data_path is set. If this parameter is missing, an error will be reported.
eod: 1 # Token ID of the eod(eos) in the dataset.
pad: -1 # Token ID of the pad in the dataset.
eod_mask_loss: False
reset_position_ids: False
create_attention_mask: False # Four columns are output. If this parameter is set to True, add attention_mask to column_names.
reset_attention_mask: False
create_compressed_eod_mask: False
eod_pad_length: 128
data_path: # Sampling weight (relative value, automatically normalized) + bin prefix (containing _text_document).
- '1'
- "/path/megatron_data_text_document"
drop_remainder: True
num_parallel_workers: 8
model:
model_type: deepseek_v3
architectures: DeepseekV3ForCausalLM
seq_length: 4096 # The value must be the same as that of train_dataset.dataloader.config.seq_length.
# For details about other model structure hyperparameters (such as hidden_size, num_hidden_layers, and MoE), see the complete example configuration (tile format) in the preceding link.
Precise meaning of
global_batch_size: The framework infers the number of gradient accumulation steps based onnum_accumulation_steps = global_batch_size // (data_parallel x local_batch_size). When gradient accumulation is not enabled (that is, the product of the three values is exactly equal),global_batch_sizeis equal tolocal_batch_size x Data parallelism degree. Onceglobal_batch_sizeis greater than this product, the excess multiple is the number of gradient accumulation steps. For the exact definition, see "Configuration File Description."
Dataset Segments
When BlendedMegatronDatasetDataLoader is running, all of the following are datasets_type, sizes, and nested config blocks (including seq_length/split/eod/pad/data_path/create_compressed_eod_mask) are required.
Field |
Description |
|---|---|
|
Dataset type. |
|
Number of samples for |
|
Length of the returned sequence. This value must be the same as that of |
|
Training/Testing/Evaluation sharding ratio (for example, |
|
Token ID of EOD (EOS)/pad, which is obtained from the tokenizer during preprocessing. |
|
List. Every two elements (sampling weight and bin prefix) form a group. The weight is a relative value and is automatically normalized (the sum does not need to be 1). The bin prefix contains the |
For details about the meaning of each field, multi-data source combination, and scenario-specific configurations such as compressed EOD mask, see "Dataset." This page provides only the minimum configuration.
Model Section
The model section of DeepSeek-V3 contains dozens of structure hyperparameters (hidden_size, num_hidden_layers, MoE routing, etc.), making manual writing both tedious and error-prone. The complete example configuration pynative_ds3.yaml provided on this page already includes all model section fields (dataclass style and structure hyperparameter tiling). You can directly download and use it. You only need to ensure that seq_length is consistent with the dataset.
Note: The configuration under
configs/deepseek3/is the static graph legacy structure (the structure hyperparameters are nested undermodel.model_config,architecturesis a list, andcontext.modeis set to0). You cannot replicate the entire section to the dynamic graph configuration. Instead, you need to move the structure hyperparameters up one level and tile them undermodel:, and changearchitecturesto a string.
Step 2: Starting Training
Dynamic graph training uses msrun to start multiple devices. Run the following command in the root directory of the MindFormers repository (place the downloaded pynative_ds3.yaml file in this directory) to start training on two devices:
msrun --worker_num=2 --local_worker_num=2 --master_port=8118 \
--join=True --log_dir=./msrun_log \
run_mindformer.py --config pynative_ds3.yaml --mode 1
--worker_num/--local_worker_num: Total number of devices/Number of devices on the local node.--config: YAML file in the previous step.--mode 1: uses dynamic graphs.
To enter dynamic graph training,
--mode 1must be explicitly passed in the startup command. (The entry reads only the command line--modeand does not read the YAML file.) Thecontextsection in the YAML file can be omitted. By default, the dynamic graph mode (mode: 1) and Ascend backend are used. If you need to adjustmax_device_memory, explicitly configure it.
For details about how to start clusters of different scales (single-device and multi-node), see "Starting Tasks."
Step 3: Viewing the Result
The training logs are output to the specified directory in --log_dir. Each worker has a subdirectory (for example, ./msrun_log/worker_0.log). Open any worker log. If the following output is displayed, the training is running properly:
{ step:[ 1/ 10], loss: 11.813965, per_step_time: 13570ms, load_balancing_loss: 1.093977, lr: 1.000000e-05, grad_norm: 13.831877, throughput: 1.36T }
{ step:[ 2/ 10], loss: 11.755612, per_step_time: 710ms, load_balancing_loss: 1.106543, lr: 1.000000e-05, grad_norm: 19.926786, throughput: 25.92T }
Key points for judgment:
stepcontinuously increases, andlossgenerally decreases (fluctuations may occur in the first few steps).grad_normis a finite value rather thanNaN.per_step_timetends to be stable after the first step (the first step includes initialization overhead and is usually significantly slower).The MoE model (such as DeepSeek-V3 in this example) will additionally print
load_balancing_loss.
In addition:
Weight: If
checkpoint.enable_saveis enabled, the weight is saved tocheckpoint.save_pathin the Safetensors format.More metrics: You can use
monitor.train_stateto collect per-parameter norms and local/device-level loss. (Note: Currently, dynamic graph monitoring metrics are output only through training logs, and the TensorBoard configuration does not take effect.) For details, see "Training Metric Monitoring and Profiling."