Logs
MindSpore Transformers dynamic graph (PyNative) training is started through scripts/msrun_launcher.sh. All logs generated during the running are collected through the worker process logs of msrun. Understanding the log directory structure and key fields is the basis for locating training failures, loss exceptions, and single-device problems.
This section describes the purpose and directory structure of logs, configurable items of msrun, meanings of key fields in training logs, and the sequence of checking logs during troubleshooting. For details about how to start a task, see Starting Tasks.
Log Overview
During PyNative training, all logs are carried by the msrun worker process logs.
msrunworker process logs: The startup scriptscripts/msrun_launcher.shpasses theLOG_DIRparameter (default value:output/msrun_log) to the--log_dirof the underlyingmsrun.msrunredirects the standard output (stdout) and standard error (stderr) of each worker (device) process to a file namedworker_{i}.log(another processscheduler.logis for scheduling). The settingto_std=True(default) of the frameworkloggermirrors all logs to standard outputs. Therefore, all framework logs, including step-by-step training metrics (such as loss, learning rate, and gradient norm), running mode banner, parallel sharding, dataset/weight loading, Python traceback, C++/HCCL bottom-layer errors, and segfault exit information, are recorded in theworker_{i}.logfile. It is also mainly used to locate problems such as process breakdown and networking failures.
Troubleshooting Entry
If the process startup fails, breaks down during running, or reports an underlying error, check the
worker_{i}.logof the corresponding rank (including the complete traceback and HCCL/driver error information).When the training can run properly, check the loss, learning rate, convergence, and time consumed by a single step, as well as the
worker_{i}.logfile, which contains all the metrics printed by the framework by step.If you suspect that a device is abnormal, compare the
worker_{i}.logandscheduler.logof the rank.
Directory Structure Example
In the default configuration (LOG_DIR of msrun_launcher.sh is output/msrun_log by default), the log flushing structure is as follows:
output/
└── msrun_log/ # msrun's --log_dir, which collects worker process logs.
├── scheduler.log # Logs of the scheduling process (networking and process startup/recycling)
├── worker_0.log # stdout/stderr of the rank 0 process.
├── worker_1.log # stdout/stderr of the rank 1 process
└── ... # Each worker has a log file, **worker_{i}.log**.
File Name and Path Source
worker_{i}.log/scheduler.log: It is generated bymsrunin the--log_dirdirectory and named after the worker or scheduler. (This is an msrun behavior and is not defined by MindFormers. For details, see Starting Tasks.)The log directory is controlled by the
LOG_DIRparameter ofmsrun_launcher.sh. The default value isoutput/msrun_log.
msrun Logs
msrun generates a worker_{i}.log file for each worker process in the directory specified by --log_dir, and writes the standard output and standard error of the process into the file. In addition, a scheduler.log file is generated to record the networking and process management. You are advised to use scripts/msrun_launcher.sh for startup. The LOG_DIR parameter specifies the log directory.
Single-Node Startup
Single node (8 devices)
bash scripts/msrun_launcher.sh \
"run_mindformer.py --config /path/to/pretrain_xxx.yaml --mode 1" \
8 8 8118 output/msrun_log False 300
After the startup, files from worker_0.log to worker_7.log and scheduler.log are displayed under the output/msrun_log/ directory.
Multi-Node Startup
In the multi-node scenario, msrun_launcher.sh is executed on each node, and the worker logs of each node are stored in the LOG_DIR directory of the node (the logs are not automatically aggregated to the primary node). The following uses a two-node system with 16 devices (eight devices per node, primary node IP 192.168.1.1) as an example:
Primary node (node 0):
bash scripts/msrun_launcher.sh \
"run_mindformer.py --config /path/to/pretrain_xxx.yaml --mode 1" \
16 8 192.168.1.1 8118 0 output/msrun_log False 300
Secondary node (node 1): The command is basically the same as that for the primary node. You only need to change the node ID parameter from 0 to 1 (the seventh parameter in msrun_launcher.sh), and keep other parameters the same. The worker_{i}.log of this node falls under the output/msrun_log/ of node 1.
Multi-Node Check Prompt
To locate a multi-node issue, log in to the corresponding node and check the
LOG_DIRof the node. For example, if you suspect that global rank 10 (local rank 2 on node 1) is abnormal, go to themsrun_logdirectory on node 1 to check theworker_10.log.
For details about msrun_launcher.sh parameters (WORKER_NUM, LOCAL_WORKER, MASTER_ADDR, MASTER_PORT, NODE_RANK, LOG_DIR, JOIN, and CLUSTER_TIME_OUT), see Starting Tasks.
Key Log Fields
In the PyNative training loop, LossCallback (_print_log of mindformers/pynative/callback/loss_callback.py) to print a line of training metrics to the standard output every several steps. Then, msrun collects the metrics to worker_{i}.log. A typical record is as follows:
[INFO] 2026-06-09 10:20:30 [.../loss_callback.py:231] _print_log: { step:[ 100/ 1000], loss: 2.345678, per_step_time: 850ms, lr: 1.000000e-04, grad_norm: 1.234000, throughput: 12.34T }
If pipeline parallelism (PP) is enabled for a training task, the loss is displayed only on the last stage. For example, if PP 2 is enabled for an 8-device training task, rank_0 to rank_3 belong to stage 0, and the corresponding log file
worker_0–3.logdoes not contain loss information. You need to view the loss information in the log filesworker_4–7.logof stage 1 (rank_4 to rank_7).
The fields are described as follows.
Field |
Example |
Description |
|---|---|---|
|
Current steps/Total steps |
Training step progress (global steps/ |
|
Loss of the current steps |
Training loss of the current steps (single value), which is used to determine the convergence trend. |
|
Duration per step |
Average training duration of the current steps. Pay attention to the performance and jitter. |
|
Current learning rate |
It reflects the warmup and decay scheduling. If the learning rate scheduler cannot obtain the current learning rate in real time, this field is not printed. |
|
Global gradient norm |
Global L2 norm of all parameter gradients. The training stability is determined based on the gradient clipping threshold. If the value is not obtained in this step, |
|
Throughput |
Training throughput of the current steps (unit: |
Supplementary Notes
Lines are concatenated per step by
LossCallback._print_logof the dynamic graph. The field name isgrad_norm(notglobal_norm), and theEpoch,loss_scale, andoverflow condfields do not exist.lossis a single value (not the two-section format of this step/moving average value). When the itemized loss function is enabled for the MoE or MTP model, fields such asload_balancing_lossandmtp_{i}_loss(for example,mtp_1_loss) are added to a line.
If more detailed metrics such as
local_norm,local_loss, or tokens-per-expert of MoE are required, you can enable them usingmonitor. (In dynamic graph mode, these metrics are output to training logs and are not written into TensorBoard.) For details, see Training Metric Monitoring and Profiling.Check whether the training is running in dynamic graph mode. During the startup phase, the banner
Running MindFormers in PYNATIVE_MODE.(fromrun_mindformer.py, corresponding to--mode 1) is printed to confirm that the training has entered the dynamic graph mode. If this banner is not displayed, the training is not started in--mode 1.
Troubleshooting Suggestions
You can check logs in the following sequence to efficiently locate the fault:
Check
scheduler.log:msrun_log/scheduler.logrecords the entire cluster networking process (worker registration, topology construction, cluster initialization, and worker deregistration) and timeout or abnormal exit information (such asNode X is timed out, please check this node's log) of each worker. It is the entry for troubleshooting networking failures and abnormal process exits.If the process crashes or an underlying error is reported, check the
worker_{i}.logof the corresponding rank: The log contains the complete Python traceback and underlying error and exit information of the HCCL and driver.If you suspect that a single device is abnormal, check the
worker_{i}.logof the corresponding rank: Compare the log of the rank with the logs of other normal ranks to locate the differences.If there is a communication or networking problem, check the
scheduler.logand eachworker_{i}.log: The information about networking setup, process startup, and process recycling is recorded inscheduler.log. HCCL errors are usually scattered in the logs of each worker, and therefore need to be compared across ranks.If the issue lies in multiple nodes, log in to the corresponding nodes and check the
LOG_DIRof the localhost: Logs of each node are not aggregated. You need to check theworker_{i}.logof the rank corresponding to themsrun_logon the abnormal nodes.