MindSpore Transformers

Introduction

  • Quick Start
  • Overall Structure
  • Model Support Library

Installation

  • Installation Guide

Training Guide

  • Training Guide

Function Features

  • Function Overview
  • Starting Tasks
  • Configuration File Description
  • Logs
  • Datasets
  • Hyperparameters and Optimizers for Training
  • Distributed Parallel Training
  • Training Graphics Memory Optimization
  • Weight Saving and Loading
  • Resumable Training
  • Training Metric Monitoring and Profiling
  • Other Training Features
  • Static Graph Features

Environment Variables

  • Environment Variables

Contribution Guide

  • MindSpore Transformers Contribution Guidelines
  • Modelers Contribution Guidelines

FAQ

  • Model-Related FAQ
  • Feature-Related FAQ

Static Graph Implementation (Deprecated)

  • Overall Structure
  • Full-process Guide to Large Models
  • Features
  • Advanced Development
  • Excellent Practice
    • Practical Case: Creating a Docker Image for MindSpore Transformers
    • Practice Case of Using DeepSeek-R1 for Model Distillation
    • Practical Case: Converting Model Weights to Megatron Model Weights
    • Practice Case: Interconnecting MindSpore Transformers with General Evaluation Tools
    • Practice Case: Using GLM4-9B for Multi-Device Model Fine-Tuning
  • Environment Variable Descriptions
MindSpore Transformers
  • »
  • Excellent Practice
  • View page source

Excellent Practice

  • Practical Case: Creating a Docker Image for MindSpore Transformers
  • Practice Case of Using DeepSeek-R1 for Model Distillation
  • Practical Case: Converting Model Weights to Megatron Model Weights
  • Practice Case: Interconnecting MindSpore Transformers with General Evaluation Tools
  • Practice Case: Using GLM4-9B for Multi-Device Model Fine-Tuning
Previous Next

© Copyright MindSpore.

Built with Sphinx using a theme provided by Read the Docs.