Pretraining, SFT and reinforcement-learning code extend the earlier model releases, with working examples focused on the 92B model and no supplied datasets.

Huawei has opened the pretraining, supervised fine-tuning and reinforcement-learning code for openPangu-2.0, extending its Ascend-native model releases into training infrastructure. Two repositories divide the work: openPangu-2.0-Training covers pretraining and SFT, while openPangu-2.0-RL handles reinforcement-learning post-training. Both are publicly listed under Apache-2.0. Training repository, RL repository.

For developers working with Huawei accelerators, the release makes more of the machinery behind model adaptation available to inspect and modify. It also gives the announcement a precise boundary: both repositories explicitly state that Huawei supplies no datasets. Access to the training software therefore does not, by itself, reproduce the models that Huawei released.

The weights arrived first

Huawei released the approximately 92-billion-parameter openPangu-2.0-Flash model on June 30, alongside basic inference code and training/inference operators. The approximately 505-billion-parameter Pro model followed on July 31 with weights, basic inference code and a technical report. The September training-code release adds another layer to that rollout. Huawei’s Flash announcement, Pro announcement.

The model cards describe both as mixture-of-experts models trained on Ascend NPUs. Flash activates approximately 6 billion parameters per token; Pro activates approximately 18 billion. Both advertise a 512K context window and approximately 34 trillion training tokens. These are Huawei’s stated specifications. Flash model card, Pro model card.

Sparse activation helps explain the attraction of MoE: each token uses only part of the model’s expert capacity. It does not make a 92B model equivalent to a dense 6B model for memory planning. The full expert weights still require storage, and training adds gradients, optimizer state, activations and communication between devices.

Architecture diagram of Huawei openPangu-2.0 Ascend training stack showing CANN, PyTorch, torch_npu, and distributed parallelism layers.
Huawei openPangu-2.0 Ascend-native training infrastructure stack spanning pretraining, supervised fine-tuning (SFT), and VERL-based reinforcement learning.

What the pretraining and SFT repository contains

The training framework documents support for starting from scratch, continuing pretraining from an existing checkpoint, and supervised adaptation with instruction, multi-turn dialogue or domain-specific data. It also lists parameter-efficient fine-tuning and multi-task training.

Its documented components cover data pipelines, distributed execution, checkpoint management and fault recovery. Data, tensor, pipeline, context and expert parallelism appear in the repository structure, alongside NPU-specific operators and monitoring tools. Those components matter because a large training job depends on distributing work and recovering from interruptions as much as on defining the neural network. Training README.

There is a useful distinction between the model family and the examples supplied with the framework. The README’s model table lists a 92B-total/6B-active configuration with 512K context and a smaller 17.36B-total/1.75B-active sample with 4K context. It does not list a ready-made 505B Pro configuration in that table. Developers should check the available configurations before treating the release as an immediately runnable recipe for every openPangu model.

The recommended environment is specific: CANN 8.5.1, PyTorch 2.9.0, torch_npu 2.9.0, Python 3.11.x and triton-ascend 3.2.1. Those dependencies establish what Ascend-native means operationally: the framework couples familiar PyTorch tooling to Huawei’s accelerator software and NPU optimizations. Software versions and model configurations.

RL builds on VERL and includes tool-using rollouts

The RL repository uses the open-source VERL framework for training orchestration. Huawei documents runtime patches and Ascend optimizations that coordinate the actor, rollout generation and reward evaluation, rather than maintaining a separate VERL fork.

The README describes GSPO/GRPO post-training in two forms. In the reasoning example, a model generates a reasoning trace and answer in one pass, with rule-based checks for correctness and format. In the Code-For-Reasoning example, the model can call a Python interpreter, inspect its output and continue reasoning across multiple turns. The published example paths target the 92B model. RL README.

This makes the release relevant to researchers studying how tool use interacts with training. The software exposes a path for rewarding answers reached through code execution, provided the team supplies suitable tasks and reference answers. It does not establish that a particular reward design will generalize beyond those tasks.

The repository also calls itself a research and demonstration framework. Its deployment notes require network isolation and explain that the code-reasoning scenario executes model-generated Python. Those are concrete operating assumptions to account for when moving from an example run to a shared training service. RL deployment notes.

Code access leaves the data question open

Huawei’s model cards describe a post-training sequence combining fast/slow SFT, specialist reinforcement learning and distillation. The newly visible infrastructure gives developers more to examine, but the existence of an RL framework does not establish that every specialist recipe, dataset and distillation setting from the original model run is present. Pro model description.

The clearest limitation comes from the repositories themselves: datasets referenced in examples are for functional testing, and Huawei provides no datasets. Recreating the original outcome would also depend on data preparation, sampling, training settings, compute and evaluation procedures. A successful launch demonstrates that the software runs in that environment; matching the released model’s quality requires separate evidence. Training usage notes, RL usage notes.

The software and weights also have separate license labels. The training repositories are listed as Apache-2.0, while the model cards identify the weights under the OpenPangu Model License Agreement Version 2.0. A description of the code’s license should not be carried over to the model weights. Flash model license reference.

For Ascend teams, the practical value is a public implementation to examine when adapting models, arranging distributed training or experimenting with tool-assisted RL. Its broader significance is the reduction in engineering work that developers must reconstruct around Huawei’s hardware. How much it improves training cost, stability or model quality remains a question for measured runs on a specified cluster.

Last Update: September 28, 2026