DeepSeek has published Ascend 950 versions of two pieces of its model infrastructure: DeepEP-Ascend for moving mixture-of-experts tokens between devices, and DeepGEMM-Ascend for matrix kernels. Their source code is inspectable, but the headline speeds are narrow kernel measurements. DeepEP’s fastest published results also rely on a manually configured proof-of-concept Huawei hardware development kit that is not publicly distributed yet.

That distinction matters more than a broad claim that the release “replaces CUDA.” A large model needs both fast math on each accelerator and fast communication across the accelerators. DeepSeek is working on both bottlenecks, but these repositories do not by themselves establish complete model training or serving parity with an NVIDIA stack.

DeepEP moves tokens; DeepGEMM does the math

In a mixture-of-experts model, routing sends token activations to selected experts and combines their outputs afterward. DeepEP-Ascend’s README says its Ascend implementation aligns its public buffer API with the NVIDIA version. The Ascend path uses HCCL/HCOMM, UBMEM and URMA, with runtime compilation through DeepJIT. That API alignment can reduce application-level rewrites, while still leaving hardware-specific kernels, streams and firmware to manage.

DeepGEMM-Ascend targets the computation side: BF16, FP8 and FP4 matrix multiplication, grouped expert GEMM, MQA logits and MegaMoE operations. It uses Ascend-specific layouts and matrix primitives. The README documents CANN 9.20, torch_npu, a C++20 toolchain and TileLang for one prenorm kernel. This is an implementation for Ascend 950, not a generic package that runs unmodified on every Huawei NPU generation.

The components therefore address different steps of the same workload. Faster GEMM cannot compensate for slow expert dispatch, and high dispatch bandwidth does not imply a fast whole-model token rate. An end-to-end claim needs the model, routing policy, batch size, network topology and full latency breakdown.

The benchmark figures are specific, and one firmware path is pending

DeepSeek reports up to 99.8% of a stated Ascend 950DT hardware limit for a BF16 dense GEMM shape of 4096 × 7168 × 16384. Its benchmark table gives 431 TFLOPS against a 432 TFLOPS limit for that case, measured with bench_msprof and cold L2. This demonstrates a strong result for one matrix shape under the listed setup. It is not a 99.8% utilization measurement for an entire model run.

For expert communication, DeepEP-Ascend reports FP8 dispatch bandwidth of 373–375 GB/s at expert-parallel size 8 and 323–327 GB/s at size 64, with its stated 16,384 tokens per rank, hidden size 7168 and top-6 routing. Using the midpoint of each range, the EP64 dispatch number is about 13% lower than EP8. The repository attributes larger-group and combine limits partly to reduction overhead and memory contention. It says dispatch reaches roughly 90–95% of its physical payload bandwidth limit through EP32 under those conditions.

The test environment is the practical catch. DeepSeek says the DeepEP measurements used an Ascend 950DT with CANN 9.2.0 and a proof-of-concept HDK supplied to it with manual configuration. That HDK is not a public release. A Huawei Q3 commercial HDK for Atlas 850E is planned for around October 15, 2026, subject to Huawei’s schedule. The README explicitly says its measurements are not results from that unreleased commercial package. Public code therefore does not yet give an outside team the same complete test path.

What developers can inspect and what remains unsettled

Both repositories contain code and test helpers. DeepGEMM-Ascend includes an explicit MIT license. The DeepEP-Ascend repository’s current top-level file list does not show a license file. Public visibility alone does not settle reuse or redistribution rights, so a team planning to ship that component should seek a clear license rather than assume it matches DeepGEMM’s terms.

DeepEP’s own ongoing-work list also matters: some Bucket collectives, expert load-balancing kernels, graph capture and CPU-backed Engram storage are incomplete or unsupported. Its validation is tied to Ascend 950DT and CANN 9.2.0; the README says the published measurements do not establish support on other Ascend generations or CANN versions.

The next convincing test would run a publicly obtainable HDK and the released code on a specified Ascend 950DT configuration, reproduce the kernel cases, then measure a complete MoE training or inference workload against a well-described baseline. Report dispatch, combine, GEMM, network time, memory use and model-level throughput separately. Until that is possible, the release is a substantive view into DeepSeek’s Ascend engineering, with strong vendor-reported kernel numbers and a clearly documented reproducibility gap.

Last Update: October 1, 2026