DeepSeek released DeepGEMM-Ascend on September 30, 2026, bringing its matrix library to Huawei Ascend 950. TileLang also added an official Ascend 950 backend that day. Together, they offer familiar library interfaces and a higher-level way to write accelerator kernels. Deployment still requires the matching Huawei software environment. DeepGEMM release notes and TileLang announcement.
The useful question for developers is how much of an application can stay familiar when its accelerator changes. There are two routes here: call an optimized library, or write a custom kernel through TileLang. Those routes solve different amounts of the migration problem.
DeepGEMM preserves the interface; data preparation still matters
DeepSeek describes the Ascend package as API-compatible with DeepGEMM, supporting BF16, FP8 and FP4 matrix operations. Huawei receives explicit credit for engineering support. Library documentation.
There is a specific migration detail worth checking: quantization scaling factors use a different representation from NVIDIA’s implementation. Pairs of UE8M0 factors along K are packed into an int16 and stored in MN-major order. Interface note.
That distinction gives application teams a concrete review point. A familiar function signature does not establish that previously prepared tensors have the right representation for another backend. Check the conversion boundary before interpreting a numerical mismatch as a kernel failure. Preserve a reference result and compare the same inputs after the required preparation.
TileLang’s new backend is distinct from the older adapter
The main TileLang repository dates official Ascend 950 support to September 30, with native code generation, automatic scheduling and synchronization, and SIMD/SIMT vector programming. Its backend acknowledgements identify DeepSeek developers as the main initial contributors and thank Huawei for collaboration. Ascend 950 backend guide.
An earlier, separate TileLang-Ascend adapter became open source on September 29, 2025; its README lists A2 and A3 as tested devices. This week’s announcement adds a specific new backend in the main project. Describing all Ascend support as newly created would lose that distinction.
The new backend reuses TileLang’s frontend while adding Ascend-specific compilation. Developers still choose shapes, types, tile sizes and block counts. The compiler can handle task scheduling, buffer rotation and synchronization. That division leaves performance-sensitive choices visible while reducing low-level coordination code. Programming responsibilities.
Fusion shows the programming benefit
The backend guide’s GEMM-plus-ReLU example sends matrix multiplication to Cube cores and the activation to Vector cores. Direct transfers move result tiles between those stages, allowing the epilogue to execute within the same kernel. The example uses BF16 inputs and FP32 accumulation and output; it requires dimensions divisible by its tile sizes. Documented example.
For a team evaluating custom operators, this is a useful starting point: verify the published example, then replace its dimensions and epilogue with the actual workload. Include edge shapes in the next stage of testing. A demonstration that explicitly excludes partial tiles cannot establish correctness for arbitrary dimensions.
TileLang powers DeepGEMM-Ascend’s mHC prenorm kernel; the acknowledgement does not describe a library-wide implementation. DeepGEMM acknowledgements.
The 99.8% result has a defined scope
DeepSeek reports 431 TFLOPS against a stated 432-TFLOPS limit for a BF16 dense GEMM with M = 4096, N = 7168 and K = 16384. The published 99.8% figure uses Ascend 950DT, CANN 9.20, bench_msprof and cold L2. It describes that kernel case. Benchmark conditions.
Whole-model improvements depend on where time was spent before the change. Consider a hypothetical workload in which GEMM accounts for 60% of execution time. Even doubling that part would reduce total time to 0.60 ÷ 2 + 0.40 = 0.70 of the original, or about a 1.43× overall speedup. This is an illustrative calculation, not a measured DeepGEMM result.
The same reasoning applies to fusion: fewer launches may help, but the benefit depends on launch overhead, data movement and the amount of work fused. Measure those components before using a kernel headline to forecast serving capacity.
Start with the exact environment and a reference output
DeepGEMM’s documented setup requires Ascend 950, CANN 9.20, torch_npu, Python 3.10 or later, and C++20 support. Requirements. TileLang’s Ascend build additionally needs compatible driver/runtime tooling and a configured compiler environment. Its guide permits an Ascend build with CUDA disabled. Installation instructions.
Evaluate one representative operator first. Record software versions, matrix dimensions, precision, numerical tolerance and input preparation. Separate compilation time from repeated execution. Then run the complete application with identical batching and quality requirements; report both throughput and latency.
For mixture-of-experts workloads, communication is another substantial variable. Klaus Fischer’s EyesTech coverage of DeepEP-Ascend explains the accompanying dispatch libraries and the reproducibility constraints documented for their benchmark setup.
DeepGEMM-Ascend is released under the MIT license. Teams can inspect and adapt the implementation subject to its notice requirements. The immediate engineering opportunity is to measure whether familiar interfaces and compiler-managed coordination reduce porting work on a real Ascend deployment. A reproducible application-level result would establish how much that opportunity translates into usable capacity.
