cd /news/ai-chips/deepseek-and-huawei-s-ascend-tools-a… · home › topics › ai-chips › article
[ARTICLE · art-145416] src=gladlabs.io ↗ pub= topic=ai-chips verified=true sentiment=· neutral

DeepSeek and Huawei's Ascend Tools Are a Software Bet, Not a Chip Story

DeepSeek and Huawei released open-source programming tools for Huawei's Ascend accelerators, announced on DeepSeek's WeChat account on Wednesday, including compute libraries, chip-to-chip communication libraries, and Ascend support for the TileLang high-level programming language, targeting the Ascend 950 generation. NextMSC counts the release as a suite of six tools, and the stated aim is to reduce reliance on Nvidia's software and hardware ecosystem amid ongoing export restrictions. The coverage frames the release as a software and toolchain play rather than a chip-specification story, since developers choose platforms based on how quickly their models reach advertised speed.

read9 min views1 publishedOct 5, 2026
DeepSeek and Huawei's Ascend Tools Are a Software Bet, Not a Chip Story
Image: Gladlabs (auto-discovered)

The GPU fans on a local box tell you things. When our writer model hands off to image generation, the card sits near 98% VRAM and the fans ramp hard, and you can hear whether the queue is healthy. Nothing in that sound has anything to do with the silicon’s peak spec. It’s about whether the software around the chip behaves.

That’s the lens I’d put on the news this week. DeepSeek and Huawei released open-source programming tools for Ascend chips, and the headline everyone ran was “less reliance on Nvidia.” The more useful reading is that this is a software story. Hardware doesn’t escape a monopoly until the software around it stops hurting.

What actually shipped #

Here is what the reporting supports, and no more.

According to Yahoo Tech’s summary of the original report, the release includes open-source libraries for AI computation and for chip-to-chip communication. It also adds Ascend support for TileLang, a high-level programming language. The stated aim is to reduce reliance on Nvidia’s software and hardware ecosystem.

SiliconReport says DeepSeek announced the toolkit on its WeChat account on Wednesday. It’s a free download, built in collaboration with Huawei, for programming Ascend accelerators. The same piece cites Bloomberg Law framing it as part of Huawei’s push to develop technology that replaces Nvidia’s.

NextMSC counts it as a suite of six tools tailored for Ascend, and ties the timing to ongoing export restrictions. The tooling targets the Ascend 950 generation, with the stated goal of making Huawei hardware easier to program and optimize.

That’s the whole verified surface. Three layers are named:

  • Compute libraries. The math kernels that do the actual work.
  • Communication libraries. The code that moves tensors between chips.
  • TileLang support. A higher-level way to write kernels without hand-tuning everything.

I haven’t run any of it. We don’t have Ascend hardware, and I won’t pretend to review code I haven’t touched. What I can do is tell you what these layers mean, because I’ve watched what happens when the equivalent layers are weak.

Why the moat is the toolchain #

A chip is a pile of math units and memory bandwidth. Nobody ships a product on that. You ship on the kernels, the compilers, the profilers, the collective-communication code, and the thousand small answers on forums when something segfaults at 2 a.m.

That’s the stack Nvidia has had a long head start on. Developers don’t pick a GPU. They pick the environment where their problem has already been solved by someone else.

So when a competitor releases hardware with good numbers, the question is never “is the silicon fast?” The question is how many days it takes to get your model running at something close to the advertised speed. If the answer is weeks, the silicon doesn’t matter.

Look at the three layers through that filter.

Compute libraries

Every serious model leans on a small set of operations: matrix multiplies, attention, normalization, activations. The first job of any new accelerator stack is making those fast, and an unoptimized kernel can leave most of a chip idle.

Open-sourcing the compute libraries matters because it lets outsiders read how those operations map to Ascend. You can see the tiling choices, the memory layout assumptions, the places the authors gave up and took a slow path. Closed vendor libraries hide exactly that.

Communication libraries

This is the layer people underestimate. Training at scale, and increasingly serving at scale, means many chips acting as one. How fast and how predictably they exchange data decides whether your cluster scales or stalls.

Chip-to-chip communication is where a platform either feels mature or feels like a science project. A shipped, open-source library here says the authors expect people to run multi-chip jobs outside Huawei’s own walls. I’d watch the issue tracker on this one more than any other.

TileLang support

A high-level kernel language is the lever for everyone who isn’t a vendor engineer. Writing a custom attention variant in raw low-level code is a specialist job. Writing it in a tile-oriented language is something a competent ML engineer can attempt on a weekend.

Reporting says Ascend support for TileLang is part of the release. If that support is real and keeps pace with new model architectures, it’s the piece that lets the community grow the kernel library without Huawei doing all the work. If it lags, it’s a demo.

Why DeepSeek is the interesting co-author #

This isn’t Huawei publishing a SDK alone, and the distinction matters. A hardware vendor’s tooling tends to be written by people who know the chip and have never had to ship someone else’s model on it under deadline.

DeepSeek is a model builder. Their incentive is different: they need the thing to run their workloads, and they have to care about the ugly path, not the benchmark path. SiliconReport describes the collaboration as a showcase of how close the two companies are. That closeness is useful for the tools. A model lab dogfooding the stack finds the bugs a vendor demo never hits.

It also carries a risk, and I’ll say it plainly. Tools shaped by one lab’s workloads can fit that lab very well and everyone else poorly. Whether the libraries generalize to a stranger’s architecture is something only outside users will find out.

What our own stack taught us about this #

Here’s where I’ll ground this in our own work, because the principle shows up even on a single consumer GPU.

We run content generation locally. At one point we considered switching our serving layer to vLLM. We ran the numbers against our real load, which is one to three concurrent requests per day, and the benchmarks said no. vLLM is tuned for far higher concurrency than we have. Ollama won on single-request latency and worked without workarounds. The silicon was identical in both cases. The software decided.

Then there’s the scheduling. Our task runner processes work sequentially on one GPU. When the writer model gives way to SDXL image generation, VRAM climbs to 98% and the queue stretches out. The card’s spec sheet says nothing about that. Our orchestration and memory handling decide it.

Here is the chart we keep coming back to, raw decode speed against what the application actually receives:

The gap between those two bars is the whole argument in miniature. Peak throughput is a property of the chip. Delivered throughput is a property of the chip plus every layer of software between it and your code. When a new accelerator vendor quotes a number, assume it’s the first bar. Your job is to find out what the second one looks like.

We’ve also been burned by version drift. We pinned the Prefect client to 3.6.29 because a mismatch caused silent failures in work-pool polling, with no crash and no error, just work that quietly didn’t happen. That’s the kind of failure a young toolchain produces at scale. On a mature stack, thousands of people have already hit it and written it up. On a new one, you’re the thousand.

None of that is Ascend-specific. It’s what adopting any less-traveled stack costs, and it’s the cost this release is trying to lower.

For more on why running models yourself changes how you weigh these tradeoffs, see our piece on why local LLMs became the backbone of development.

What this does and doesn’t change for you #

Let me be blunt about who this is for.

If you’re a developer or tinkerer running models on a 24GB consumer card in your office, this release does nothing for you this month. Ascend accelerators aren’t something you’re going to order and slot into a desktop. Your CUDA workflow keeps working.

If you work on inference engines, kernel libraries, or compilers, it’s different. You now have an open-source reference for a second large accelerator ecosystem. That’s worth reading even if you never deploy on it, because seeing how someone else solved the kernel mapping teaches you things about your own target.

If you work in or near China’s AI industry, the calculation is more immediate. NextMSC ties the release to ongoing export restrictions on Nvidia processors. When one vendor is partly off the table, a usable alternative stack stops being a nice idea and becomes a requirement.

What I’d check before believing the pitch

If you’re evaluating this seriously, don’t read the announcement. Read the repositories. Here’s what I’d look at, in order:

  1. Commit cadence after launch. One big drop followed by silence is a press release. Steady commits for months is a project.
  2. Issues from people outside the two companies. Who is filing them, and do the maintainers respond?
  3. Coverage of current model architectures. If the kernels only handle last year’s attention variants, you’ll be writing your own.
  4. Multi-chip behavior. Run a job across several chips and watch for stalls. Single-chip demos hide the hard problems.
  5. Documentation honesty. Good docs say what’s unsupported. Bad docs imply everything works.

None of those require trusting anyone’s benchmark. They require an afternoon and a willingness to be unimpressed.

The bigger pattern #

Open-sourcing is the right move for a challenger, and it’s the only move that works. A closed toolchain from a smaller player asks developers to bet their schedule on a vendor’s goodwill. An open one lets them fix their own problems. That’s how a second ecosystem gets built: by making it cheap for strangers to contribute.

It also changes the competitive shape. Nvidia’s advantage was never one library. It was the accumulated weight of years of other people’s fixes. You can’t copy that, but you can seed it, and publishing the compute, communication, and kernel-language layers together is a sensible seed.

Will it work? The honest answer is that nobody knows yet, and anyone who says otherwise is selling something. The reporting tells us what was released. It doesn’t tell us how the tools perform on your workload, how many bugs sit under the surface, or whether the community shows up. Those answers come from months of use.

Where I land #

This is a real release from two companies with strong reasons to make it work, and I’d treat it as a serious attempt rather than a gesture. Compute libraries, communication libraries, and TileLang support are the right three layers to open first.

But “serious attempt” isn’t “ready.” The gap between a toolkit existing and a toolkit being pleasant is where ecosystems are won or lost, and that gap is measured in the unglamorous stuff: silent failures, version drift, undocumented limits. We see small versions of it on one local GPU every week.

So watch the repos, not the headlines. If outside developers are shipping fixes in six months, the Nvidia story gets more complicated. If the commits dry up, it was a press release. Either way, the answer shows up in the issue tracker long before it shows up in a benchmark.

── more in #ai-chips 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-and-huawei-…] indexed:0 read:9min 2026-10-05 · —