All articles
AI

DeepSeek and Huawei: It’s Time for Cheaper AI

DeepSeek's open-source tools for Huawei Ascend could lower the cost of running AI. Will consumers see the savings? A closer look at benchmarks, falling API prices, and subscription bills.

Concept illustration for DeepSeek Ascend pricing: a chip on a green circuit board beside a receipt and coins
AI-generated concept illustration, not a photograph of a Huawei Ascend chip.
Also available in中文Español

My wish list for AI starts with something unglamorous: make it cheaper.

On September 30, 2026, DeepSeek announced open-source computing and communication tools for Huawei Ascend chips. Reuters reported the announcement, citing DeepSeek’s account of Huawei’s support for development and the companies’ joint work on an Ascend 950 system.

The DeepSeek Ascend tools are for teams running models on compatible hardware. As a consumer, my question is simpler: if a provider saves money, do I get to pay less?

Technical progress should show up on our bills, too. To make that case, though, we need to distinguish chip efficiency, the cost of running a service, and the price customers pay.

Where the DeepSeek Ascend tools could cut costs

A sentence typed into an AI app becomes a lot of computation on a server. Large models often run across several chips, so data must travel between them. If either computation or communication falls behind, expensive hardware can spend time waiting.

DeepGEMM-Ascend optimizes matrix multiplication and other basic operations that models perform repeatedly. Developers can call the implementations in the library instead of writing them from scratch. The project is currently developed and validated on Ascend 950 hardware.

DeepEP-Ascend handles communication between chips. For example, a model may route work to different “expert” modules spread across several processors. Data must reach those modules, and their results must come back together. This library makes those transfers more efficient.

TileLang helps engineers write the underlying compute programs, leaving some scheduling and hardware details to the compiler. It was already open source in 2025. The September 30 update added native support for Ascend 950.

Adapting software costs money as well. Nvidia’s CUDA toolkit includes compilers, compute libraries, and debugging tools. Moving to different hardware means checking the programs built around that environment. In the DeepSeek Ascend release, DeepGEMM-Ascend keeps the original DeepGEMM interfaces, leaving room to reuse existing code. Whether a complete model runs reliably still needs separate testing.

These tools offer ways to reduce computation, communication, and adaptation overhead. Measuring the savings takes a deployment: with comparable answer quality and response requirements, how many resources does it take to complete a batch of tasks? Open-source code alone does not provide that cost breakdown.

95% of peak performance does not set a subscription price

One eye-catching figure comes from DeepSeek’s FlashMLA technical write-up. The project reports that, under typical workloads for its V4.1 model, an Ascend sparse-attention kernel reaches 95% of the tested hardware’s theoretical peak while processing input.

The reference point is that chip’s own theoretical peak, and the measurement covers a particular computation. Producing a complete answer involves other steps. Providers also pay for machines, electricity, and operations. The result makes this optimization worth examining, but it does not tell us how much cheaper the whole service becomes.

Even downloading the code does not guarantee the same result. The DeepEP-Ascend environment notes say its published performance figures used a proof-of-concept HDK supplied to DeepSeek, with additional manual configuration. That configuration is not a publicly distributed release. The documentation also lists unfinished and experimental features.

Those deployment conditions belong alongside the headline numbers. The documents do not provide a comparison of total service costs at equal answer quality and response requirements. Predicting a price cut from one benchmark gets ahead of the evidence.

AI prices have already been falling

Calling for cheaper AI should not erase the price cuts that have already happened.

Stanford HAI’s 2025 AI Index recorded a striking change. For models meeting GPT-3.5’s performance threshold on the MMLU knowledge benchmark, the price of a million tokens fell from US$20 in November 2022 to US$0.07 in October 2024, roughly 1/286 of the earlier price.

Tokens are the units models use to process text; an API is the interface through which other applications call a model. These figures compare API prices at a fixed benchmark threshold. Chapter 1 of the report weights input and output prices 3:1. It measures how cheaply that level of tested capability became available. It is not a price history for one subscription, nor evidence that all those models work equally well on every task.

API pricing at the GPT-3.5 MMLU threshold fell from US$20 per million tokens in November 2022 to US$0.07 in October 2024, about 1/286 of the earlier price.
Historical data | API prices at the GPT-3.5 MMLU threshold (64.8%), with input/output prices weighted 3:1. These are two reported observations, not one subscription’s price history. Source: Stanford AI Index 2025, Chapter 1, Figure 1.3.22; underlying data from Epoch AI and Artificial Analysis.

A more recent example appears in DeepSeek’s own September 10 changelog: the company announced lower API prices with the launch of V4.1 Flash. That predates the September 30 open-source announcement. The available documents do not establish how the tools contributed to those costs, so we cannot simply credit that price cut to the DeepSeek Ascend release.

Price competition is already happening. The next question is whether ordinary users can buy those increasingly affordable capabilities in a product that suits them.

Why a cheaper API might not mean a cheaper subscription

Developers generally pay for API usage. In DeepSeek’s pricing table, charges depend on inputs, outputs, cache hits, and time of use. An AI app subscription may bundle model access with document processing, storage, and other features. Those are different things to price.

Consider a hypothetical bill: a service costs $100 to operate, of which model calls account for $30. Halve the cost of those calls, with everything else unchanged, and the total becomes $85. One item got 50% cheaper; total costs fell 15%. These are illustrative figures, not any company’s actual accounts.

Hypothetical costs: $30 in model calls plus $70 in other costs total $100. Halving model-call costs leaves $15 plus $70, or $85: a 15% total reduction, not actual company data.
Hypothetical example | Model calls fall from $30 to $15; other costs stay at $70. The total drops from $100 to $85. These figures explain the arithmetic, not DeepSeek’s or any other company’s accounts, and do not forecast subscription prices.

Pricing comes next. A business can lower its monthly fee, raise usage limits, fund new features, or keep more profit. Without its cost and plan details, an unchanged fee does not prove it has pocketed every saving. Equally, “the model is smarter” does not establish that existing customers are getting better value for their everyday work.

My hope for the DeepSeek Ascend tools is that they help teams offer more competitively priced services on Huawei hardware. When users can switch to a cheaper product of comparable quality, other providers face more direct pressure on prices.

That is a judgment about future competition, to be tested against actual products and offers. The DeepSeek Ascend announcement alone tells us nothing about the future prices of platforms that do not use these tools.

Smarter AI can command a premium. AI that is already good enough should come with affordable options. Someone who wants to tidy up a few emails or organize some documents should not have every conversation about value redirected to the latest flagship model’s benchmark score.

Figure and sourceHow to read it
FlashMLA: 95% of theoretical peak
Reported by the project
One computation approaches the tested chip’s theoretical peak.
This does not measure a whole-service saving or a subscription price cut.
AI Index: US$20 → US$0.07
Per million tokens, Nov 2022 → Oct 2024
API pricing fell at a fixed MMLU threshold.
This compares different models, not one subscription, and predates the Ascend release.
Example costs: $100 → $85
Hypothetical arithmetic
Halving model-call costs, initially 30% of the total, saves 15% overall if other costs stay unchanged.
These are not company accounts or a price forecast.
What each figure actually tells us

Price the job that actually gets done

For consumers, the useful comparison is what it costs to finish the same task.

Take a long document. Does the plan support that much text, or must you buy extra usage? If the result needs another attempt, does the retry cost more? Your actual monthly spend, plus the time spent checking and fixing the output, gets closer to the real cost than an advertised “from” price per call. That was also my concern in the earlier Claude Mods article: how much work is still left for the person after delivery?

A subscription that completes more tasks to an acceptable standard for the same monthly fee can be better value. But usage limits, overage charges, and paid extras should be clear before purchase. Showing only the lowest possible number leaves customers to discover the conditions halfway through a job.

The next time a release promises a major efficiency gain, I will turn to the pricing page first.

Sources checked October 6, 2026. News details are based on Reuters and project documentation; historical pricing draws on Stanford’s AI Index. Reuters reported company statements and did not independently test the hardware. I have not benchmarked Ascend hardware; performance figures are the project’s claims. The cost example is hypothetical, and the consumer-pricing argument is the author’s analysis.