Alibaba Creates RISC-V Chip Capable of Running 27B AI Model With No GPU



Uploaded image As China continues on its warpath towards total semiconductor independence, Alibaba recently demonstrated its latest RISC-V chip running a 27B AI model with no GPU. What exactly did Alibaba demonstrate, and what does this indicate about the future of AI inference?

Alibaba Develops RISC-V Chip Capable of Running 27B AI Model

Recently, Alibaba demonstrated its 64-core XuanTie C950 RISC-V processor running the 27-billion-parameter Qwen 3.8 27B language model entirely on the CPU.

What makes this particularly interesting is that the system does not require a discrete GPU or an instruction-set translation layer to run the model. Instead, the processor has been designed from the ground up to accelerate the types of mathematical operations used by modern AI models.

The C950 is manufactured using TSMC's 5nm process and combines 64 RISC-V CPU cores with dedicated vector and matrix acceleration engines. These engines are designed to accelerate tensor operations, which make up a significant part of the computation required by large language models. According to Alibaba, the processor can achieve more than 30 tokens per second when running Qwen 3.8 27B, while achieving a time-to-first-token of around 1.9 seconds.

While this might not sound particularly impressive when compared with high-end GPUs, it needs to be considered in context. The entire 27-billion-parameter model is being executed on a CPU platform without additional GPU acceleration.

The C950 also includes dedicated acceleration for different numerical formats, including FP16, FP8 and INT4, alongside Alibaba's custom RISC-V matrix extensions. These allow the processor to perform the large number of lower-precision mathematical operations needed for modern AI inference much more efficiently than a conventional CPU.

However, Alibaba is not simply developing the processor.

The company is also developing the software and AI models that run on it, including the Qwen family of models. This creates an interesting hardware-software optimisation loop, where both the processor and the AI model can be designed with each other in mind.

In many ways, this resembles NVIDIA's approach of combining specialised hardware with its software ecosystem, except Alibaba is building its system around RISC-V, an open instruction set that allows companies to customise processor architectures without licensing the underlying instruction set itself.

What Does this Indicate for Future AI Inference?

One interesting observation that is becoming increasingly true is that engineers are slowly moving away from GPUs for AI inference.

While GPUs have become the king of AI computation, this isn't necessarily because GPUs are inherently the best architecture for AI. Rather, GPUs became dominant because they could perform massively parallel calculations far better than conventional CPUs.

But when dedicated AI accelerators enter the picture, things change.

NPUs and other AI accelerators can be designed specifically around the mathematical operations required by neural networks. Instead of building a processor that can perform almost anything, engineers can create hardware specifically designed to execute AI models as efficiently as possible. And this is exactly what we are now seeing with processors such as the C950.

Rather than using a CPU and then adding a GPU beside it, engineers are increasingly integrating AI-specific acceleration directly into the processor. This can reduce the need for separate hardware while also reducing the energy and latency associated with moving data between different processors.

This doesn't mean GPUs will disappear. GPUs remain extremely effective for AI training, where enormous amounts of parallel computation are required, and they are also useful for systems that need to run many different models and workloads. Their flexibility remains one of their biggest advantages.

But inference is different. Once an AI model has been trained, the priority often changes from raw computational power to efficiency, cost and latency. A dedicated AI accelerator can potentially perform the required calculations using far less power than a large GPU.

This could become particularly important for local devices. Running AI directly inside a computer, smartphone, vehicle or industrial machine means that every watt matters, and there is little reason to install a massive GPU if a specialised processor can perform the same task using a fraction of the energy.

But the C950 also demonstrates something else. China is not simply developing AI models or semiconductor hardware independently. It is increasingly developing the entire stack, from processor architecture and fabrication through to AI accelerators, software and the models themselves.

That could allow Chinese companies to build AI systems without relying on hardware or software ecosystems controlled by companies elsewhere in the world.

At the same time, open-source AI models can cross these boundaries. If engineers in the West can freely download, modify and train models developed in China, those models can become part of Western engineering projects regardless of where they originated.

This creates an interesting possibility where the AI ecosystem becomes increasingly separated from the traditional semiconductor ecosystem. Engineers could use Chinese-developed models on Chinese-developed processors, while other engineers use those same models on completely different hardware.

Ultimately, Alibaba's C950 demonstrates that large AI models do not necessarily need enormous GPUs to become useful. As specialised AI hardware continues to improve, we could see inference move away from the traditional GPU-centric model and towards CPUs, NPUs and highly specialised accelerators.

And if that happens, the most important processor in an AI system might not be the one with the most computing power, but the one that can perform the required computation using the least amount of energy.


You may also like

Robin Mitchell

About The Author

Robin Mitchell is an electronics engineer, entrepreneur, and the founder of two UK-based ventures: MitchElectronics Media and MitchElectronics. With a passion for demystifying technology and a sharp eye for detail, Robin has spent the past decade bridging the gap between cutting-edge electronics and accessible, high-impact content.

Avnet Silica IoT Podcast
Avnet Silica At The Edge
DigiKey
Avnet Silica At The Pulse