The Information : China’s AI Lab Zhipu Weighs Custom Chip As Demand for its GLM

China’s AI Lab Zhipu Weighs Custom Chip As Demand for its GLM Model Soars

The Takeaway
  • Zhipu AI explores custom chip design due to soaring demand and U.S. export controls.
  • U.S. export controls and blacklist force Zhipu to seek Chinese chip partners.
  • GLM-5.2 model sees 27x token usage surge, straining Zhipu’s compute resources.

Zhipu AI, the Chinese AI lab behind the highly regarded GLM series of open-source AI models, is weighing designing its own AI chip as surging demand and U.S. export controls make computing resources a growing constraint, according to three people with direct knowledge of the plan.

The Beijing-based company, one of China’s leading AI labs, recently made preliminary inquiries with some Chinese chip design houses about the possibility of working on a bespoke AI processor optimized for running its models, according to two of the three people. The discussions are still at an early stage, and Zhipu has not selected a designer, according to the two people.

Zhipu currently relies on a combination of chips from Chinese tech giant Huawei, other locally made chips and some Nvidia chips. The company in January released an image-generation model trained entirely on Huawei chips, the first major image model to use only Chinese chips for training.

Still, Huawei’s chips have their own constraints. U.S. export controls continue to limit access to the advanced equipment Chinese foundries need to produce Huawei chips, while Zhipu must do extra software engineering work to run its models efficiently on Huawei’s hardware.

Zhipu, also known as Z.ai, is focused solely on finding a design partner in China because of its designation on a U.S. blacklist, two of the three people said. The U.S. Commerce Department has put Ziphu on a list of foreign companies prohibited from procuring U.S. technologies. If the project moves ahead, Zhipu is also likely to manufacture the chip at Chinese foundries, the two people added.

That blacklist doesn’t prevent U.S. companies from using Zhipu’s models, which are increasingly popular. Coinbase CEO Brian Armstrong, for instance, recently said in an X post that Coinbase was experimenting with GLM-5.2 as part of a broader effort to save money on AI costs. GLM-5.2 is also available on Oracle Cloud Infrastructure, allowing enterprise customers to deploy the open-weight model on Oracle computing clusters.

While anyone can download an open-source model for free and run it at their own cost, for AI application developers, the benefit of procuring access from the model maker directly is that they can keep the expense of running the models low as the model maker shoulders the bulk of the compute costs.

But the rising demand has strained Zhipu’s computing resources. On U.S. startup Vercel’s model aggregator platform, for instance, GLM-5.2 has been the fastest-growing model since its release last month, with daily token usage surging as much as 27 times during the first week of launch.

Zhipu, founded by researchers at the prestigious Tsinghua University in 2019, has been trying to generate more revenue from selling enterprise customers cloud-based access to its AI models, while reducing its reliance on deploying models for customers at their own data centers. This is because such on-site deployment usually generates just one-off revenue, because once a model is installed and up and running at a customer’s data center, little additional service will be required from the model maker. By comparison, selling cloud-based access lets the model maker charge for revenue on a continuous basis,

A custom AI processor would be a longer-term solution rather than a quick fix for Zhipu’s computing crunch. The company would need to build or expand a semiconductor team, choose a design partner, test the processor and adapt its software before its models could run on the chip. The whole endeavor could take more than two years.

Zhipu’s chip plan mirrors similar pushes by leading AI developers in the U.S. and China. These custom AI chips, known in the industry as application-specific integrated circuits, are processors designed to perform specific tasks tailored to specific models, as opposed to general-purpose AI chips like those Nvidia produces.

Major AI developers, including Google, OpenAI, ByteDance and Alibaba, have come up with their own custom chips to wean themselves off outside suppliers and reduce the costs of running their own models. The latest example is OpenAI’s Jalapeño, which it developed together with Broadcom and will deploy to run its GPT models by the end of this year. Other popular ASICs include Google’s tensor processing units and Amazon’s Trainium.

For Zhipu and other Chinese AI developers, U.S. export restrictions on advanced chips are another incentive to build their own bespoke chips.

Zhipu was the world’s first large language model developer to go public, and its stock price has soared more than 100 times since its debut on the Hong Kong stock exchange in January. With a market capitalization of around $100 billion, the company is planning a dual listing in the Shanghai Stock Exchange’s tech-heavy Star Market.

Other custom chip efforts by Chinese companies include Alibaba, whose T-Head unit has developed processors for running and training AI models, and ByteDance, which has pursued several custom AI chip projects to support its AI infrastructure. Baidu’s Kunlunxin chip arm, which grew out of the company’s decade-old internal AI hardware work, is targeting a $50 billion valuation for a Hong Kong listing, The Information reported.