Per The Information, Zhipu AI is also (after DeepSeek) exploring a custom ASIC after GLM-5.2 usage reportedly jumped 27x in one week.
A custom ASIC removes flexibility, but it can cut power draw and per-token cost.
Nvidia GPUs are strong general-purpose machines, but inference at scale has different economics. A fixed model can run better on silicon designed around its own repeated operations.
Zhipu has not chosen a partner, and the project may take more than 2 years.
The pattern is now bigger than one Chinese lab or one model launch. Chinese AI companies are trying to make software, hardware, and deployment less separable.