At the World Artificial Intelligence Conference 2026 (WAIC 2026), many businesses said that reducing token processing costs (the basic data unit that AI processes) is becoming a common goal of the entire AI industry chain. From chips, computing infrastructure to AI and smart routing models (automatic AI model selection systems suitable for each task) are being optimized to improve efficiency and promote commercialization.
According to Xinhua, at WAIC 2026 and the Global AI Governance Summit, many businesses said that reducing token processing costs is becoming a focus as AI Agents significantly increase operating costs.
In the field of AI chips, tensor processors (TPUs) are assessed to have much better energy and cost efficiency per token than graphics processors (GPUs) in large-scale AI training and inference tasks.
Yang Gongyifan - founder of TPU chip company Zhonghao Xinying - said that the company's goal is to halve the cost per token in the next generation chip and continue to reduce it in subsequent generations to promote AI commercialization.
In computing infrastructure, Liu Haifeng - Chairman of the Board of Directors of the Chinese Academy of Sciences (CAS) - said that electricity costs account for more than 20% of the total life cycle costs of computing centers. According to him, combining wind power, solar power and energy storage can significantly reduce operating costs.
AI model developers also focus on reducing operating costs. Tencent said the Hunyuan HY3 AI model has performance equivalent to many larger models but significantly lower usage costs. Meanwhile, Meituan said LongCat-2.0 is the first model in the industry to apply a free mechanism for tokens stored in cache. These are data that has been pre-processed by AI and saved for reuse when encountering similar requirements, helping users not have to pay multiple processing costs.
According to the AI model evaluation organization SuperCLUE, the Kimi K3 model achieved a score 18% higher than the model ranked second, while the average cost to complete each task is about 16% lower. This result shows that new AI models not only improve processing capabilities but also increasingly save usage costs.
In addition, the "smart model routing" technology is also becoming a topic of interest at WAIC 2026. According to PPIO company, this technology can select an AI model suitable for each task, thereby improving processing efficiency by about 20% and reducing token costs by 50-60%.
