Bolun Zhihui Raises Angel Funding for Heterogeneous AI Clusters

A new Beijing start-up is betting that the next bottleneck in Chinese artificial intelligence will not be the supply of any single processor, but the difficulty of making many different processors work together. Bolun Zhihui, an AI infrastructure company established in April, has completed a tens-of-millions-of-yuan angel round, according to Zhidx. The company says it will use the funds for core technology iteration, product work, and team expansion as it pursues a market for heterogeneous-compute scheduling and AI inference services.

The round is small by the standards of China’s headline-grabbing model companies and chip listings. Its focus is nevertheless revealing. Enterprises may own or lease GPUs from several generations and vendors. A cluster can include Nvidia hardware alongside domestic accelerators, each with different software support, memory behavior, and suitability for a given model. The challenge is to convert that mixed equipment into a dependable service instead of leaving capacity fragmented across incompatible systems.

Bolun Zhihui says it has developed a Token-Oriented Large-scale Distributed Architecture, known as TOLD, to address that problem. Zhidx reported that the company has tested the system on a 10,000-GPU heterogeneous cluster combining Nvidia and domestic chips. That testing claim comes from the company and should not be treated as an independent measure of performance. It does, however, show the scale of the operating environment Bolun believes it must handle if it is to win enterprise customers.

AI Infrastructure Is Becoming a Scheduling Problem

The first image many people have of AI infrastructure is a warehouse full of powerful accelerators. That picture misses a growing operational problem. A model may run efficiently on one processor but require a different software stack on another. Some hardware may be suited to high-throughput batch work, while other equipment is needed for requests that demand rapid responses. Capacity can be available on paper but difficult to use efficiently in practice.

Bolun Zhihui’s proposed approach is to treat tokens as the unit of service. Inference providers ultimately sell or allocate model output, and the system must decide which resources should process each request. A scheduling layer could, in principle, assign work based on a model’s requirements, the hardware available, and the service level a customer expects. That does not eliminate the work of adapting models to new chips, but it can make a mixed environment easier to operate.

Zhidx said the company offers two principal business lines: private deployments for enterprises with their own compute resources and self-operated token factories that provide inference capacity. The difference matters. A private deployment focuses on helping a customer use equipment it already controls. A token-factory model turns underlying compute into an external service. Both depend on reliable orchestration, model compatibility, and a clear understanding of the cost of each workload.

The theme has become more prominent as China develops a more diverse hardware base. EastFrontier’s recent report on Mucang’s Series B examined a related infrastructure layer: high-speed network chips for large AI clusters. Networking determines whether machines can exchange data quickly enough. Scheduling determines whether the available machines are assigned work intelligently. A workable AI stack needs both.

A Young Company Enters a Crowded Infrastructure Race

Bolun Zhihui is a very young company. Zhidx said it was founded in Beijing in April 2026. Its CEO, Liu Jinzhi, previously worked at Intel and Motorola, while CTO Yang Guang has worked at Siemens, IBM, and Oracle, according to the report. The company also says its team includes people with backgrounds at Huawei, Tianshu Zhixin, and other technology businesses. Those biographies help explain the company’s focus on cloud-native systems, AI infrastructure, and heterogeneous hardware, but they do not yet amount to a proof of commercial scale.

The start-up’s need is clear. China’s AI market is producing more model providers, more specialized accelerators, and more enterprises that want to deploy AI without rebuilding their systems for every new chip or model. The economic opportunity lies in reducing the friction among those layers. A company that can make a cluster more useful may offer value even if it does not manufacture a GPU itself.

There are also hard limits to what software orchestration can achieve. No scheduling platform can turn a weak processor into a leading accelerator, compensate for insufficient memory, or instantly create mature libraries where they do not exist. It must work within the performance and compatibility of the hardware it manages. That is why the startup’s claim of testing across Nvidia and domestic GPUs is noteworthy but incomplete. Customers will want evidence that the platform works across their own models, equipment, security requirements, and traffic patterns.

The funding round reflects an investor view that this engineering work may become more important as China’s infrastructure becomes more heterogeneous. EastFrontier’s story on Enflame’s new Shanghai IPO timetable showed how domestic accelerator makers are seeking capital to grow. If more domestic chips reach customers, the need for systems that help them coexist with existing hardware will rise as well.

From Cluster Testing to Repeatable Customer Use

Bolun Zhihui’s immediate task is to turn an architecture into a product that customers can trust. That means showing that its platform can schedule work accurately, protect data, handle failures, and provide usable tools for model deployment. Enterprises will compare its offer not only with competing start-ups, but also with cloud providers, hardware vendors, and internal engineering teams that may build their own orchestration layers.

The firm is also working on an entry-level intelligent-hardware product, according to Zhidx. That ambition points toward a broad view of AI infrastructure that includes cloud services, local devices, and the pathways between them. It may eventually create more opportunities, but it also adds execution risk for a company still proving its core platform.

The angel funding gives Bolun Zhihui time to build. It does not settle whether its TOLD architecture will become a widely used layer in China’s AI infrastructure. The company’s advantage, if it can establish one, will come from making a complex, mixed-GPU environment simpler for customers to operate. As AI services move from one-off model experiments toward continuous inference workloads, that kind of operational expertise may become a valuable part of the country’s AI stack.