ZStack AIOS Completes Full-Stack Adaptation for Ascend 910C! A Major Upgrade for the High-End Computing Ecosystem

As large model applications accelerate into production, enterprise demand for co

Released Sep 11, 2026
Tag
Blogs

As large model applications accelerate into production, enterprise demand for computing power is diversifying across performance, scale, and delivery models — and AI infrastructure is steadily moving toward heterogeneous computing running in parallel. Whether different chips can be quickly integrated into a unified platform while continuing to leverage existing model deployment and operations systems is becoming a key factor in how enterprises assess the long-term evolution capability of their AI infrastructure.


01 / FULL-STACK ADAPTATION

910C Scales into Production: ZStack AIOS Completes Full-Stack Adaptation


As Ascend's new-generation high-end AI chip designed for large model training and inference, the 910C is entering large-scale deployment alongside products such as the Atlas 900 A3, and has become one of the most closely watched mainstream options in China's domestic AI computing market.


As domestic computing power moves from project validation to production deployment, enterprise attention to the 910C is also extending from raw performance to platform adaptation, model compatibility, and sustained operations. Once a chip enters the data center, drivers, runtime environments, inference frameworks, and models must be matched layer by layer, and computing resources must be integrated into unified monitoring, inspection, and fault localization systems. Whether new hardware can smoothly integrate into an existing platform and quickly deliver stable model services is becoming a key criterion for enterprises evaluating and adopting the 910C.


In response, ZStack AIOS has completed adaptation for the Huawei Ascend 910C, bringing it into a unified process spanning computing resource management, model deployment, inference services, and operations. This adaptation covers the full chain from hardware onboarding to model service launch, focusing on three key questions: Can computing power be managed in a unified way? Can models run smoothly? Can business services continue to connect without disruption?


Computing power under control: Unified onboarding of 910C resources


Once new chips enter the data center, the first challenge is resource visibility and unified management. If device status, resource utilization, and upper-layer inference services are scattered across different tools, administrators must switch between multiple interfaces, lengthening the workflow for resource allocation, routine inspection, and troubleshooting.


ZStack AIOS presents the operational status of 910C devices and inference services in a unified view, and incorporates them into existing systems for resource allocation, status monitoring, inspection, and anomaly localization. Platform administrators can monitor newly added NPU resources and the services they carry from a single view, without creating a separate hardware management console. When NVIDIA GPUs, Ascend NPUs, and other types of computing power coexist, enterprises can continue using one management system — avoiding further fragmentation of resource views and operations workflows as hardware types multiply.


Models that run: Unified delivery of runtime environments and frameworks


Hardware onboarding alone does not mean models can run directly. Drivers, runtime environments, inference frameworks, and models must form matching version combinations. If every project has to prepare and verify these on site, deployment cycles lengthen and the cost of troubleshooting environment inconsistencies grows.


ZStack AIOS delivers the relevant runtime environments and inference frameworks together with the platform. Through a unified entry point, users can select models, inference frameworks, and 910C or other computing resources to create and manage inference services — without assembling a deployment environment themselves. The platform also supports offline deployment for enterprise sites with strict network and data security requirements, reducing repetitive work such as component installation, dependency preparation, and version verification, and shortening the path from 910C resource onboarding to model service launch.


Business that stays connected: No re-adaptation needed for upper-layer applications


Enterprise applications typically call underlying computing power through model services. If a new chip means new service endpoints and access methods, hardware changes propagate upward — triggering extra work such as interface modifications, integration testing, and application releases, and widening the range of teams and systems affected by a single computing-power change.


In ZStack AIOS, after underlying computing power changes, upper-layer applications continue to access model services the same way as before. Application teams don't need to rework interfaces with every chip change, and platform administrators, model deployment engineers, and application developers can keep collaborating under their existing division of responsibilities. Enterprises can contain computing-power changes within the platform, maintaining business continuity while smoothly bringing the 910C into their existing AI business systems.


02 / WORKLOAD VALIDATION

From Onboarding to Inference: ZStack AIOS Completes 910C Large Model Workload Testing


After bringing the 910C under unified management, ZStack AIOS further validated — through real model workload testing — whether the complete chain from runtime environment and inference framework to model service works stably, so that the 910C's inference capability can be effectively hosted and invoked on the platform.


This test used the DeepSeek-V4-Flash model. Under specific environment and workload conditions, 16 logical NPU devices achieved a peak output throughput of 5,587 tokens/s in high-concurrency testing, along with 64K long-context workload testing. The results show that ZStack AIOS has fully connected the chain from 910C resource discovery and runtime environment delivery to large model inference services, and can sustain inference workloads featuring high concurrency and long contexts. After adopting the 910C, enterprises can carry out model deployment and business validation on an already-adapted platform chain, reducing the integration work between hardware onboarding and inference service launch.


*Note: The results above were obtained under specific test environments and workload configurations. Actual performance may vary depending on model version, quantization method, concurrency strategy, input/output lengths, and software/hardware configurations.*


03 / PLANNING CHECKLIST

Four Questions to Answer Before Planning Enterprise AI Infrastructure


When preparing to move large models from pilot to production, or planning a domestic or hybrid computing platform, enterprises should confirm the following in advance:


01 — Can existing and newly added GPUs and NPUs be managed and operated in a unified way?

02 — Can different models, inference frameworks, and runtime environments be delivered quickly?

03 — After underlying computing power changes, do upper-layer applications require re-adaptation?

04 — When new chips and models are added in the future, will another platform need to be built?


ZStack AIOS hosts heterogeneous computing management, model deployment, and inference services on a unified platform, helping enterprises connect the full chain from resource onboarding to business invocation while maintaining continuity in both management practices and application interfaces.


If these questions are already on your AI infrastructure roadmap, contact ZStack — we'll help you evaluate the right AI infrastructure plan based on your existing computing power, models, and business scenarios.


Ready to modernize your infrastructure?

Talk to our experts and see how ZStack can accelerate your cloud journey.

Most popular

Start Free Trial

Full-featured private cloud — single server free for one year, unlimited nodes for three months.

Start Free Trial
Evaluation

Schedule a Demo

See ZStack in action with a live walkthrough tailored to your use case and migration goals.

Request Demo
Resources

Get More Resources

Access white papers, migration guides, case studies, and technical documentation to plan your ZStack deployment.

Browse Resources