Boosting GPU Utilization by 1.3x: ZStack AIOS Helps Enterprises Efficiently Unlock Heterogeneous Compute Powe

ZStack AIOS is a next-generation AI infrastructure operating system for enterprises.

Released Sep 24, 2026
Tag
Blogs

The number of GPUs in enterprises has increased, yet the situation of business waiting for compute power has not noticeably improved. 

This is becoming a real challenge for many AI infrastructure projects. Early-purchased NVIDIA GPUs are still in use, domestic GPUs are gradually entering data centers, and even within the same brand there are different generations and models. 

Resources are scattered across different servers, clusters, and teams. Some GPUs run at sustained high load while others sit idle; small inference or development tasks occupy an entire card, while new training and inference tasks still have to wait in queue. 

From the procurement ledger, the enterprise "has the cards"; from the business experience, the compute power available on demand remains insufficient. 

The question truly worth the enterprise's attention is: before continuing to purchase GPUs, have the existing resources been adequately organized and can they be supplied in time according to business needs?

01 / WHY LOW UTILIZATION 

Why is business still waiting for compute power when there are so many GPUs?

The coexistence of multiple brands and models increases the complexity of resource management, but the type of chip itself does not directly lead to low utilization. The problem usually occurs at three points in the resource supply chain.

The first is a mismatch between resource location and task location

Different brands, projects, or departments build their own resource pools, and resource ownership and usage boundaries remain fixed over time. One resource pool is fully queued while another still has idle capacity, yet the waiting tasks cannot directly use it. What the enterprise owns is multiple isolated local resource pools, rather than an integrated compute pool that can supply resources according to task needs.

The second is a mismatch between resource granularity and task requirements.

Large-scale training and high-concurrency inference require entire cards or even multi-card performance. Tasks such as development and debugging, small-model inference, Embedding, and Rerank often use only a portion of the VRAM and computing capacity. If the platform can only allocate whole cards, the resource specifications requested by tasks will consistently exceed actual demand. The card is occupied, yet its remaining capacity cannot be handed over to other tasks.

The third is a disconnect between monitoring and supply.

The fact that the platform can see temperature, load, and faults does not mean idle resources will automatically return to an allocatable state. In the absence of quota, priority, reclamation, and metering rules, GPUs still depend on manual allocation. To avoid future difficulty in requesting resources, teams tend to hold onto resources for long periods, further amplifying idleness.

Improving GPU utilization requires connecting the entire resource supply chain. 

The platform must not only grasp the location and status of resources, but also choose the appropriate delivery method based on the task, allow released resources to return to the pool promptly, and then continuously supply them through scheduling and operational rules.

02 / AIOS OVERVIEW 

ZStack AIOS: Boosting GPU utilization by over 1.3x

Through unified management of heterogeneous GPUs, on-demand partitioning, dynamic reclamation, and collaborative scheduling, ZStack AIOS  helps enterprises raise GPU utilization from around 30% under traditional approaches to over 70% — an improvement of more than 1.3x. 

The same batch of GPUs can support more training, inference, and development tasks, shortening the time business waits for compute power while making fuller use of the data center assets already invested.

ZStack AIOS is a next-generation AI infrastructure operating system for enterprises. 

Targeting the problems of resource fragmentation, whole-card waste, and low supply efficiency that arise after multiple brands and models of GPUs coexist, AIOS connects resource management, GPU virtualization and partitioning, multi-engine hosting, distributed scheduling, monitoring and operations, and resource operations in a single platform. 

Today, AIOS covers multiple GPU brands and more than 30 hardware form factors, with unified management capability for tens of thousands of GPUs. 

The platform supports multiple compute hosting environments, including virtual machines, containers, and bare metal, and provides resource delivery methods such as GPU passthrough, vGPU, MIG, dGPU, container whole-card, and VRAM partitioning. 

Once devices are onboarded to the platform, administrators can uniformly grasp GPU models, status, allocation relationships, and real-time load. 

Training, inference, or development tasks can obtain the appropriate GPU specifications and hosting method according to their performance, isolation, elasticity, and runtime environment requirements. Resources released by tasks return to the resource pool and continue to participate in subsequent scheduling.

What underpins this improvement is a technical system that runs through resource identification, specification provisioning, task scheduling, and day-to-day operations.

03 / UNIFIED VIEW

First, bring scattered heterogeneous GPUs into a single management framework

AIOS brings scattered GPUs into a unified resource view, centrally presenting device model, status, allocation relationships, load, video memory, temperature, and power consumption — providing a complete basis for resource allocation and scheduling.

Chips such as NVIDIA, Ascend, Hygon, and Alibaba PPU each retain their own software stacks, drivers, operators, and adaptation environments. Within its compatibility scope, AIOS manages these resources in a unified way, allowing administrators to see how different resource pools are being used and to select computing power that meets each task's requirements. Once device status and allocation relationships enter the same view, the differences in idle and busy states across resource pools can directly participate in scheduling, and idle resources return to business provisioning more quickly.

04 / GPU PARTITIONING

Second, let a single GPU provide resources according to task needs

Unified management solves the question of "where the resources are." To improve utilization, we also need to answer "what specification a task should receive." AIOS offers GPU passthrough, vGPU, MIG, dGPU, as well as container full-card and video-memory partitioning, each suited to different performance, isolation, elasticity, and hardware conditions.

GPU Passthrough — For large-model training, multi-card inference, or high-performance computing, GPU passthrough hands the physical device exclusively to a virtual machine, reducing virtualization overhead — ideal for scenarios where native performance is the priority. Here, exclusive full-card allocation is the sensible choice.

vGPU and MIG — When a business needs strong isolation or fixed specifications, native vGPU or MIG can be used depending on the GPU vendor and model. They differ in licensing, hardware support range, and partitionable specifications; the platform can offer the appropriate option based on the actual device conditions.

dGPU (CUDA API interception and forwarding) — For CUDA AI workloads inside virtual machines, AIOS's dGPU uses CUDA API interception and forwarding to dynamically create resources on supported NVIDIA GPUs according to video-memory specification templates. When a virtual machine starts, it obtains the required video memory; when it is shut down or unloaded, the video memory is immediately returned to the resource pool — there is no need to fix a card into halves, quarters, or eighths in advance. This approach suits development environments, small-model inference, Embedding, Rerank, and multi-project sharing scenarios, reducing long-term full-card occupation.

Container video-memory partitioning (granularity down to 1%) — In supported container environments, AIOS can also perform fine-grained video-memory partitioning with granularity as fine as 1%, letting a single physical card host multiple instances simultaneously.

Which approach to choose depends on the task's real requirements for performance, isolation, elasticity, and cost. Multiple resource forms reduce the gap between requested specification and actual usage, letting large tasks obtain full performance while allowing lightweight tasks to avoid occupying an entire GPU long term.

05 / SCHEDULING

Third, use multi-engine and collaborative scheduling to handle different tasks

Partitioning releases remaining resources within a card; scheduling determines whether those resources can be delivered to the business in time. Above multiple hosting environments — virtual machines, containers, and bare metal — AIOS matches resource states with task requirements.

Native-performance tasks: tasks requiring native performance can choose full-card or bare-metal resources. Environment isolation and development toolchains: tasks requiring environment isolation, development toolchains, or the Windows ecosystem can use virtual machines. Elastic inference and fast delivery: elastic inference and rapid-delivery scenarios can use containers with the corresponding GPU resource configuration methods.

In larger-scale environments, distributed and collaborative scheduling can combine resource availability, task specifications, and priorities to place tasks on compatible and appropriate nodes or resource pools. Queue and priority mechanisms also help enterprises organize workloads across different time periods — for example, guaranteeing online inference while using off-peak periods to carry training or batch-processing tasks.

Heterogeneous scheduling matches compatible resources around task requirements. Models, drivers, and software environments are still adapted according to specific chip conditions, while the platform is responsible for placing tasks on nodes or resource pools that meet the requirements. Enterprises can organize multiple kinds of computing power in a unified system while preserving the operating conditions required by different chips and workloads.

06 / GOVERNANCE

Fourth, sustain long-term sharing through monitoring, quotas, and metering

Once GPU sharing enters production, resource governance directly affects utilization.

Monitoring and alerting — AIOS provides a unified view of the allocation and running status of physical GPUs, vGPUs, and dGPUs, monitoring utilization, video memory, temperature, and power consumption. Administrators can set alert thresholds, push messages to enterprise collaboration platforms or Webhooks, and locate faulty hardware through device information. Resource anomalies are discovered more quickly, and chronically low-load devices enter recycling and optimization workflows more easily.

Quotas, permissions, and metering — Multi-tenancy, permissions, quotas, and metering further solidify the sharing rules. Different departments and projects receive clear resource quotas; the platform records usage and, when necessary, allocates scarce resources through approval and priority mechanisms. GPUs no longer depend on manual "card occupation" and ad-hoc coordination, and enterprises can carry out cost allocation and capacity planning accordingly.

This capability determines whether a sharing mechanism can run over the long term. Without monitoring and operational constraints, even the finest partitioning can fragment again.

07 / CUSTOMER CASE

Customer practice: letting existing computing power continuously enter real business

When building its AI intelligent-computing platform, a provincial university of finance and economics already owned servers, storage, and dozens of NVIDIA A40 and NVIDIA L20 GPUs. Different research groups and departments have different computing-power needs. Some tasks require full-card performance, while teaching experiments and lightweight applications are better suited to fine-grained resources. The school also needs real-time visibility into GPU status and load, and needs to connect model capabilities to campus applications through a standard inference API. How to make full use of existing equipment while meeting the differentiated computing-power needs of multiple business lines became the core problem the platform had to solve.

AIOS brought the existing servers, storage, and GPUs into a unified platform, providing GPU passthrough, vGPU, and video-memory partitioning according to scenario, and improving resource operations efficiency through load monitoring and alerting. On top of a unified computing-power foundation, the platform supports teaching, research, and campus applications through a standard inference API, transforming existing resources from scattered equipment into AI services that can be delivered continuously.

Before the next round of GPU procurement begins, first clarify one thing: is the existing computing power visible, partitionable, schedulable, recyclable, and measurable? AIOS turns scattered GPU devices into computing resources that can be continuously provisioned according to task needs, helping enterprises improve equipment utilization, accelerate resource delivery, and clarify costs and sustain operations in an environment where multiple brands and models coexist over the long term.

If your data center is also facing both uneven idle-and-busy resource distribution and business queueing at the same time, you are welcome to contact ZStack. We can start from your existing resources, task types, and provisioning methods to help you find the computing-power space that can still be released.

Ready to modernize your infrastructure?

Talk to our experts and see how ZStack can accelerate your cloud journey.

Most popular

Start Free Trial

Full-featured private cloud — single server free for one year, unlimited nodes for three months.

Start Free Trial
Evaluation

Schedule a Demo

See ZStack in action with a live walkthrough tailored to your use case and migration goals.

Request Demo
Resources

Get More Resources

Access white papers, migration guides, case studies, and technical documentation to plan your ZStack deployment.

Browse Resources