Watch on YouTube
Watch on Vimeo
GPUs represent a massive capital investment, yet enterprise AI initiatives frequently struggle with underutilized hardware, fragmented access, and low compute efficiency. Unlocking true return on investment requires hardware-aware virtualization and dynamic resource orchestration. In this technical deep-dive, we explore how VCF optimizes enterprise GPU infrastructure for production AI pipelines. We detail the architectural mechanics of vGPU pooling, fractional GPU sharing, dynamic resource allocation, and direct device assignment. Through live technical demonstrations, learn how VCF enables multiple training and inference jobs to share underlying physical hardware concurrently without performance degradation. Shawn Kelly, a principal architect at Broadcom focused on AI, opened by explaining that GPUs suit AI because their thousands of cores handle the massive parallel arithmetic that large language models require, though CPUs can still serve models under roughly 15 billion parameters. He described the two main ways to use GPUs on VCF: pass-through (DirectPath I/O or Dynamic DirectPath I/O), where the VM talks directly to the GPU, and NVIDIA vGPU, which adds a vGPU manager on the host to present virtual GPUs. He noted that under the covers both approaches are pass-through, with vGPU and Enhanced DirectPath I/O using “mediated pass-through” for setup and migration before handing work directly to the GPU. As a result, virtualization overhead is usually under 2 percent and sometimes a fraction of a percent, and in a few cases the ESX scheduler even outperforms bare-metal Linux, a trade-off he considered well worth the agility and lifecycle benefits.
Kelly said the gap between the two modes has narrowed now that pass-through supports time slicing and MIG, leaving vMotion as the main capability exclusive to vGPU, which also requires NVIDIA AI Enterprise licensing. In response to delegate questions, he recommended vGPU when live migration matters, such as for a researcher’s GPU workstation, and pass-through for model serving behind an AI gateway, where redundant Kubernetes nodes handle availability. He suggested bare metal only for training large models from scratch, since most enterprises do inference or fine-tuning, both well suited to virtualization. For models too large for one GPU, he covered NVLink bridges and the NVSwitch fabric in eight-GPU systems, which can hold models approaching 2.8 trillion parameters in a single host, and VCF device groups that let administrators split those eight GPUs into pairs, individual GPUs, or one large pool.
He then walked through fractional GPU options for VMs and VKS Kubernetes nodes, including time slicing, MIG (spatial slicing), and on newer Blackwell and RTX Pro 6000 cards, time slicing within a MIG slice. Time slicing gives each workload the full set of cores in turn while dividing memory into fixed shares, letting active workloads benefit when others are idle, though users can grow accustomed to borrowed performance and complain when it disappears. MIG carves the GPU into isolated slices with dedicated streaming multiprocessors, cache, and frame buffer, eliminating noisy neighbors and supporting multi-tenancy, and slices can be combined in different sizes, with different GPUs in the same host running different modes. Kelly closed with GPU dashboards showing temperature, memory, and compute utilization by cluster, host, and workload, noting that most customers he talks to run below 50 percent utilization and need this visibility to reclaim idle capacity. He added that AMD and Intel GPUs are supported in pass-through mode, but slicing and a vGPU equivalent for those vendors are still in development.
Personnel: Shawn Kelly
Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.
We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.
Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!