Watch on YouTube
Watch on Vimeo
Frontier AI has compressed threat timelines, allowing attackers to weaponize vulnerabilities in hours rather than months. For enterprise infrastructure, running production AI requires a fundamental engineering pivot: prioritizing continuous platform hardening over rapid feature delivery. In this session, we break down how VCF adapted its architecture to counter AI-driven threats. We walk through our security delivery model, combining monthly express updates with predictable quarterly releases to maintain zero-day resiliency without operational friction. We showcase how Broadcom uses Frontier AI internally to simulate multi-vector attacks against VCF before adversaries do. Finally, we demonstrate VCF’s hypervisor-level virtual patching, showing how inline network microsegmentation instantly insulates sensitive AI workloads from exploit vectors without disruptive downtime. Bob Plankers opened by describing how much the threat landscape has shifted since late 2025, when AI became capable enough to act as a force multiplier for attackers and defenders alike. He explained that AI excels at analyzing source code, which has driven a flood of CVEs against open-source projects like the Linux kernel and created supply chain pressure for vendors such as Broadcom. AI can also find configuration weaknesses, chain together low- and moderate-severity vulnerabilities that organizations never bothered to patch, produce flawless multilingual phishing, clone voices for social engineering, and generate custom malware for each victim, as seen with Brickstorm. He noted that AI-driven attacks can move very slowly to stay under detection thresholds, and that anyone with a modest GPU can run an uncensored open model to build a hacking bot.
Plankers urged customers to practice “security always” rather than claiming “security first,” to actually implement zero trust, and to avoid the survivorship bias of assuming no breach means their defenses work. He stressed that traditional controls such as patching, logging, identity management, isolation, and recovery remain effective, and that every VCF and vSphere feature maps to confidentiality, integrity, or availability. He highlighted snapshots as a recovery tool that has improved dramatically, with vSAN ESA using B-trees instead of delta files, consolidation no longer stunning VMs, and compression and deduplication keeping costs low. In his lab, the protection and recovery appliance stored more than a thousand snapshots of twelve VMs in about 25 GB, letting him patch weekly and revert quickly when something broke. He then covered live patching, which uses partial maintenance mode and a fast suspend-and-resume to swap the virtual machine monitor in microseconds without rebooting or evacuating hosts, while cautioning that large updates like 9.1.0 to 9.1.1 still require traditional vMotion and that change advisory boards remain a hurdle. VCF 9.1 also introduced a deprivileged user-level virtual machine monitor, so an attacker who escapes a VM lands with no permissions rather than as root.
On the organizational side, Plankers said Broadcom joined Anthropic’s Project Glasswing and its Mythos preview in April 2026, and that AI is now part of its secure development lifecycle, with different models and effort levels surfacing different findings. The VCF roadmap has been reoriented around security, with monthly express patches, quarterly releases like 9.1.1 that ship security features early rather than waiting for 9.2, and a commitment to avoid further major architectural changes in the 9.x series so customers can upgrade to 9.1. Broadcom signs its code, runs AI and conventional code reviews and scanners, pursues FIPS 140-3 and Common Criteria validation, and publishes a CVE disposition CSV and SPDX-format SBOM with each build so customers can reconcile scanner findings, though it keeps release notes vague to avoid educating attackers. In response to delegate questions, he said 9.1 hardens the management plane through process isolation and removing unneeded components, that VM escape vulnerabilities typically appear about once per major version and are less severe on 9.1, and that a predictable monthly patch date is a goal Broadcom is still working toward. He closed by comparing VCF to the star fort of Palmanova, Italy, built in response to the siege cannon, as a trusted place that can be defended when necessary.
Personnel:
Watch on YouTube
Watch on Vimeo
GPUs represent a massive capital investment, yet enterprise AI initiatives frequently struggle with underutilized hardware, fragmented access, and low compute efficiency. Unlocking true return on investment requires hardware-aware virtualization and dynamic resource orchestration. In this technical deep-dive, we explore how VCF optimizes enterprise GPU infrastructure for production AI pipelines. We detail the architectural mechanics of vGPU pooling, fractional GPU sharing, dynamic resource allocation, and direct device assignment. Through live technical demonstrations, learn how VCF enables multiple training and inference jobs to share underlying physical hardware concurrently without performance degradation. Shawn Kelly, a principal architect at Broadcom focused on AI, opened by explaining that GPUs suit AI because their thousands of cores handle the massive parallel arithmetic that large language models require, though CPUs can still serve models under roughly 15 billion parameters. He described the two main ways to use GPUs on VCF: pass-through (DirectPath I/O or Dynamic DirectPath I/O), where the VM talks directly to the GPU, and NVIDIA vGPU, which adds a vGPU manager on the host to present virtual GPUs. He noted that under the covers both approaches are pass-through, with vGPU and Enhanced DirectPath I/O using “mediated pass-through” for setup and migration before handing work directly to the GPU. As a result, virtualization overhead is usually under 2 percent and sometimes a fraction of a percent, and in a few cases the ESX scheduler even outperforms bare-metal Linux, a trade-off he considered well worth the agility and lifecycle benefits.
Kelly said the gap between the two modes has narrowed now that pass-through supports time slicing and MIG, leaving vMotion as the main capability exclusive to vGPU, which also requires NVIDIA AI Enterprise licensing. In response to delegate questions, he recommended vGPU when live migration matters, such as for a researcher’s GPU workstation, and pass-through for model serving behind an AI gateway, where redundant Kubernetes nodes handle availability. He suggested bare metal only for training large models from scratch, since most enterprises do inference or fine-tuning, both well suited to virtualization. For models too large for one GPU, he covered NVLink bridges and the NVSwitch fabric in eight-GPU systems, which can hold models approaching 2.8 trillion parameters in a single host, and VCF device groups that let administrators split those eight GPUs into pairs, individual GPUs, or one large pool.
He then walked through fractional GPU options for VMs and VKS Kubernetes nodes, including time slicing, MIG (spatial slicing), and on newer Blackwell and RTX Pro 6000 cards, time slicing within a MIG slice. Time slicing gives each workload the full set of cores in turn while dividing memory into fixed shares, letting active workloads benefit when others are idle, though users can grow accustomed to borrowed performance and complain when it disappears. MIG carves the GPU into isolated slices with dedicated streaming multiprocessors, cache, and frame buffer, eliminating noisy neighbors and supporting multi-tenancy, and slices can be combined in different sizes, with different GPUs in the same host running different modes. Kelly closed with GPU dashboards showing temperature, memory, and compute utilization by cluster, host, and workload, noting that most customers he talks to run below 50 percent utilization and need this visibility to reclaim idle capacity. He added that AMD and Intel GPUs are supported in pass-through mode, but slicing and a vGPU equivalent for those vendors are still in development.
Personnel: Shawn Kelly
Watch on YouTube
Watch on Vimeo
A private cloud delivers its economic promise only when IT teams actively manage capacity, rightsizing, and operational visibility. Without proper governance, overprovisioning and dark capacity can quickly erode the financial advantages of on-premises infrastructure. This practitioner-focused session breaks down how to build an operational FinOps practice natively within VCF. We step away from high-level ROI calculators and jump directly into the console to demonstrate real-world tokenomics management, capacity reclamation, automated rightsizing policies, workload placement optimization, and chargeback and showback frameworks. Architects and sysadmins will gain clear, actionable methods to eliminate cloud waste, optimize core density, and run a highly efficient private cloud platform that pays for itself over time. Kamesh Subramanian, who leads product management for VCF Operations, described a customer maturity path of assess, optimize, and plan. In the assess stage, a built-in cost engine calculates total cost of ownership using preloaded drivers for hardware, facilities, labor, licenses, and maintenance, which customers can adjust for their own discounts. Optimization turns reclaimable resources into dollar figures, and he recounted how a large airline got business units to give back unused capacity only after sending each leader a “shameback” bill showing half a million dollars wasted in a month. More mature organizations move to showback or chargeback with rate cards, which service providers can meter down to seconds, while planning covers infrastructure, workload, and migration scenarios translated into both capacity and cost, a need he said has grown with recent hardware price swings.
Subramanian explained that cost drivers roll up into per-unit base rates, so any service, whether a VM, Kubernetes cluster, database, or load balancer, can be priced much like a public cloud instance, and VCF Automation blueprints show their cost before deployment. Answering delegate questions, he said spending can be tracked against forecasts and budgets using the FinOps Foundation’s FOCUS specification and custom metrics at any organizational level, and that base rates are recalculated daily, with dynamic thresholds alerting administrators only when hardware changes shift costs meaningfully. He then turned to AI, noting that Broadcom is a founding member of the Linux Foundation’s Tokenomics Foundation. He broke tokenomics into layers: infrastructure teams maximizing GPU utilization, an emerging AI operator role that chooses providers and models, routes requests, and sets token or dollar limits, application teams driving consumption, and a value realization layer asking whether AI spending actually moves the business, which one charity customer measures by how many refugees it can help. A delegate called it the best tokenomics breakdown they had seen.
In a live demo, Subramanian showed total environment cost with potential savings, ad hoc analysis of VM and cluster costs over twelve months in more than a hundred currencies, and comparisons such as AI clusters versus traditional clusters. He cautioned that potential savings often differ from realized savings, since policies like required snapshot retention limit what can actually be reclaimed, and that cumulative savings figures can become a vanity metric. The demo also covered cost-driver breakdowns, reclamation opportunities by resource type, and chargeback views by organization and project, with quota and budget tracking through VCF Automation. On the AI side, he showed observability metrics such as time to first token and P95 latency tied to token consumption per model and per prompt, revealing that some internal experiments generating code from markdown files ran for hours and consumed enormous numbers of tokens. He finished with agent views that map agents, sub-agents, models, and tool calls to trace spending spikes, arguing that combining observability with token costs gives everyone, not just a FinOps team, the financial literacy the AI era requires.
Personnel: Kamesh Subramanian
Thank you for being part of the Tech Field Day community! Our mailing list is a great way to stay up to date on our events and technical content, and we appreciate your signup.
We promise that we’ll never spam you, send ads, or sell your information. This list will only be used to communicate with our community about our events and content. And we’ll limit it to no more than one message per week.
Although we only need your email address, it would be nice if you provided a little more information to help us get to know you better!