Friday, August 7, 2026

VMware Deficiency - DRS Not Considering Network Throughput

We had a customer complaining about dropped packets.

So, we looked. Yup, dropped packets.

Why did  we not know? Well, apparently Aria Operations was not alerting on any dropped packets thresholds. But hey - there are a myriad of reasons that packets can be dropped, and some packets NEED to be dropped! So when you look at dropped packets, you really need to get deep into the statistics after an initial look at the situation. Are they sent (Tx) packets? Receive (Rx)? On all vmnics or just one? What protocols? And so on.

In this case, the packets being dropped were receive (Rx) packets.  From the uplink switch.

We then noticed that in this heterogeneous cluster we had,  DRS (VMware's Distributed Resource Scheduler) wanted to stack up a large number of VMs on one host. I think there were 44 VMs on that one host, and 24 on another, and then 3 hosts that had no workloads at all on them.

We had to examine host rules, affinity and anti-affinity rules, et al. We finally called in VMware. Turns out, "this is just the way DRS works". DRS - which I mention in past posts - no longer does a "balanced water level" approach to placement. It will pick the host it likes and as long as it meets a criteria, it will continue placing workloads on that host until no more meet the criteria in which case it will go to the next best host. This supposedly reduces the DRS overhead, reduces live migrations, et al.

But - DRS is ONLY LOOKING AT MEMORY AND CPU!!! IT IS NOT LOOKING AT NETWORK THROUGHPUT. 

 SO IF YOU HAVE 45 VMS TRYING TO INGEST PACKETS, THE NUMBER OF INTERRUPTS GENERATED OVERWHELMS THE RING BUFFER AND THE PACKETS GET DROPPED.

Now, having poll mode drivers might help this situation. We are not using poll mode drivers. 

So, the only way to fix this, is with rules to ensure too many of these VMs are not sitting on a single hypervisor on a single vmnic.

 

No comments:

NVIDIA GPU Xid 79 Lost Bus Connection Issue - Fixed It Appears

I didn't want to, but I finally disconnected the server, brought it onto the table and pulled the card. I did a full vacuum to get all d...