Artificial Intelligence Technical Community Group

HAMi: Optimize Your Heterogeneous AI Cluster

Capacity: 100
virtual
Event date
Sep 29, 26
01:00 PM - 02:00 PM +08
Registration is open until Sep 29, 2026 at 12:59 PM +08.
Location
Virtual event
About this event

GPU jobs on Kubernetes usually hit two opposite problems. Large jobs need topology-aware placement: a poor mapping across interconnects cuts GPU-to-GPU bandwidth and hurts training or inference performance. Small jobs need device sharing so one accelerator can be split across workloads and utilization goes up.

This talk shows how the CNCF incubating project HAMi addresses both. I will walk through topology-aware scheduling and device-sharing feature for different GPU vendors, and how to apply them on a Kubernetes AI cluster.

Tags
Kubernetes Topology Scheduling
Organizers