Artificial Intelligence Technical Community Group

Distributed LLM Inference on Kubernetes with llm-d (CNCF Technical Community Group Meeting)

Capacity: 100
virtual
Event date
Jul 31, 26
08:00 AM - 09:00 AM PDT
Location
Virtual event
About this event

Distributed LLM Inference on Kubernetes with llm-d

Part of the CNCF Artificial Intelligence Technical Community Group Meeting Series.

Maroon Ayoub, Senior Principal ML Engineer at Red Hat and llm-d co-lead, walks through llm-d, the Kubernetes-native framework for distributed LLM inference and a CNCF Sandbox project backed by IBM, Google and Red Hat.

Cache locality, request routing and disaggregation decide your latency and your GPU bill. llm-d brings cloud-native patterns to that problem.

What You'll Learn

  • What llm-d is and how it fits the Kubernetes inference stack
  • KV-cache aware routing, sending requests where the context already lives
  • Disaggregated prefill/decode for better GPU utilization
  • Operating large-scale LLM inference with cloud native patterns

Project: https://github.com/llm-d/llm-d


Joining Details

Register for this event to receive the meeting link.

Tags
#cnai #Kubernetes #llm-d #AI-ML #Cloud-Native #Inference
Organizers