Cloud Native Austin

Ultimate AI: Go with your own intelligence

Capacity: 50
in-person
Event date
Oct 22, 26
06:00 PM - 08:00 PM CDT
Location
701 Brazos St, Austin, Texas, US
About this event

ABOUT THIS EVENT

🚀 Join us at Station Austin for Food, Drinks, and Kubernetes on Thursday, October 22, 2026!

Station Austin (formerly Capital Factory) is generously hosting the Cloud Native Austin meetup!

Everyone is welcome whether you are new to Kubernetes or already experienced, come connect with the Austin cloud native community. We'll share the latest updates, practical insights, and real world learnings from Kubernetes, CNCF projects, and modern platform engineering.

Featured Talk

🎤 Bill Kennedy Ultimate AI: Go with your own intelligence

Expect good conversations, practical takeaways, food, drinks, and a strong community vibe.

📍 Venue Station Austin 701 Brazos St, Austin, TX 78701

Arrival Instructions Upon arrival at the building, please proceed to 1st Floor, Apollo Room. We'll be in Apollo Room, 1st Floor (look for Capital Factory signs) after 5:45 pm.

Parking We know parking in Downtown is tricky! So, you can park in the building garage for just $8.00 (validation parking tickets will be distributed)! Street parking will still be an option. More information on parking here: https://www.capitalfactory.com/parking/

🏢 Venue Sponsored by Station Austin Thank you to Station Austin for sponsoring Cloud Native Austin! Station Austin is the center of gravity for entrepreneurs in Texas. They meet the best entrepreneurs in Texas and introduce them to their first investors, employees, mentors, and customers. To sign up for a Station Austin membership, click here: https://stationaustin.org/commons/

🤖 Sponsored by Cast AI Thank you to Cast AI for sponsoring Cloud Native Austin! Cast AI is the Kubernetes automation platform that keeps application performance and cloud costs on autopilot. It continuously analyzes how workloads actually behave and automatically tunes CPU, memory, and infrastructure provisioning — cutting cloud spend without hand-tuning your clusters. Learn more at https://cast.ai/

Sponsored by Ardan Labs Thank you to Ardan Labs for sponsoring Cloud Native Austin!

Looking forward to seeing you there!


AGENDA

6:30 PM — Ultimate AI: Go with your own intelligence In-Person · Bill Kennedy

Running open-source models is about much more than loading weights and sending a prompt. Every model server must make the same kinds of decisions: which model and quantization fit the available hardware, how requests are admitted and batched, how prompts are rendered, how context is cached, how tokens are sampled, and how concurrent generations share compute without sharing state.

In this one hour lecture, I will review the kronk architecture of the inference stack and small working programs. From there we follow a generation request through the entire serving lifecycle: admission control, model-specific prompt rendering, incremental message-cache selection, scheduling, slot assignment, KV-cache restoration, prompt prefill, batched decoding, sampling, parsing, and streaming.

By the end of the talk, you will be able to reason at a high-level about how an open-source model server works, configure one intentionally, diagnose performance and capacity problems, and use Kronk to run models on hardware you control.

Agenda
  1. 6:30 PM CDT

    Ultimate AI: Go with your own intelligence

    in-person

    Running open-source models is about much more than loading weights and sending a prompt. Every model server must make the same kinds of decisions: which model and quantization fit the available hardware, how requests are admitted and batched, how prompts are rendered, how context is cached, how tokens are sampled, and how concurrent generations share compute without sharing state.

    In this one hour lecture, I will review the kronk architecture of the inference stack and small working programs. From there we follow a generation request through the entire serving lifecycle: admission control, model-specific prompt rendering, incremental message-cache selection, scheduling, slot assignment, KV-cache restoration, prompt prefill, batched decoding, sampling, parsing, and streaming.

    By the end of the talk, you will be able to reason at a high-level about how an open-source model server works, configure one intentionally, diagnose performance and capacity problems, and use Kronk to run models on hardware you control.

Organizers