Cloud Native Community Japan

Cloud Native Community Japan - AI Infra Meetup #1

Capacity: 180
hybrid
Event date
Oct 1, 26
06:00 PM - 09:30 PM JST
Registration is open until Oct 1, 2026 at 6:00 PM JST.
Location
Call for Speakers

講演スピーカー(プロポーザル)を募集します!

第1回ミートアップを一緒に盛り上げてくださるスピーカーを募集します。 上記の通り初回開催として、「CNCFのAI Infraインフラとしての現在地を把握しよう」をテーマ、幅広いトピックをカバーするセッションを募集します。 上記に挙げた技術トピックに関する解説・検証実績、自社でのAI/ML基盤の運用事例、アップストリームへのコントリビューション経験など、大小問わず大歓迎です! 「こんなテーマで話してみたい」「まだ検証段階だけど知見を共有したい」という方も、ぜひお気軽にご応募ください。

  • 募集セッション枠:
    • 一般セッション(30分程度)
    • ライトニングトーク(LT)/ショートセッション(10分程度)
    • 提出されるプロポーザルのDuration(セッション時間)の設定を確認の上、上記の時間の範囲で提出をお願いいたします

■募集対象トピック(一例)

  • スケジューリング(DRA, WAS, TAS, Kueue など)
  • オーケストレーション(JobSet, LeaderWorkerSet など)
  • デプロイスタック(KServe, NVIDIA Dynamo、llm-d, agentgateway, Envoy AI Gateway など)
  • Agent Workload(kagent, AgentSandbox など)
  • AI Conformance
  • その他AI Workloadに関連する内容

※上記に限らず、関連トピックの提案を歓迎します。

イベント概要

  • 開催日: 2026年10月1日
  • 開催地: LINEヤフー(ガーデンテラス紀尾井町)
  • 講演時間: 5分〜30分
  • 講演内容: AI workloadに関連するトピックであればなんでもOK!
    • ※営利的な宣伝はお控えください。(1ページの採用募集程度であればOK)

スケジュール

  • 応募締切: 2026年8月28日 23:59

プロポーザルの提出方法

  1. プロポーザルを作成 : Proposals ダッシュボードで作成してください。詳しくはこちら

Duration(セッション時間)の設定 :

  • 一般セッションの場合は30分
  • ライトニングトーク(LT)/ショートセッションの場合は10分
  1. イベントに提出 : 本イベントページの 「Submit session proposal」 をクリックし、作成したプロポーザルを選択してください

  2. 提出を確認 : Submissions ダッシュボードで提出状況を確認してください。詳しくはこちら

お問い合わせ:cncj@googlegroups.com

Submissions open: Jul 3, 10:00 AM - Aug 28, 11:59 PM JST, 2026
About this event

AI Infra SIGの設立を記念し、記念すべき第1回目のミートアップを開催します! テーマは「CNCFのAI Infraインフラとしての現在地を把握しよう」です! 最先端のCloud Native AIインフラに興味があるエンジニアの皆様、ぜひご参加・ご登壇ください。

To celebrate the launch of the AI Infrastructure SIG, we are excited to host our very first meetup. Our inaugural theme is “Understanding the Current State of CNCF as the Foundation for AI Infrastructure.” If you’re an engineer interested in the latest developments in Cloud Native AI infrastructure, we would love to have you join us. We also welcome talk proposals from anyone eager to share their knowledge and experience with the community.

AI Infra SIG (Special Interest Group)

Cloud Native Platform Engineering Japanは、CNCFの日本チャプターであるCNCJ配下のSIGとして発足した、AI Infra領域のコミュニティです。 AI Infra SIG (Special Interest Group)では、AI/機械学習ワークロードを動かすためのベストプラクティスや最適化・現場の実践知の共有にフォーカスし、生成AIや自律型エージェントをはじめとするAI/機械学習技術の進化を支えるインフラとしてのKubernetesやCloud Nativeエコシステムについて取り扱います。 国内のエンジニア、研究者、プラットフォーム開発者、企業、組織がCloud Native × AIのフロンティアを切り拓いていくためのコミュニティとして、ご参加をお待ちしております。

🔗 CNCF BLOG: CNCJ AI Infra SIG発足のお知らせ & 第1回 Meetup 開催・スピーカー募集開始!

Cloud Native Platform Engineering Japan is an AI infrastructure community established as a Special Interest Group (SIG) under the Cloud Native Community Japan (CNCJ), the official Japan Chapter of the Cloud Native Computing Foundation (CNCF). The AI Infrastructure SIG focuses on sharing best practices, optimization techniques, and real-world operational knowledge for running AI and machine learning workloads. Its scope covers Kubernetes and the broader Cloud Native ecosystem as the infrastructure powering the next generation of AI technologies, including Generative AI and autonomous agents. We welcome engineers, researchers, platform developers, companies, and organizations across Japan who are interested in advancing the frontier of Cloud Native × AI. Join us in building and shaping the future of AI infrastructure together.

🔗 CNCF BLOG: Launch of the AI Infra SIG under the CNCF Japan chapter: First meetup and call for speakers

📝 Registration (Free)

LF Accountを作成のうえログイン後、「Attend Event」ボタンより参加登録が可能です。 受付時間は17:30〜18:30です。18:30以降はイベント会場への入場はできませんのでご注意ください。

After creating an LF Account and logging in, you can register to participate by clicking the "Attend Event" button. Registration desk is open from 17:30 to 18:30. Please note that entry will not be permitted after 18:30.

🚶 Access to the Veneue

LINEヤフー株式会社 紀尾井町オフィス (東京都千代田区紀尾井町1-3 東京ガーデンテラス紀尾井町 紀尾井タワー) → 受付は紀尾井タワー2階カウンターに設置しております。 東京メトロ南北線 「永田町駅」9a出口直結 東京メトロ半蔵門線 「永田町駅」7番出口より徒歩2分 東京メトロ丸の内線/銀座線 「赤坂見附駅」D出口より徒歩1分 ※駅からオフィス入口まで少し距離が離れておりますので、お時間に余裕をもってお越しください。

LY Corporation – Kioicho Office (Kioi Tower, Tokyo Garden Terrace Kioicho, 1-3 Kioicho, Chiyoda-ku, Tokyo) → Please check in at the reception counter on the 2nd floor of Kioi Tower. Direct access from Exit 9a of Nagatacho Station (Tokyo Metro Namboku Line); 2-minute walk from Exit 7 of Nagatacho Station (Tokyo Metro Hanzomon Line); 1-minute walk from Exit D of Akasaka-mitsuke Station (Tokyo Metro Marunouchi/Ginza Lines). ※Please allow yourself plenty of time for your arrival, as there is some distance between the stations and the office entrance.

👥 Event Organizers

  • Masaya Aoyama, Senior Software Engineer / Product Owner, CyberAgent, Inc.
  • Sunyanan Choochotkaew (Pang), Senior Research Scientist, IBM Research - Tokyo
  • Toru Komatsu (Toru), Engineer, Preferred Networks, Inc
  • Shingo Omura, Principal Architect of AI Infrastructure, LY Corporation
  • Kenta Tada, Member, CNCF End User Technical Advisory Board & eBPF Foundation Governing Board
  • Kenji Tagashira, Senior Manager, Fsas Technologies Inc
Agenda
  1. 6:00 PM - 6:10 PM JST

    Opening

    hybrid
    SPEAKERS
  2. 6:10 PM - 6:40 PM JST

    AI を支えるクラウドネイティブエコシステム — AI/ML インフラを構成する技術の現在地

    hybrid

    AI の普及に伴い、クラウドネイティブエコシステムは、従来の一般的なウェブサービスやバッチ処理とは異なる要件に直面しています。希少なアクセラレータの共有、分散ワークロードの実行、モデルやデータの配置、状態を考慮した推論リクエストの処理など、AI/ML ワークロードに適した新しい仕組みが求められています。

    本セッションでは、スケジューリングやサービングを主な切り口として、関連する Kubernetes/CNCF 周辺のプロジェクトや OSS の動向を概観します。Workload-Aware Scheduling (WAS), Kueue, Gateway API, Inference Extension などの技術が、AI/ML インフラにおいてどのような課題を解決し、どのような役割を担うのかを紹介します。

    さらに、Preferred Networks の社内 AI/ML 基盤における実践を題材に、さまざまな要素技術がどのように連携し、AI/ML の実験・開発・学習・推論を支えるのかを説明します。AI/ML インフラに関心のある開発者やプラットフォームエンジニアを対象とした、エコシステムの全体像をつかむための入門セッションです。

    SPEAKERS
  3. 6:40 PM - 7:10 PM JST

    Inference as a Service for Chat and Non-Chat Models on On-Prem Kubernetes

    hybrid

    SB Intuitionsでは、NVIDIA GB200 NVL72 で構成されたオンプレミスKubernetesクラスタを推論基盤として運用しています。 この上には、自社でフルスクラッチ開発したLLMおよびSTT/TTSモデルの商用ホスティングと、社内・グループ会社向けの検証用の大規模open-weightモデル提供という、特性の異なる多数のモデルが同居しています。

    これらのモデル群を効率よくk8sクラスタでホスティングするためのInferenceプラットフォームをGateway API, Gateway API Inference Extension (GAIE), llm-dという、upstreamで標準化が進むスタックで構築しています。一方でこれらは "1モデル1InferencePool"、"OpenAI互換の /v1/chat/completions API" を暗黙の前提としており、社内の要求を満たせないケースがありました。本セッションではこのプラットフォームの技術要素と構築する際にあった技術的な課題について取り上げます。

    具体的には、

    • 多品種のモデルを共通の方法でデプロイ,払い出しするための仕組みを自前で用意する必要があった
    • 2026年8月時点では、audioを入力にとる /v1/audio/transcriptions のようなエンドポイントをそのままでは扱えない
    • vLLMにロードされた個々のLoRAアダプタ単位でのアクセス制御は標準では提供されていない

    といった課題に対し、どのように対処したのかを紹介します。

    SPEAKERS
  4. 7:10 PM - 7:40 PM JST

    Beyond Fast Inference: Balancing Performance and Operational Excellence for Agentic AI Platform with NVIDIA Dynamo

    hybrid

    With the rapidly growing demand for Agentic AI, it is essential to design a unified system that not only processes individual requests at high performance but also handles routing that takes long-term context and KV cache into account, manages model outputs involving tool usage, supports deployments spanning multiple GPUs and nodes on Kubernetes, and ensures operations that expect failures and updates.

    In this session, we will provide an overview of the evolution of NVIDIA Dynamo since v1.0 - which reached general availability (GA) in March 2026 - as a set of features that constitute the Agentic AI infrastructure. As for the features for agents, we will introduce a mechanism that normalizes model-specific outputs into OpenAI-compatible responses in the case of tool calling and reasoning. Also, we will discuss how the agent harness can pass execution intent - such as priority and expected output length - as "Agent Hints," which Dynamo then incorporates into routing and scheduling. In long agent sessions, KV cache handling and routing become particularly important. We will explain the evolution of KV-aware routing, which has advanced in response to these demands. Related to this, there are options for routing deployment. In addition to the approach using the Dynamo Frontend, we will also introduce methods for working with Kubernetes Gateways together via the Gateway API Inference Extension (GAIE) and the Endpoint Picker Plugin (EPP). Furthermore, we will clarify the role of Grove in Kubernetes orchestration and discuss the deployment and scaling of AI workloads, including disaggregated serving. Finally, we’ll provide an overview of request migration, graceful shutdown, load control, fast model loading, trace replay, simulation, and other features from the perspective of the resiliency and validation required in real-world operations.

    We'll provide the first step toward viewing Dynamo not merely as a high-speed inference server, but as an agent-aware, cloud-native, distributed and disaggregated inference platform that balances performance and operability.

    SPEAKERS
  5. 7:40 PM - 7:45 PM JST

    Break

    hybrid
  6. 7:45 PM - 7:55 PM JST

    Multi-Tenant Network Security for KServe Serverless Mode with Istio Service Mesh

    hybrid

    This lightning talk will be conducted in English, but the presentation slides will have Japanese content. This talk presents an overview of KServe's network architecture in Serverless mode, which is backed by the Knative Serving platform and a service mesh, along with network security considerations for intra-cluster communication. By default, KServe doesn't offer any network security option, but we can implement network traffic restrictions at the service mesh level. This talk will showcase Istio's AuthorizationPolicy to control traffic from clients in tenant namespaces to KServe's AI inference endpoints in other namespaces.

    SPEAKERS
  7. 7:55 PM - 8:05 PM JST

    More Freedom on the Same Shared GPU Cluster: A Small Team’s Experience with vCluster

    hybrid

    日本語タイトル:共有GPUクラスタのまま、自由を増やす——少人数運用で試したvCluster 私たちは約4名のチームで、データセンター設備からKubernetesまで、社内向けオンプレGPU基盤を運用しています。 当初はAI研究者によるNotebookやGPUジョブの利用が中心で、Namespace単位の分離で十分でした。しかし、AIアプリケーション開発チームが増えるにつれ、CRDやOperatorを含むHelmチャートなど、Namespaceの権限だけでは扱いにくいアプリケーションを検証したいという要求が生まれました。 そこで、共有GPUクラスタを維持しながら利用者の自由度を高めるため、vClusterを導入しました。HostのNode情報を同期することで、GPUノードの指定には従来のYAMLをほぼそのまま利用でき、既存ワークロードを大きく変更せずに移行できました。 本LTでは、NodeやIngressなど、HostとvClusterの間で何を共有するかという設定を取り上げ、既存の操作性を維持しながら分離を追加するために得られた知見を紹介します。

    SPEAKERS
  8. 8:05 PM - 8:15 PM JST

    kagent から見る、AI エージェント時代の Kubernetes 運用を考える

    hybrid

    AI エージェントが主流となりつつある時代、Kubernetes の運用はどう変わるのでしょうか。本セッションでは、CNCF Sandbox の kagent を実際に動かしながら、その問いを初心者目線で一緒に考えて行きます。

    kagent を使うと、AI エージェントは特別な基盤ではなく Kubernetes 上の “普通のワークロード” として動きます。まず「エージェントとは何か(プロンプト+モデル+ツールの組み合わせ)」と「MCP(Model Context Protocol)という “道具のつなぎ方” 」を整理します。

    そのうえでデモとして、CrashLoopBackOff・ImagePullBackOff・Service の設定ミスという “よくある3つの障害” を、エージェントがチャットで調査して直す様子をお見せします。ユーザー(デモ実行者)が行うことは「障害を仕組む」ことと「結果を確認する」ことだけです。調査も修復もエージェントが行います。生成 AI や MCP の予備知識は不要です。

    そして最後に、動かしてみて見えてきた「AI エージェント時代の Kubernetes 運用」を考えます。運用者の役割が「操作する」から「意図を伝えて承認する」へ移っていくこと。実際に手を動かした実感をもとにお話しします。

    SPEAKERS
  9. 8:15 PM - 8:20 PM JST

    クロージング

    hybrid
  10. 8:20 PM - 9:30 PM JST

    懇親会

    in-person
Organizers