Suresh Michael
All workshops
LLM Engineering

Large Language Model Inferencing in Production

A technical session on serving LLMs for real: usage patterns, managed services versus self-hosting, vector stores, LLMOps, and running inference on Kubernetes with Kubeflow.

Delivered
October 12, 2025
Format
Technical talk
Audience
Platform & ML engineers
SlidesView on SlideShare →

What we covered

Getting a model to answer is a weekend. Getting it to answer at volume, at a cost you can defend, is the work. This session is about the second one.

  • Usage patterns, and how the shape of your traffic decides everything downstream.
  • Managed AI services and Model Garden: what you give up and what you stop maintaining.
  • Vector stores in the cloud, and choosing one against your retrieval pattern rather than a benchmark.
  • MLOps and LLMOps — what carries over from the former, and what genuinely doesn't.
  • Where AI meets Kubernetes, and the Kubeflow architecture underneath it.
  • Serving LLMs on Kubeflow, end to end.
  • AI agents, and what they add to the inference bill.

Format

A technical talk aimed at the people who get paged, working through each layer of the serving stack.

Who it's for

Platform and ML engineers responsible for LLM infrastructure, and architects deciding between a managed endpoint and a cluster of their own.

Share this workshop
Copied