What we covered
Getting a model to answer is a weekend. Getting it to answer at volume, at a cost you can defend, is the work. This session is about the second one.
- Usage patterns, and how the shape of your traffic decides everything downstream.
- Managed AI services and Model Garden: what you give up and what you stop maintaining.
- Vector stores in the cloud, and choosing one against your retrieval pattern rather than a benchmark.
- MLOps and LLMOps — what carries over from the former, and what genuinely doesn't.
- Where AI meets Kubernetes, and the Kubeflow architecture underneath it.
- Serving LLMs on Kubeflow, end to end.
- AI agents, and what they add to the inference bill.
Format
A technical talk aimed at the people who get paged, working through each layer of the serving stack.
Who it's for
Platform and ML engineers responsible for LLM infrastructure, and architects deciding between a managed endpoint and a cluster of their own.