Inference engineering: the subtle art of token generation

October 26, 2026 / 4:00 pm - 4:30 pm

In this talk, I will discuss the challenges in serving LLMs on on-premises or cloud infrastructure. I will go over the mechanics of LLM inference and how it affects throughput and the solutions major inference frameworks offer to optimize token generation. 

skip to content

Schedule

4:00 pm

Inference engineering: the subtle art of token generation

Most agentic solution starts off using proprietary model APIs. There comes a moment when you reach the limits of proprietary models. Whether it is concern about control of your data or data privacy, improving latency and availability, or reducing your cost. The solution is serving (fine-tuned) open models. In this talk, I will discuss the challenges in serving LLMs on on-premises or cloud infrastructure. I will go over the mechanics of LLM inference and how it affects throughput and the solutions major inference frameworks offer to optimize token generation. 

Guests

Roy van Santen Machine Learning Engineer Xebia