Time and room for this session are published in October. when the timetable goes live.
Everyone reaches for a frontier LLM by default — and then the bill, the latency, and the data-privacy questions arrive. But public routing research (RouteLLM, Microsoft's Hybrid LLM) shows that 40–60% of typical queries can be handled by a smaller model with no measurable drop in quality — and for specialized, bounded tasks, that number easily explodes toward the vast majority.
At Swiss Post, we put this to the test. We built a centrally self-hosted platform around a Small Language Model, served with vLLM on Kubernetes with GPUs, using OCR as our real-world case. The result: lower cost, better latency, full data sovereignty — and far less operational complexity than the "you can't self-host in production" narrative suggests.
This talk makes the case for SLMs as a serious production strategy, not a compromise. We share a candid comparison of cost, accuracy, and performance against LLM alternatives, how we kept sensitive citizen data fully in-house, and the architecture that makes self-hosting genuinely approachable.
The core message: an LLM is not a Swiss Army knife. For a huge share of real workloads, a specialized, self-hosted SLM is cheaper, faster, more private — and easier to ship than you'd expect. We close with the question that reframed our thinking: how many of your requests actually need an LLM in the end?