Speaker
Abstract
The past year has seen an absolute explosion in the use of AI and agents in particular, a trend that is guaranteed to accelerate going forward. The sheer usage and amount of agents puts an immense amount of pressure on the underlying cloud infra they run on; it is not uncommon to hear of services creating 10s of millions such sandboxes in a matter of weeks. In the face of such disruptive scale, what can be done to ensure that we don’t keep throwing money at the problem by building more and increasingly larger data centers to cope with the load?
To answer the question, in this talk we’ll cover our years-long journey aimed at severely optimizing and increasing the efficiency of how workloads are deployed on the cloud, beginning with research and OSS work into unikernels (specialized, ultra-efficient virtual machines) and covering the basics of virtualization and isolation primitives. From that basis, we will describe how we leveraged that work to build a cloud platform that can start any workload in a few milliseconds, and cram up to 1M+ such lightweight VMs into a single, off-the-shelf server, allowing for millions of strongly-isolated agents to be hosted in a rack, rather than an entire data center. Finally, we will show a live demo of this in action, including details of a k8s integration that retains these millisecond semantics.
Topics
QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Native Languages - and Wasm Hosted by Werner Schuster Kernel Developer @Wolfram, InfoQ Editor Functional ProgrammingFrom the same track
Monday 16 March
10:35 Windsor (5th Fl.) Session Deterministic Simulation Testing A Deterministic Simulation Testing (DST) Journey: From WASM in Go to State Machines in Rust Alfonso Subiotto Software Engineer @Polar Signals Deterministic simulation testing finds bugs by exploring random execution paths, injecting failures, and letting you replay any failure with a single starting seed. 11:45 Rutherford (4th Fl.) Unconference Unconference: Native Languages 14:45 Windsor (5th Fl.) Session Use<’lifetimes> For<’what> Ethan Brierley Senior Engineer @TrueLayer and Co-Organiser of Rust London As Rust has become more ergonomic, lifetimes have become more nuanced. By thinking of lifetimes as sets of loans, rather than using the traditional "regions of code" definition, this talk explores advanced lifetime concepts such as variance and higher ranked lifetimes. 15:55 Windsor (5th Fl.) Session Platform Engineering Fixing the AI Infra Scale Problem by Stuffing 1M Sandboxes in a Single Server Felipe Huici CEO and Co-Founder @Unikraft, Founder and Maintainer of the Linux Foundation Unikraft Open Source Project The past year has seen an absolute explosion in the use of AI and agents in particular, a trend that is guaranteed to accelerate going forward. 17:05 Windsor (5th Fl.) Session performance Understanding and Tuning System Performance with CPU Hardware Counters Bryan Boreham Distinguished Engineer @Grafana Labs, Member of the Prometheus Team, Expert in Distributed Systems and Computer Performance Counters are fundamental to monitoring: how many requests were processed, how many CPU-seconds consumed, how many bytes sent over a network. Very likely you are already monitoring your applications and operating systems via the hundreds or thousands of counters they expose.