Wrangling Telemetry at Scale: A Guide to Self-Hosted Observability

QCon London 2026

Session Observability

Wrangling Telemetry at Scale: A Guide to Self-Hosted Observability

Tuesday Mar 17 / 03:55PM GMT, Mountbatten (6th Fl.) at The QEII Centre, London

Abstract

Observability is supposed to help you tame complexity, but your Observability stack can quickly become just as complex as the systems it's meant to watch. For most teams, the answer is to pay someone else to deal with it. But bills grow, auditors ask awkward questions, and sometimes you just run out of road with your SaaS provider. In those instances, you have to turn to running it yourself.

Drawing on a decade of experience in building, maintaining, and operating self hosted monitoring and Observability stacks, in this talk, I will explain what it actually means to run your own stack, what the tooling landscape is, where it shines, and where the open source world struggles behind the SaaS experience.

Along the way, I'll cover options for all your telemetry types, with concrete recommendations on what to use and what to avoid, and insights on how to tie them together into one coherent debugging canvas, with a look at where the Observability world is going next.

Interview

My session is about self hosted Observability, and the options and unique challenges it presents. More than that, it's an insight into what's going under the hood in an observability system, and aims to contextualise that to better enable engineers to understand and work with their telemetry going forward

Especially in the world of AI, our systems are becoming more complex by the day. Observability is the answer to tackling that complexity, but you have to do it right, which means knowing how to get the best out of your telemetry systems

Lots of developers struggle with tying telemetry together into a cohesive debugging strategy. Try as we might, the "three pillar" idea still exists, and is sub optimal for the modern distributed system.

Ways to tie telemetry of different types. In particular "exemplars" are an underused aspect of the modern metrics system.

Topics

Observability Platform Engineering Insights
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 March

10:35 Windsor (5th Fl.) Session Sociotechnical Leadership Orienting, Understanding, Playing, Thriving: Debugging your Organisation Hazel Weakly Fellow @Nivenly Foundation; Director, Haskell Foundation; Experienced Leader Focusing on Organizational Change, Developer Experience, and Resilience Engineering 11:45 Whittle (3rd Fl.) Session Distributed Tracing How Eve Online Leverages Head Based Sampling to Observe "Fun" Nicholas Herring Technical Director, Eve Online @CCP Games, Refiner of Internet Spaceships and Explorer of Feral Gordian Knots of Python 13:35 Rutherford (4th Fl.) Unconference Unconference: Debugging Distributed Systems 14:45 Fleming (3rd Fl.) Session Can Claude Fix Itself? Using LLMs for Incident Response Alex Palcuie Member of Technical Staff in AI Reliability Engineering @Anthropic, Previously Staff Site Reliability Engineer on Google Cloud Platform 15:55 Mountbatten (6th Fl.) Session Observability Wrangling Telemetry at Scale: A Guide to Self-Hosted Observability Colin Douch Site Reliability Engineer @DuckDuckGo 17:05 Mountbatten (6th Fl.) Session Observability Are We All on the Same Page? Let’s Fix That - With AI Assistance Luis Mineiro Director of Digital Foundation @ASOS.com, SRE Charmer, Previously @Delivery Hero and @Zalando