Speaker
Abstract
A unique pattern in video game software is real-time interactions to express the personality of users.
Here we will talk about how we instrument the universe of New Eden to identify the traffic that matters, even the "fun" parts!
We'll explore how we achieve performant observability in a 20 year old legacy system running along side modern technologies:
- 100% Head base sampling ecosystem
- Blend of dynamic and deterministic sampling techniques
- The blessing and curse that is the Exponential Moving Average (EMA)
- Observing "fun" without breaking the bank
Topics
Distributed Tracing
sampling
legacy systems
video games
real-time systems
76%
senior dev or higher
1:11
speaker ratio
60+
practitioners
QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Tuesday 17 March
10:35 Windsor (5th Fl.) Session Sociotechnical Leadership Orienting, Understanding, Playing, Thriving: Debugging your Organisation Hazel Weakly Fellow @Nivenly Foundation; Director, Haskell Foundation; Experienced Leader Focusing on Organizational Change, Developer Experience, and Resilience Engineering Debugging is both an art and a science. But more than that, it's an activity undertaken with deep intention: to understand and improve your systems. In the purely technical realm, we have an extraordinary range of tooling and techniques that can help us tackle this problem. 11:45 Whittle (3rd Fl.) Session Distributed Tracing How Eve Online Leverages Head Based Sampling to Observe "Fun" Nicholas Herring Technical Director, Eve Online @CCP Games, Refiner of Internet Spaceships and Explorer of Feral Gordian Knots of Python A unique pattern in video game software is real-time interactions to express the personality of users. Here we will talk about how we instrument the universe of New Eden to identify the traffic that matters, even the "fun" parts! 13:35 Rutherford (4th Fl.) Unconference Unconference: Debugging Distributed Systems 14:45 Fleming (3rd Fl.) Session Can Claude Fix Itself? Using LLMs for Incident Response Alex Palcuie Member of Technical Staff in AI Reliability Engineering @Anthropic, Previously Staff Site Reliability Engineer on Google Cloud Platform Can you throw an LLM at a production incident and expect useful results? A candid look from someone who runs a distributed AI system and reaches for Claude before reaching for a dashboard. Surprises, failures, and why the answer matters for every engineer carrying a pager. 15:55 Mountbatten (6th Fl.) Session Observability Wrangling Telemetry at Scale: A Guide to Self-Hosted Observability Colin Douch Site Reliability Engineer @DuckDuckGo Observability is supposed to help you tame complexity, but your Observability stack can quickly become just as complex as the systems it's meant to watch. For most teams, the answer is to pay someone else to deal with it. 17:05 Mountbatten (6th Fl.) Session Observability Are We All on the Same Page? Let’s Fix That - With AI Assistance Luis Mineiro Director of Digital Foundation @ASOS.com, SRE Charmer, Previously @Delivery Hero and @Zalando In distributed systems, incidents rarely fail because of missing signals - they fail because the right people aren’t mobilised quickly enough, and teams struggle to build a shared understanding under pressure.