Are We All on the Same Page? Let’s Fix That - With AI Assistance

QCon London 2026

Session Observability

Are We All on the Same Page? Let’s Fix That - With AI Assistance

Tuesday Mar 17 / 05:05PM GMT, Mountbatten (6th Fl.) at The QEII Centre, London

Abstract

In distributed systems, incidents rarely fail because of missing signals - they fail because the right people aren’t mobilised quickly enough, and teams struggle to build a shared understanding under pressure. Customer-facing teams absorb pages for downstream failures, ownership blurs, and valuable time is lost coordinating humans rather than solving problems.

 

This talk takes a sociotechnical view of debugging distributed systems. Beyond traces and metrics, effective incident response depends on how teams are structured, how ownership is defined, and how people collaborate during failure. Routing incidents to the team closest to the problem is necessary - but it’s only the starting point.

 

We explore how trace causality can be used to align alerts with ownership, and how AI can act as a teammate during incident response - helping teams converge on probable causes, correlate signals, and surface remediation options directly in the moment of paging. Rather than replacing engineers, AI reduces cognitive load and accelerates shared understanding when it matters most.

 

Drawing on real-world experience, we’ll show how combining intelligent routing, collaborative debugging practices, and AI-assisted investigation transforms incident response from a noisy escalation chain into a coordinated team effort. The result isn’t just faster recovery - it’s teams that can confidently own their systems in production.

 

The approach is vendor-neutral and architecture-agnostic, applicable across modern microservice environments and observability stacks.

Interview

The session explores incident response as a sociotechnical problem, not just a technical one. Most incidents fail because teams can’t mobilize quickly enough or align on the root cause under pressure - not because monitoring is missing signals. Senior developers will learn practical patterns to combine trace causality with intelligent alert routing so the right people own the right problems, and how AI can act as a thinking partner to accelerate that shared understanding.

As we head into 2026, the complexity of distributed platforms is growing faster than the operational capacity of the teams building them. Alert fatigue and unclear ownership are becoming the primary brake on incident response. Organizations are already investing heavily in observability infrastructure, but without the sociotechnical glue - structured ownership models and intelligent routing - those signals drown teams rather than help them. 2026 is the year to move past “better metrics” to “better human coordination.”

The classic challenges: symptom-based alerting works great until you’re paging the same team for ten different root causes deeper in the distributed system. Trace data exists, but teams lack the ownership structures or automation to route alerts based on causality rather than topology. And when an incident happens, people spend a lot of time gathering context and coordinating rather than solving. After figuring out where the actual problem may be, effective mitigation requires skills that are not abundant - that context assembly is where AI can add genuine value, not as a replacement but as a thinking partner.

Audit your alert rules against your actual team ownership model. Most organizations alert based on technical symptoms (service error rates) without aligning that to which team actually owns fixing it. Start with symptom-based alerting on your critical paths, then layer in one adaptive paging rule that follows trace causality to the root service owner. You’ll immediately see fewer pages and faster resolution.

QCon brings together practitioners who’ve actually built and operated these systems at scale. The talks move beyond theory - they’re grounded in real constraints: cost, team capacity, organizational structures, and the messy reality of production systems. That’s exactly the context you need when thinking through incident response, because no solution works in isolation from your team’s structure and culture.

Topics

Observability Alerting monitoring Paging Troubleshooting debugging RCA AI-Assisted
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 March

10:35 Windsor (5th Fl.) Session Sociotechnical Leadership Orienting, Understanding, Playing, Thriving: Debugging your Organisation Hazel Weakly Fellow @Nivenly Foundation; Director, Haskell Foundation; Experienced Leader Focusing on Organizational Change, Developer Experience, and Resilience Engineering 11:45 Whittle (3rd Fl.) Session Distributed Tracing How Eve Online Leverages Head Based Sampling to Observe "Fun" Nicholas Herring Technical Director, Eve Online @CCP Games, Refiner of Internet Spaceships and Explorer of Feral Gordian Knots of Python 13:35 Rutherford (4th Fl.) Unconference Unconference: Debugging Distributed Systems 14:45 Fleming (3rd Fl.) Session Can Claude Fix Itself? Using LLMs for Incident Response Alex Palcuie Member of Technical Staff in AI Reliability Engineering @Anthropic, Previously Staff Site Reliability Engineer on Google Cloud Platform 15:55 Mountbatten (6th Fl.) Session Observability Wrangling Telemetry at Scale: A Guide to Self-Hosted Observability Colin Douch Site Reliability Engineer @DuckDuckGo 17:05 Mountbatten (6th Fl.) Session Observability Are We All on the Same Page? Let’s Fix That - With AI Assistance Luis Mineiro Director of Digital Foundation @ASOS.com, SRE Charmer, Previously @Delivery Hero and @Zalando