Can Claude Fix Itself? Using LLMs for Incident Response

QCon London 2026

Session

Can Claude Fix Itself? Using LLMs for Incident Response

Tuesday Mar 17 / 02:45PM GMT, Fleming (3rd Fl.) at The QEII Centre, London

Abstract

Can you throw an LLM at a production incident and expect useful results? A candid look from someone who runs a distributed AI system and reaches for Claude before reaching for a dashboard. Surprises, failures, and why the answer matters for every engineer carrying a pager.

Interview

It's a field report on using LLMs for incident response. I run a production AI system and these days I reach for Claude before I reach for a dashboard. It's still taboo to say this, and sometimes it would have been better to just open the dashboard, but I want to discuss when it is and when it isn't.

If you decide to skip this because it's yet another talk about AI, I totally understand. I sometimes think there's too much talk about AI and not enough doing. This one's a field report, I'm definitely not going to try and convince you that Claude will solve all your problems.

Loss of control feels very scary. What once felt like a comfortable on-call rotation where you knew all the nooks and crannies now includes a large language model that sometimes finds the issue faster than you can and sometimes feels like an overconfident junior.

Curiosity for experimenting. We're all learning together.

The audience has been paged at 3am and has worked with mission-critical systems. I can skip the intro and go straight to the interesting part.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 March

10:35 Windsor (5th Fl.) Session Sociotechnical Leadership Orienting, Understanding, Playing, Thriving: Debugging your Organisation Hazel Weakly Fellow @Nivenly Foundation; Director, Haskell Foundation; Experienced Leader Focusing on Organizational Change, Developer Experience, and Resilience Engineering 11:45 Whittle (3rd Fl.) Session Distributed Tracing How Eve Online Leverages Head Based Sampling to Observe "Fun" Nicholas Herring Technical Director, Eve Online @CCP Games, Refiner of Internet Spaceships and Explorer of Feral Gordian Knots of Python 13:35 Rutherford (4th Fl.) Unconference Unconference: Debugging Distributed Systems 14:45 Fleming (3rd Fl.) Session Can Claude Fix Itself? Using LLMs for Incident Response Alex Palcuie Member of Technical Staff in AI Reliability Engineering @Anthropic, Previously Staff Site Reliability Engineer on Google Cloud Platform 15:55 Mountbatten (6th Fl.) Session Observability Wrangling Telemetry at Scale: A Guide to Self-Hosted Observability Colin Douch Site Reliability Engineer @DuckDuckGo 17:05 Mountbatten (6th Fl.) Session Observability Are We All on the Same Page? Let’s Fix That - With AI Assistance Luis Mineiro Director of Digital Foundation @ASOS.com, SRE Charmer, Previously @Delivery Hero and @Zalando