Speaker
Abstract
Let's be honest - observability can suck. Ever feel like you're swimming in dashboard soup? You know the feeling: tons of single-use dashboards, building new ones during every incident only to lose them in the chaos, and spending ages creating visualizations that no one ever looks at again. Even with all the right tools, something still feels off.
This talk shares the journey of how our team tackled this exact problem while building an on-call tool with high expectations for reliability. You'll hear how we went from feeling lost in our own tooling to becoming confident enough to sleep soundly while on-call. Through building clear system boundaries, a thoughtful approach to dashboards, and running hands-on drills with our team, our observability approach became the key thing that makes our system feel reliable.
Key takeaways include:
- Principles for building dashboards that are useful, long lived, and widely adopted
- How to layer your stack, by connecting your observability tooling to build a debugging flow that puts UX first
- Creating a culture change towards observability, getting your team invested through hands-on drills
QCon London 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Wednesday 9 April
10:35 Fleming (3rd Fl.) Session resiliency Timeouts, Retries and Idempotency In Distributed Systems Sam Newman Microservice, Cloud, CI/CD Expert, Author of "Building Microservices" and "Monolith to Microservices", 20+ Years Experience as a Developer The definition of insanity is doing the same thing over and over again” - this quote attributed to Einstein warns us of the danger of magical thinking, hoping that trying something just one more time will achieve success when before we failed. But is this really insanity? 11:45 Fleming (3rd Fl.) Session architecture From Confusion to Clarity: Advanced Observability Strategies for Media Workflows at Netflix Sujana Sooreddy, Naveen Mareddy Managing media workflows at the Netflix scale is both thrilling and daunting. With millions of workflow executions across hundreds of types and over 500 million CPU hours consumed quarterly, costs can skyrocket, and encoding issues can disrupt the streaming experience. 13:35 Whittle (3rd Fl.) Session APIs Scaling API Independence: Mocking, Contract Testing & Observability in Large Microservices Environments Tom Akehurst CTO and Co-Founder @WireMock, 20+ Years Building Enterprise Systems Microservices promise faster deployments and team autonomy. In reality, engineers are often blocked waiting for APIs, dealing with broken sandboxes, or wrangling test environments. 14:45 Whittle (3rd Fl.) Session From Dashboard Soup to Observability Lasagna: Building Better Layers Martha Lambert Product Engineer @incident.io, Building Reliable and Observable Systems Let's be honest - observability can suck. Ever feel like you're swimming in dashboard soup? You know the feeling: tons of single-use dashboards, building new ones during every incident only to lose them in the chaos, and spending ages creating visualizations that no one ever looks at again. 15:55 Fleming (3rd Fl.) Session architecture Platforms for Secure API Connectivity With Architecture as Code Jim Gough Distinguished Engineer, API Platform Lead Architect @Morgan Stanley, Co-Author of Optimizing Java As microservices and complex platforms become the standard, ensuring secure connectivity while maintaining a smooth developer experience is a significant challenge. Traditional security models often introduce friction, slowing down innovation and deployment.