Building High-Fidelity Data Streams

QCon London 2023

Session

Building High-Fidelity Data Streams

Monday Mar 27 / 01:40PM BST, Whittle (3rd Fl.)

Abstract

Low latency data streaming technology and practices remain a hot and trending topic among data engineers today. At its core, it promises to deliver data in near real time in order to provide snappy data-driven user experiences. This experience comes in many forms including low latency updates to social news feeds, near-real time payment fraud prevention, time-relevant recommender systems used in flash sales, self-driving car route planning, and more. Our need to stay engaged has made low latency data streams a critical part of modern data architectures. 

While it may seem trivial to get a data streaming POC up and running, productionalizing such a system under strict SLAs with the aid of a lean engineering team requires making the right choices but also learning from mistakes along the way. At Datazoom, we built a lossless streaming data system that guarantees sub-second (p95) event delivery at scale with better than three nines availability – we measure availability in terms of the on-time delivery of events. Come to this talk to learn how you can build such a system soup-to-nuts.

Interview

I currently serve as the Chief Architect and Head of Engineering at Datazoom, a company that offers a video data platform that captures video playback telemetry data. This data can be used to understand how customers experience and interact with video. At Datazoom, we build both client SDKs and a cloud-based analytics platform.

In this talk, I explain how engineers can build a low-latency, high-fidelity data streaming system using open source software and public cloud technologies combined with recommended best practices. My talk focuses on the non-functional requirements (e.g. the -ilities) of such a system including but not limited to scalability, performance, reliability, observability, availability, etc…

This talk will take a ground-up approach to building such a system. My talk requires little background knowledge beyond basic familiarity with various AWS technologies & Apache Kafka. The ideal target audience would be composed of engineers, ranging from beginner to intermediate, interested in building a high-fidelity streaming system.

This talk will serve as an architect’s guide to building a high-fidelity streaming system. While it may leave out specific details for lack of time, it will provide enough information to get an architect 80% of the way to building a similar system.

AI-managed data infra – it is sorely needed in order to reduce the onerous task of operating data infrastructure at scale.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 27 March

10:35 Churchill (Ground Fl.) Session scalability Scaling Google's Global Cloud L7 Load Balancer James Spooner Principal Engineer, Load Balancing @Google 11:50 Rutherford (4th Fl.) Unconference Unconference: Architectures You've Always Wondered About Shane Hastie Global Delivery Lead @SoftEd, Lead Editor for Culture & Methods @InfoQ 13:40 Whittle (3rd Fl.) Session Building High-Fidelity Data Streams Sid Anand Fellow, Cloud & Data Platform @Walmart, Apache Airflow Committer/PMC, Ex-Netflix, LinkedIn, eBay, Etsy, & PayPal 14:55 Fleming (3rd Fl.) Session Microservices Tales of Kafka @Cloudflare: Lessons Learnt on the Way to 1 Trillion Messages Andrea Medda, Matt Boyle 16:10 Churchill (Ground Fl.) Session scalability Zoom: Why Does It Work? Ian Sleebe Senior Solutions Architect @Zoom 17:25 Fleming (3rd Fl.) Session Microservices Banking on Thousands of Microservices Suhail Patel Senior Staff Engineer @Monzo Leading the Platform and Data Functions, Previously @Citymapper