Abstract
In a world where AI and ML are rapidly evolving, the need for efficient Realtime Feature Platforms has never been greater. But the journey to create one is far from straightforward.
In the talk, Ivan Burmistrov will share how ShareChat - the largest social network in India - built their own Realtime Feature Platform serving more than 1 billion features per second, and how they managed to make it cost-efficient.
Ivan will cover the challenges the team faced along the way, how they managed to overcome them and which ones are still not fully resolved. The talk will also cover the experience in using relatively new technologies such as ScyllaDB and RedPanda and why such technologies are crucial for building a cost efficient system. Additionally, Ivan will share how the system leverages Apache Flink in the very core of the data pipeline.
This talk will provide insights for anyone interested in real-time data pipelines and Realtime Feature Platforms, in particular.
Interview
I'm a Principal Software Engineer at ShareChat, working on infrastructure for a recommendation system, with a particular focus on a realtime feature store and other data-related systems.
Building a realtime feature store has been a lot of fun, and there are a lot of real-life lessons I've learned - which are hard to find online. So, the motivation is to share something worthwhile.
Anyone who is curious about realtime data processing and low-latency systems.
I'd like attendees to walk away with an understanding of what it takes to build a realtime feature store, and generic tips and tricks on realtime data processing.
Topics
QCon London 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Tuesday 9 April
10:35 Windsor (5th Fl.) Session database Powering User Experiences with Streaming Dataflow Alana Marzoev Founder & CEO @ReadySet Streaming dataflow provides a unique solution to scaling OLTP applications by allowing for an efficient cache implementation that does not diverge from the relational model of the underlying data store. 11:45 Windsor (5th Fl.) Session ML Feature Store The Harsh Reality of Building a Realtime ML Feature Platform Ivan Burmistrov Principal Software Engineer @ShareChat In a world where AI and ML are rapidly evolving, the need for efficient Realtime Feature Platforms has never been greater. But the journey to create one is far from straightforward. 13:35 Windsor (5th Fl.) Session architecture High Performance Time-Series Database Design With QuestDB Vlad Ilyushchenko Co-Founder & CTO @QuestDB, OG Author of PSY-Probe, Geek In this talk we will explore the world of time series and unique set of problems time series present to the developers. We will discuss the engineering principles behind QuestDB's design, focusing on high performance. 14:45 Whittle (3rd Fl.) Session architecture Improving Developer Experience Using Automated Data CI/CD Pipelines Noémi Ványi, Simona Pencea Validating your code against actual production data can be challenging. We have all been at least once on the receiving end of a "test1" email subject because somebody somewhere did a test with the production database. 15:55 Windsor (5th Fl.) Session Building Databases Rockset - Building a Modern Analytics Database on Top of RocksDB Igor Canadi Founding Engineer and Architect @Rockset, Previously at RocksDB and Facebook RocksDB, a key-value store built on the foundation of Log-Structured Merge-Tree data structures and originally open-sourced by Facebook, has played a significant role in shaping data systems over the past decades. 17:05 Rutherford (4th Fl.) Unconference Unconference: Innovations in Data Engineering An unconference is a participant-driven meeting. Attendees come together, bringing their challenges and relying on the experience and know-how of their peers for solutions.