Speaker
Abstract
There are a few common and mostly well-known challenges when architecting for data. For example, many data teams struggle to move data in a stable and reliable way from operational systems to analytics systems. At the same time, they must manage complex and often costly infrastructure landscapes. These issues hinder companies from effectively leveraging their data for business purposes.
Drawing from real-world experience, this presentation will explore how we address these challenges by building reliable and scalable data platforms with reasonable costs. It will also cover solutions to help operational teams provide their data and to observe the flow towards analytics systems. In addition to discussing architectures and design considerations the presentation will also highlight tools and techniques used to implement these platforms.
Interview
In one sentence, I would say: availability of data for analytical use, regardless of whether it is a dashboard, machine learning or GenAI. In addition to data platforms and infrastructures, this often includes the transfer of data from operational systems to analytical systems, which is often still treated rather neglected and offers many pitfalls.
Many data initiatives fail to obtain reliable and consistent data. There are many different approaches to solving this problem, some of which are borrowed from well-known software engineering best practices. I would like to present a few of these solutions and, above all, show how we have implemented them in practice. Without the costs exploding.
Of course, the talk is primarily aimed at data architects and data engineers. But I also invite all software engineers to take a look at the topic. They too are increasingly coming into contact with the supply of data and can make a big difference. Plus knowing about whats possible there could make their lives easier.
Ideas on how I can provide data better in the sense of more stable and correct. Tangible possibilities of which tools and processes can be used to implement this. Without sinking into complexity and ultimately costs.
I believe that one of the big issues will be the reduction of complexity. Nowadays, modern IT landscapes are a patchwork of barely comprehensible building blocks that are somehow held together. In addition to maintenance, this also makes it difficult to test and implement new features in order to quickly adapt to the market. In this context, automation (e.g. through AI) will play a major role on both the business and technical side.
I was already familiar with duckdb, but in 2023 I understood in Hannes Mühleisen's presentation what possibilities the technology offers and what diverse applications it enables. That was the first time I was able to imagine the changes in data architectures that are possible thanks to the new, lean processing frameworks.
Topics
QCon London 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Modern Data Architectures Hosted by Fabiane Nardon Data Expert, Java Champion & Data Platform Director @totvsFrom the same track
Wednesday 9 April
10:35 Whittle (3rd Fl.) Session Data Architecture Reliable Data Flows and Scalable Platforms: Tackling Key Data Challenges Matthias Niehoff Head of Data and Data Architecture @codecentric AG, iSAQB Certified Professional for Software Architecture There are a few common and mostly well-known challenges when architecting for data. For example, many data teams struggle to move data in a stable and reliable way from operational systems to analytics systems. 11:45 Whittle (3rd Fl.) Session AI/ML Achieving Precision in AI: Retrieving the Right Data Using AI Agents Adi Polak Director, Advocacy and Developer Experience Engineering @Confluent, Author of "Scaling Machine Learning with Spark" and "High Performance Spark 2nd Edition" In the race to harness the power of generative AI, organizations are discovering a hidden challenge: precision. 13:35 Fleming (3rd Fl.) Session Panel: Modern Data Architectures 14:45 Mountbatten (6th Fl.) Session AI/ML The Data Backbone of LLM Systems Paul Iusztin Senior ML/AI Engineer, MLOps, Founder @Decoding ML Any LLM application has four dimensions you must carefully engineer: the code, data, models and prompts. Each dimension influences the other. That's why you must learn how to track and manage each. The trick is that every dimension has particularities requiring unique strategies and tooling. 15:55 Whittle (3rd Fl.) Session Data Architecture Beyond the Warehouse: Why BigQuery Alone Won’t Solve Your Data Problems Sarah Usher Data & Backend Engineer, Community Director, Mentor Many organizations mistake the adoption of a data warehouse, like BigQuery, as the golden ticket to solving all their data challenges. But without a robust data strategy and architecture, you’re simply shifting chaos into the cloud.