Speaker
Abstract
Managing vector data entails storing, updating, and searching collections of large and multi-dimensional pieces of data. Some believe that this justifies the creation of a new class of data systems specialized for this. Others would contend that such systems would eventually need to provide services provided by database system, including e.g., transaction management, role-based access control, and integration of vector search predicates in complex queries.
Recent research (PDX - Partition Dimension Across) has shown that already highly optimized vector search kernels can profit from columnar storage. This talk gives a sneak preview of our ongoing work in this area, including optimized vector ingest, tailored vector indexing, and integrated evaluation of queries and vector predicates in the DuckDB system.
Topics
QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Modern Performance Optimization Hosted by Chaitanya Bhandari Distributed Systems Engineer @TigerBeetleFrom the same track
Wednesday 18 March
10:35 Windsor (5th Fl.) Session AI/ML Machine Learning at the Edge of Scale and Speed: Nanosecond Inference at the CERN Large Hadron Collider Thea Klaeboe Aarrestad Particle Physics and Real-Time ML @CERN @ETH Zürich The CERN Large Hadron Collider (LHC) produces O(10,000) exabytes of raw data annually from high-energy proton collisions. Handling this volume under strict compute and storage limits requires real-time event filtering capable of processing millions of collisions per second. 11:45 Windsor (5th Fl.) Session Data Systems Vector Search on Columnar Storage Peter Boncz Professor @CWI, Co-Creator of MonetDB, VectorWise and MotherDuck, Database Systems Researcher, and Entrepreneur Managing vector data entails storing, updating, and searching collections of large and multi-dimensional pieces of data. Some believe that this justifies the creation of a new class of data systems specialized for this. 13:35 Mountbatten (6th Fl.) Session architecture Not Just I/O: Using Async/Await for Computational Scheduling Orson Peters Senior Engineer of Query Execution @Polars, (Co-)Author of Stdlib Sort in Rust & Go In the past two years I have developed a new query execution engine for Polars, which not only tries to execute as much of your query in parallel as possible, but in a streaming fashion as well, such that you can process data sets which do not fit in memory. 14:45 Mountbatten (6th Fl.) Session Data Management Looking Under the Hood: Data Processing Systems Performance Tricks (and How to Apply Them to Your Code) Holger Pirk Associate Professor for Data Management Systems at Imperial College London and Avid Runner — Minimizing Cache Misses, Thread Divergence and Aerobic Decoupling Modern data processing systems—databases, analytics engines, vector stores, and stream processors—hide an extraordinary amount of performance engineering beneath their abstractions. 15:55 Windsor (5th Fl.) Session compilers Automatically Retrofitting JIT Compilers Laurence Tratt Shopify / Royal Academy of Engineering Research Chair in Language Engineering @King's College London We as a community have attempted, multiple times, to speed up languages such as Lua, Python, and Ruby by hand-writing JIT compilers. Sometimes we've had short-term success, but the size, and pace of change, of their standard implementations has proven difficult to keep up with over time.