Reliable Retrieval for Production AI Systems

QCon London 2026

Session AI/LLM

Reliable Retrieval for Production AI Systems

Tuesday Mar 17 / 10:35AM GMT, Whittle (3rd Fl.) at The QEII Centre, London

Abstract

Search is central to many AI systems. Everyone is building RAG and agents right now, but few are building reliable retrieval systems.

Drawing from our real world RAG system built on 10K+ documents, used by 300+ users, we found that most RAG failures can be traced back to two things: indexing and retrieval. This talk shares what matters when building production retrieval systems, starting from effective document parsing, chunking, indexing to reliable search and retrieval pipeline. I will present specific implementation details, toolings, and how we solve the challenges encountered.

Here is one truth: you know who the real boss of NLP is? A PDF! Retrieval is only as good as the documents you index. We will cover how to handle document layout nightmares and the decisions around parsing and chunking strategies. Your system might not even need chunking, adding it could hurt performance. I will show when to chunk and how to find the "chunking sweet spot.”

Once your documents are indexed, the next challenge is search itself. Many retrieval systems see the entire world as just strings. But real user queries carry non-textual signals such as time or numbers which also needs to be encoded. We demonstrate how search can be combined with temporal scoring to capture user's time intent and when agentic search could help in retrieval.

Finally, let's not forget evals. If you can't trust your evals, how can you trust your AI system? I will share our approach to building good evaluation sets from working with stakeholders to capture real failure modes, to using bootstrapping to determine how many samples you actually need.

Topics

AI/LLM Search and Retrieval rag
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 March

10:35 Whittle (3rd Fl.) Session AI/LLM Reliable Retrieval for Production AI Systems Lan Chu AI Tech Lead and Senior Data Scientist 11:45 Fleming (3rd Fl.) Session AI Rewriting All of Spotify's Code Base, All the Time Jo Kelly-Fenton, Aleksandar Mitic 13:35 Churchill (Ground Fl.) Session AI/ML Refreshing Stale Code Intelligence Jeff Smith CEO & Co-Founder @Neoteny AI, AI Engineer, Researcher, Author, Ex-Meta/FAIR 14:45 Churchill (Ground Fl.) Session AI Beyond Context Windows: Building Cognitive Memory for AI Agents Karthik Ramgopal Distinguished Engineer & Tech Lead of the Product Engineering Team @LinkedIn, 15+ Years of Experience in Full-Stack Software Development 15:55 Fleming (3rd Fl.) Session applied ai Building an AI Gateway Without Frameworks: One Platform, Many Agents Amit Navindgi, Jatin Aneja 17:05 Whittle (3rd Fl.) Session Async Agents in Production: Failure Modes and Fixes Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist