Deploy MultiModal RAG Systems with vLLM

QCon London 2025

Session AI/ML

Deploy MultiModal RAG Systems with vLLM

Tuesday Apr 8 / 10:35AM BST, Whittle (3rd Fl.)

Abstract

While text-based RAG systems have been everywhere in the last year and a half, there is so much more than text data. Images, audio, and documents often need to be processed together to provide meaningful insights, yet most RAG implementations focus solely on text. Think automated visual inspection systems understanding both manufacturing logs and production line images, or robotics systems correlating sensor data with visual feedback. These multimodal scenarios demand RAG systems that go beyond text-only processing.

In this talk, we'll talk about how one can build a MultiModal RAG system that helps solve this problem. We'll explore the architecture that makes it possible to run such a system and demonstrate how to build one using Milvus, LlamaIndex, and vLLM for deploying open-source LLMs on your own infrastructure.

Through a live demo, we'll showcase a real-world application processing both images and text queries. Whether you're looking to reduce API costs, maintain data privacy, or simply gain more control over your AI infrastructure, this session will provide you with actionable insights to implement MultiModal RAG in your organization.

Interview

I focus on GenAI usage, going from simple RAG systems to full Agentic ones. I also highlight how search works at Scale. I am a big open source fan so most of my work is focused around that. 

To showcase that you can deploy open source apps that can be very good. The idea is to showcase to people that they can be in control and not dependent on closed source systems.

People interested in moving from OpenAI and they want to control their GenAI stack. Also people interested in Multimodality.

Learn how specific open source tools like vLLMs can match or exceed proprietary solutions while giving you full control over your AI stack and SLAs.

I believe that the combination of open source AI models and rapid development tools will enable more customized, sovereign AI solutions.

People will have their own unique version of their software and likely not rely as much on the typical apps we used to have. 

Topics

AI/ML vLLM K8s
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 8 April

10:35 Whittle (3rd Fl.) Session AI/ML Deploy MultiModal RAG Systems with vLLM Stephen Batifol Developer Advocate @Zilliz, Founding Member of the MLOps Community Berlin, Previously Machine Learning Engineer @Wolt, and Data Scientist @Brevo 11:45 Rutherford (4th Fl.) Unconference Unconference: AI and ML for Software Engineers 13:35 Whittle (3rd Fl.) Session AI/ML AI for Food Image Generation in Production: How & Why Iaroslav Amerkhanov Senior Data Scientist @Delivery Hero, Founder of T4lky, Creator & Host of EPAM Podcast, Speaker 14:45 Whittle (3rd Fl.) Session AI/ML Foundation Models for Ranking: Challenges, Successes, and Lessons Learned Moumita Bhattacharya Senior Research Scientist @Netflix, Previously @Etsy 15:55 Churchill (Ground Fl.) Session AI/ML Building Embedding Models for Large-Scale Real-World Applications Sahil Dua Senior Software Engineer, Machine Learning @Google, Stanford AI, Co-Author of “The Kubernetes Workshop”, Open-Source Enthusiast 17:05 Whittle (3rd Fl.) Session AI/ML How to Unlock Insights and Enable Discovery Within Petabytes of Autonomous Driving Data Kyra Mozley ML Engineer @Wayve, Previously Security & AI PhD Candidate @Royal Holloway University