Navigating LLM Deployment: Tips, Tricks, and Techniques

QCon London 2024

Session AI/ML

Navigating LLM Deployment: Tips, Tricks, and Techniques

Monday Apr 8 / 11:45AM BST, Mountbatten (6th Fl.)

Abstract

Self-hosted Language Models are going to power the next generation of applications in critical industries like financial services, healthcare, and defence. Self-hosting LLMs, as opposed to using API-based models, comes with its own host of challenges - as well as needing to solve business problems, engineers need to wrestle with the intricacies of model inference, deployment and infrastructure. In this talk we are going to discuss the best practices in model optimisation, serving and monitoring - with practical tips and real case-studies.

Interview

At TitanML our focus is on making Generative AI applications easier to develop, deploy and serve. A large focus of our work recently is making it easier to build applications that involve both RAG and JSON constrained outputs. 

Almost every business is trying to build and deploy LLM applications at the moment, however very few of them have successfully got these applications into production. Our teams are experts in deploying and serving LLM apps so we have a lot of tips and tricks to help other developers avoid common pitfalls. 

This session is interesting for those working with or thinking of building with Generative AI, especially self-hosted open source AI. It is not a 'code-along' session, however there may be some technical concepts. 

I want this persona to realize that deploying LLMs within your own environment is a viable option and is not as scary as it might appear!

Topics

AI/ML LLM Deployment Inference Infrastructure
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon London 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 8 April

10:35 Windsor (5th Fl.) Session AI/ML Retrieval-Augmented Generation (RAG) Patterns and Best Practices Jay Alammar Director & Engineering Fellow @Cohere & Co-Author of "Hands-On Large Language Models" 11:45 Mountbatten (6th Fl.) Session AI/ML Navigating LLM Deployment: Tips, Tricks, and Techniques Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist 13:35 Mountbatten (6th Fl.) Session AI/ML Reach Next-Level Autonomy with LLM-Based AI Agents Tingyi Li Enterprise Solutions Architect @AWS 14:45 Mountbatten (6th Fl.) Session AI/ML LLM and Generative AI for Sensitive Data - Navigating Security, Responsibility, and Pitfalls in Highly Regulated Industries Stefania Chaplin, Azhir Mahmood 15:55 Whittle (3rd Fl.) Session AI/ML The AI Revolution Will Not Be Monopolized: How Open-Source Beats Economies of Scale, Even for LLMs Ines Montani Co-Founder & CEO @Explosion, Core Developer of spaCy 17:05 Fleming (3rd Fl.) Session AI/ML How Green is Green: LLMs to Understand Climate Disclosure at Scale Leo Browning First ML Engineer @ClimateAligned