Rafał Kuć
Software engineer, trainer, consultant and author from time to time – some would say that he is an all in one battle weapon concentrated on information retrieval, performance and user search experience. However he also likes all the other cool stuff that is happening in the IT world. Likes to share his knowledge by giving talks at various meet ups and conferences.
Sessions
Which GPU for Local LLMs?
Short Talk
16. June 2025, 10:40 - 11:00
Kesselhaus
You’re using local LLMs. For example, to power RAG. You want to deploy them in production, but you don’t know where: which type of GPU? How large should it be? Should you use a larger model but quantize more aggressively?
Our benchmark results and their interpretation will give you some answers.