Riccardo Zoncada is a software engineer at xtream who explores adopting Small and Large Language Models for local computing. He focuses on the technical challenges of running AI inference at the edge.
Recently, he has been untangling the fast-moving ecosystem of LLM inference servers. Through his research and technical talks, he breaks down the concepts and helps developers navigate the tools required to spin up language models locally.