> Markdown version of [/jobs/48414-senior-ai-serving-engineer-backend](https://www.wearedevelopers.com/jobs/48414-senior-ai-serving-engineer-backend). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Serving Engineer, Backend - **Company:** Sciforium - **Location:** San Francisco, United States - **Experience:** Expert - **Salary:** €190,000.0 - €250,000.0 - **Skills:** Python, Rust - **Published:** August 21, 2026 - **Apply:** https://jobs.ashbyhq.com/Sciforium/12b31a2e-efd4-43b3-b610-64d37b4cade3 ## About the Role Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications. ## **About the role** This role offers a unique opportunity to work on the core systems that power Sciforium’s multimodal AI models. You’ll help build the model serving platform working across C++, Python, runtime execution, and distributed infrastructure to create a fast, reliable engine for real-time AI applications. You’ll gain hands-on experience with performance engineering, learn how large AI models are optimized and deployed at scale, and collaborate closely with ML researchers and experienced systems engineers. If you enjoy low-level programming, care deeply about performance, and want exposure to the full AI stack, this role provides both high-impact work and strong growth potential. ## **What you'll do** - Build the model serving platform, including API, Control Plane, Billing, Monitoring, and distributed inference features. - Collaborate with ML researchers to integrate new multimodal models into production workflows. - Write reliable, maintainable code with strong testing and documentation practices. - Provide operational support for keeping our production services highly performant, available and reliable - Help troubleshoot complex issues across runtime, service, and GPU layers, working closely with other engineers. ## **Ideal candidate profile** - Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience) - 3+ years of software engineering experience, with a focus on infrastructure or machine learning systems. - Strong proficiency in C++/Python/Go/Rust - Experience with Kubernetes, Containerization - Experience in building large scale ML/MLOps infrastructure - Strong collaboration and communication skills, with the ability to work effectively across engineering and ML teams. - Comfortable working from the office and contributing to a fast-moving, high-ownership team culture. ## Description ## **Nice-to-have** - Experience with ML systems engineering, open source inference engine like vLLM, Sglang, or TRT-LLM - Proficiency in CUDA or ROCm and experience with GPU profiling tools - Contributions to open-source ML or HPC infrastructure ## About Sciforium [Company profile](https://www.wearedevelopers.com/companies/4292-sciforium) ### More Jobs at Sciforium - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) - [LLM Dataset Engineer](https://www.wearedevelopers.com/jobs/48419-llm-dataset-engineer) - [Model Implementation Engineer](https://www.wearedevelopers.com/jobs/48421-model-implementation-engineer) - [GPU Kernel Engineer](https://www.wearedevelopers.com/jobs/48412-gpu-kernel-engineer) - [ML Engineer](https://www.wearedevelopers.com/jobs/48422-ml-engineer) ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [CUDA in Python](https://www.wearedevelopers.com/videos/1294-cuda-in-python) - [Rust in the Real World: Adoption, Migration, and Tradeoffs](https://www.wearedevelopers.com/videos/100147-rust-in-the-real-world-adoption-migration-and-tradeoffs) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Full Stack Web Apps With Nothing But Python](https://www.wearedevelopers.com/videos/417-full-stack-web-apps-with-nothing-but-python) - [Eternal Sunshine of the Spotless Programming Language](https://www.wearedevelopers.com/videos/847-eternal-sunshine-of-the-spotless-programming-language) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Building AI Solutions with Rust and Docker](https://www.wearedevelopers.com/magazine/494-building-ai-solutions-with-rust-and-docker) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)