> Markdown version of [/jobs/ext/2714953-sr-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2714953-sr-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Machine Learning Engineer - **Company:** Illumio - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Big Data, Cloud Computing, Distributed Systems, Memory Management, Python (Programming Language), Performance Tuning, Data Processing, System Availability, Large Language Models, Multi-Agent Systems, Database Optimization, Prompt Engineering, Apache Spark, Event Driven Architecture, Kubernetes, Apache Flink, Real Time Data, Apache Kafka, TensorRT, Virtual Agents, Terraform, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/sr-machine-learning-engineer-illumio-com-8147297 ## About the Role * 5-8 years of experience in backend engineering using Java, Python, or Go. * Expertise in distributed systems, asynchronous architectures (Kafka), and large-scale data processing (Spark/Flink). * Hands on experience with agentic frameworks (e.g., AutoGen, CrewAI, or custom orchestration layers), RAG, MCP, fine tuning models and prompt engineering. * Agentic observability using Langfuse, Evals frameworks for Testing/Resilience Bonus Points: * Advanced IaC: Expertise in building reusable Terraform modules and managing complex multi-region cloud deployments. * Vector DB Optimization: Deep experience in indexing strategies (HNSW vs IVF) and performance tuning for high-concurrency vector databases at scale. * AI Ops: Experience with LLM deployment optimization (e.g., vLLM, TensorRT-LLM) or managing proprietary model inference endpoints. This position involves access to software/technology that is subject to U.S. export controls. Any job offer made will be contingent upon the applicant's capacity to serve in compliance with U.S. export controls #LI-TD1 #LI-ONSITE ## Description As a Senior Software Engineer, you will architect high-scale distributed systems that process massive data volumes to fuel our Agentic AI ecosystem. You will lead the development of autonomous agents that don't just provide analytics, but take action-driving complex automation and insights for our enterprise customers. Your Impact: * Asynchronous Systems: Architect and optimize high-throughput, event-driven systems using Apache Kafka to handle real-time data flows. * Data Processing at Scale: Build and maintain large-scale data pipelines using Apache Spark or Flink to provide the high-volume analytics that power our AI. * Agentic Systems at Scale: Design sophisticated AI Agents capable of autonomous planning, memory management, and high-reliability tool-use across distributed environments. * Infrastructure & Orchestration: Lead the architectural design of containerized services on Kubernetes, ensuring high availability and scalability across Cloud Infrastructure (AWS/Azure/GCP). ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Agentic AI Systems for Critical Workloads](https://www.wearedevelopers.com/videos/1592-agentic-ai-systems-for-critical-workloads) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)