> Markdown version of [/jobs/ext/1889718-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1889718-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Illumio - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Big Data, Software as a Service, Cloud Computing, Distributed Systems, Memory Management, Electronic Signatures, Internet Security, Python (Programming Language), Machine Learning, Performance Tuning, Software Engineering, Systems Architecture, Data Processing, Scripting, Google Cloud, System Availability, Large Language Models, Multi-Agent Systems, Concurrency, Database Optimization, Prompt Engineering, Apache Spark, Malware, HybridCloud, Event Driven Architecture, Kubernetes, Apache Flink, Real Time Data, Apache Kafka, Data Management, TensorRT, Virtual Agents, Terraform, Data Pipelines - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/sr-machine-learning-engineer-sunnyvale-ca--65987dba-9d51-4f4c-a760-41f07c3bff4a ## About the Role * 5-8 years of experience in backend engineering using Java, Python, or Go. * Expertise in distributed systems, asynchronous architectures (Kafka), and large-scale data processing (Spark/Flink). * Hands on experience with agentic frameworks (e.g., AutoGen, CrewAI, or custom orchestration layers), RAG, MCP, fine tuning models and prompt engineering. * Agentic observability using Langfuse, Evals frameworks for Testing/Resilience Bonus Points: * Advanced IaC: Expertise in building reusable Terraform modules and managing complex multi-region cloud deployments. * Vector DB Optimization: Deep experience in indexing strategies (HNSW vs IVF) and performance tuning for high-concurrency vector databases at scale. * AI Ops: Experience with LLM deployment optimization (e.g., vLLM, TensorRT-LLM) or managing proprietary model inference endpoints., Amazon Web Services (AWS), Apache Kafka, Apache Spark, Architectural Design, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Automation, Cloud Computing, Concurrency, Data Management, Data Processing, Database Optimization, Distributed Computing, Ecosystems, GCP (Good Clinical Practices), High Availability, High Reliability, High Throughput, Hybrid Cloud, Internet Security, Java, Leadership, MCP - Microsoft Certified Professional, Machine Learning, Memory Management, Microsoft Windows Azure, Performance Tuning/Optimization, Philosophy, Python Programming/Scripting Language, Ransomware, Security Attacks, Software Engineering, Software as a Service (SaaS), System Architecture ## Description Our guiding philosophy in Engineering is to get things right through practicing disciplined engineering, focusing, not cutting corners, and of course having fun while we are at it. We believe in enabling ownership at all levels of the organization and empowering teams. If you thrive in this culture, come join us!As a Senior Software Engineer, you will architect high-scale distributed systems that process massive data volumes to fuel our Agentic AI ecosystem. You will lead the development of autonomous agents that don't just provide analytics, but take action-driving complex automation and insights for our enterprise customers. Your Impact: * Asynchronous Systems: Architect and optimize high-throughput, event-driven systems using Apache Kafka to handle real-time data flows. * Data Processing at Scale: Build and maintain large-scale data pipelines using Apache Spark or Flink to provide the high-volume analytics that power our AI. * Agentic Systems at Scale: Design sophisticated AI Agents capable of autonomous planning, memory management, and high-reliability tool-use across distributed environments. * Infrastructure & Orchestration: Lead the architectural design of containerized services on Kubernetes, ensuring high availability and scalability across Cloud Infrastructure (AWS/Azure/GCP)., This position involves access to software/technology that is subject to U.S. export controls. Any job offer made will be contingent upon the applicant's capacity to serve in compliance with U.S. export controls #LI-TD1 #LI-ONSITE Our Commitment Illumio believes that an environment of unique backgrounds, experiences, viewpoints, and individual contributions creates a culture of belonging, drives our future, and makes us stronger together in support of our customers and their success. All official job offers from our company are extended directly by our recruitment team and will be sent through an official E-Signature document for your review and signature. Please be aware that we do not ask for any personal information in the process of extending offers of employment, such as financial details or social security numbers. Upon acceptance of any offer, we will request such information as part of the onboarding process prior to or on your first day of employment, and only after completing a background check through an authorized third-party vendor. If you receive any communication asking for personal details outside of these processes, please contact us immediately to verify the authenticity of the request. Your security is important to us, and we are committed to a safe and transparent hiring experience. For roles in San Francisco and Los Angeles: Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Illumio will consider for employment qualified applicants with arrest and conviction records. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Full Spectrum File Uploads](https://www.wearedevelopers.com/videos/870-full-spectrum-file-uploads) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)