> Markdown version of [/jobs/ext/3398867-staff-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3398867-staff-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Infrastructure Engineer - **Company:** Cognizant Technology Solutions Corporation - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $130,000.0 - $145,000.0 - **Contract:** Permanent contract - **Skills:** Big Data, Business Software, Cloud Engineering, Distributed Systems, Machine Learning, Open Source Technology, Reliability Engineering, Tensorflow, Azure Machine Learning, Software Engineering, Data Processing, System Availability, Apache Spark, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Machine Learning Operations, Data Pipelines, Apache Beam - **Published:** September 17, 2026 - **Apply:** https://dejobs.org/x/x/186B016B05BD4154BB41F1D1ECA4E7A7/job/ ## About the Role * 10+ years of software engineering experience, including 5+ years focused on machine learning infrastructure, MLOps, or platform engineering. * Deep expertise designing and operating distributed systems and large-scale data processing platforms. * Proven experience building and supporting production ML platforms that power mission-critical business applications. * Hands-on experience with technologies such as Apache Spark, Apache Beam, feature stores, and large-scale data pipelines. * Strong knowledge of ML serving and inference platforms, including TorchServe, TensorFlow Serving, Triton, or similar technologies. * Experience with containerization, orchestration platforms, cloud-native architectures, and infrastructure automation. * Demonstrated success leading complex technical initiatives and influencing architecture decisions across engineering organizations. These will help you stand out * Experience designing and optimizing GPU-based infrastructure for machine learning workloads. * Background in high-performance computing environments. * Expertise in ML observability, monitoring, reliability engineering, and operational excellence. * Contributions to open-source machine learning infrastructure, platform engineering, or MLOps projects. * Master's degree in computer science, Engineering, or a related technical field. * Experience supporting security, governance, and compliance requirements within ML ecosystems. We're excited to meet people who share our mission and can make an impact in a variety of ways. Don't hesitate to apply, even if you only meet the minimum requirements listed. Think about your transferable experiences and unique skills that make you stand out as someone who can bring new and exciting things to this role. ## Description As a Staff Machine Learning Infrastructure Engineer , you will make an impact by architecting and advancing scalable machine learning infrastructure that enables the development, deployment, and operation of ML models at enterprise scale. You will be a valued member of the ML Platform Engineering team and work collaboratively with data scientists, machine learning engineers, platform engineers, and cross-functional technology stakeholders to deliver reliable, high-performance ML systems., * Architect and evolve end-to-end machine learning platforms supporting data processing, feature management, training, and model serving. * Design and implement scalable infrastructure that enables efficient development, deployment, and operation of ML workloads. * Drive technical strategy, standards, and best practices for machine learning infrastructure, automation, and platform reliability. * Lead the design of real-time and batch inference solutions that support high-volume production workloads. * Mentor engineers, lead technical reviews, and partner with data science teams to accelerate ML adoption and innovation across the organization. Work model At Cognizant, we strive to provide flexibility wherever possible, and we are here to support a healthy work-life balance through our various well-being programs. Based on this role's business requirements, this is an onsite position requiring 5 days per week in a client or Cognizant office in Sunnyvale, CA . The working arrangements for this role are accurate as of the date of posting. This may change based on the project you're engaged in, as well as business and client requirements. Rest assured; we will always be clear about role expectations. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)