> Markdown version of [/jobs/ext/1319908-mlops-ai-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1319908-mlops-ai-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # MLOps & AI Infrastructure Engineer - **Company:** Altera Corporation - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $149,100.0 - $215,925.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Airflow, Amazon Web Services, Artificial Neural Networks, Microsoft Azure, Cloud Computing, Computer Programming, Continuous Integration, Information Engineering, DevOps, Programming Tools, Machine Learning, Performance Tuning, Tensorflow, Zero Trust Network Access, Azure Machine Learning, Software Engineering, Unstructured Data, Management of Software Versions, AI Infrastructure, Reinforcement Learning, Google Cloud, Cloud Platform System, High Performance Computing, Feature Engineering, Pytorch, Transfer Learning, Large Language Models, Deep Learning, Generative AI, Cloudformation, Containerization, Data Lakes, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Data Lineage, HuggingFace, Free and Open-Source Software, Slurm, Machine Learning Operations, Virtual Agents, Terraform, Devsecops, Docker, Data Generation - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/b82dc382-16f6-45bf-8d2d-d41799c5deea ## About the Role Strong ownership mindset - you drive ML initiatives from prototype to production without being asked Bias toward automation: if you do it twice, you automate it Ability to bridge research and engineering - translating papers into production-grade systems Thrives in fast-paced, ambiguous environments typical of deep-tech and semiconductor companies Clear communicator who can explain complex ML concepts to non-technical stakeholders, * Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or related field and 10+ years of industry experience * 10+ years of experience across ML engineering, data science, and MLOps - including frameworks (PyTorch, TensorFlow, JAX, Hugging Face) and production model deployment at scale * 8+ years of experience experience with parallelism strategies (FSDP, DeepSpeed, data/model parallelism) * 10+ years of experience and proficiency in Python programming * 8+ years of experience in cloud ML platforms (AWS, Google Cloud Platform, Azure), Docker/Kubernetes, and CI/CD pipelines * 5+ years of hands-on experience with MLflow, W&B, or Neptune for tracking and reproducibility, * Phdin Computer Science, Machine Learning, Statistics, or related field * Experience applying ML/AI to semiconductor, EDA, or chip design domains (e.g., timing prediction, place & route optimization, DRC closure) * Familiarity with HPC schedulers such as LSF or Slurm and GPU cluster management for training workloads * Knowledge of LLM fine-tuning, Retrieval-Augmented Generation (RAG) architectures, and AI agent frameworks such as LangChain or AutoGen * Experience with graph neural networks (GNNs) or geometric deep learning for circuit and netlist analysis * Background in reinforcement learning for optimization problems * Exposure to zero-trust security, DevSecOps, and compliance automation for ML systems * Experience working with large-scale simulation pipelines and synthetic data generation * Experience at organizations such as NVIDIA, AMD, Intel, Google DeepMind, or similar AI/HPC-focused companies * Published research or open-source contributions in ML, MLOps, or AI for EDA * Experience building AI-powered developer tools or copilot-style products * Familiarity with Synopsys, Cadence, or Siemens EDA toolchains and associated data formats ## Description We are looking for a Senior MLOps & AI Infrastructure Engineer to architect, build, and operationalize machine learning systems at scale. This role sits at the intersection of data science, software engineering, and infrastructure - combining deep ML expertise with the DevOps/MLOps discipline required to ship models reliably into production. You will partner closely with software, data, and infrastructure teams to design end-to-end ML pipelines, automate model lifecycle management, and deliver AI-powered capabilities across our EDA, HPC, and cloud environments. Key Responsibilities: ML Platform & Pipeline Engineering Design, build, and maintain scalable ML pipelines for training, evaluation, and deployment across cloud and on-prem HPC environments Build MLOps infrastructure including experiment tracking, model registry, feature stores, and automated retraining workflows Implement CI/CD/CT (Continuous Training) pipelines for ML models using tools such as Kubeflow, MLflow, Airflow, or similar Containerize ML workloads with Docker and orchestrate at scale using Kubernetes and GPU node pools Model Development & Optimization Develop, fine-tune, and deploy large-scale models including LLMs, GNNs, and reinforcement learning agents for EDA and chip design applications Apply advanced techniques: transfer learning, quantization, pruning, distillation, and RLHF for production-grade model efficiency Implement A/B testing frameworks and shadow deployments for safe model rollout Benchmark and optimize model inference performance on GPU/TPU clusters Data Engineering & Feature Management Build and maintain data pipelines for large-scale structured and unstructured datasets (terabyte-scale) Collaborate with data teams to design feature engineering systems and maintain data quality for ML training Implement data versioning and lineage tracking (DVC, Delta Lake, or similar) Infrastructure & Operations Manage cloud ML infrastructure on AWS (SageMaker), Azure (AML), or Google Cloud Platform (Vertex AI) with cost and performance optimization Automate infrastructure provisioning using Terraform or CloudFormation for GPU-backed ML environments Build monitoring, alerting, and observability systems for model performance drift, data quality, and system health Support HPC schedulers (LSF, Slurm) for large-scale distributed training jobs Collaboration & Leadership Partner with research scientists to productionize experimental models with engineering rigor Mentor junior engineers and define ML engineering best practices across the organization ## Related Videos - [MLOps - What’s the deal behind it?](https://www.wearedevelopers.com/videos/392-mlops-what-s-the-deal-behind-it) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)