> Markdown version of [/jobs/ext/2652673-machine-learning-infrastructure-engineer-technology](https://www.wearedevelopers.com/jobs/ext/2652673-machine-learning-infrastructure-engineer-technology). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Infrastructure Engineer, Technology - **Company:** POINT72, L.P. - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $185,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Profiling, Software Debugging, Digital Architecture, Distributed Systems, Python (Programming Language), Key Management, Machine Learning, Reinforcement Learning, Google Cloud, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Machine Learning Operations, Terraform - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7e0c71dee7aff1a9 ## About the Role * Bachelor's or master's degree in computer science, electrical engineering, or a related technical field * 3-7 years of experience building and maintaining scalable compute or machine learning infrastructure systems * Deep understanding of distributed systems, container orchestration (Kubernetes), and public cloud platforms such as AWS, Google Cloud Platform, or Azure * Hands-on experience with machine learning operations and infrastructure tools such as MLflow, Ray, Airflow, Kubeflow, and Terraform * Strong understanding of reinforcement learning concepts and their infrastructure implications * Proficiency in Python and systems-level programming in one or more languages such as Go, C++, or Rust * Strong debugging, performance profiling, and optimization skills across GPU and CPU compute stacks * Experience implementing monitoring, observability, and cost-optimization for GPU/accelerator-based compute environments * Excellent collaboration and communication skills with a systems-thinking mindset * Commitment to the highest ethical standards ## Description * Design and implement high-performance infrastructure to support large-scale generative AI and machine learning workloads, enabling faster model iteration and real business impact * Design and operate distributed systems for model training, hyperparameter tuning, inference, and data preprocessing pipelines to deliver reliable end-to-end machine learning (ML) workflows * Collaborate with ML researchers and engineers to produce models, optimizing compute utilization, training throughput, and inference latency * Develop and automate deployment, orchestration, and CI/CD pipelines for models and data workflows using container orchestration and infrastructure-as-code (IaC) * Implement observability, monitoring, and cost-management strategies for GPU and accelerator compute environments to maintain predictable performance and spend * Evaluate, integrate, and benchmark emerging hardware and software technologies across cloud and on-prem environments to improve scalability and throughput * Drive security, compliance, and operational runbooks for GenAI infrastructure including access controls, secrets management, and incident response procedures * Troubleshoot, profile, and optimize performance across GPU and CPU compute stacks to remove bottlenecks and increase reliability * Document architecture, operational practices, and mentor engineers to expand team capability and accelerate adoption of production-ready GenAI infrastructure ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)