> Markdown version of [/jobs/ext/3018324-remote](https://www.wearedevelopers.com/jobs/ext/3018324-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** MAG 24 LLC - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Salary:** $350,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Big Data, Cloud Computing, Cloud Engineering, Continuous Integration, DevOps, Operational Databases, Multi-Cloud, Kubernetes, Infrastructure Automation Frameworks, Machine Learning Operations, Terraform, Data Pipelines - **Published:** September 20, 2026 - **Apply:** https://www.careerjet.com/jobad/us95b5f69f2e09afdfe9b50fa73eedd098 ## About the Role * 8+ years of experience in production infrastructure, platform engineering, DevOps, or SRE * 3+ years of engineering leadership experience * Deep expertise with AWS, GCP, or multi-cloud production environments * Strong Terraform or comparable infrastructure-as-code experience * Strong knowledge of Kubernetes and containerised infrastructure * Experience designing and operating modern CI/CD platforms * Demonstrated success building highly available, observable, and resilient systems * Strong understanding of infrastructure security, compliance, and operational risk * Experience scaling infrastructure and engineering teams in fast-moving environments * Excellent written and verbal communication and ability to influence technical strategy * AI/ML infrastructure or large-scale data-platform experience is highly valuable * Familiarity with model training, inference, evaluation, or data-pipeline infrastructure is advantageous * Experience with FedRAMP, GovCloud, CMMC Level 2, or comparable regulated environments is beneficial ## Description We are sharing a full-time opportunity for an experienced Director of Infrastructure Engineering with deep expertise in AWS, GCP, infrastructure as code, CI/CD, platform engineering, reliability, security, and technical leadership to build and scale infrastructure supporting production AI systems. The role combines hands-on infrastructure engineering with strategic leadership across cloud architecture, developer platforms, observability, reliability, security, and engineering operations., Cloud Infrastructure & Platform Strategy * Own multi-cloud architecture across AWS and GCP * Define infrastructure strategy around scalability, reliability, security, and cost * Build and maintain infrastructure as code using Terraform or comparable tooling * Develop reusable platform abstractions, automation, and internal infrastructure tooling * Improve developer productivity while maintaining strong operational standards Reliability, Delivery & Observability * Design and improve CI/CD systems for fast, reliable, and secure software delivery * Establish SLOs, error budgets, incident-response processes, and on-call practices * Lead disaster-recovery and resilience initiatives * Build observability across metrics, logs, traces, alerting, and operational signals * Use production data and postmortems to improve reliability and reduce deployment risk Security, Operations & Leadership * Embed security into cloud architecture, platform tooling, and software-delivery workflows * Support compliance with frameworks such as ISO 27001, SOC 2, and CMMC * Lead and develop a high-performing Infrastructure or Platform Engineering team * Mentor engineers and influence infrastructure strategy across technical and executive stakeholders * Balance long-term platform strategy with hands-on production and incident-management responsibilities ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)