Remote

MAG 24 LLC
New York, NY, United States
5 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$350,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Big Data Cloud Computing Cloud Engineering Continuous Integration DevOps Operational Databases Multi-Cloud Kubernetes Infrastructure Automation Frameworks Machine Learning Operations
+2 more
Terraform Data Pipelines

Job description

We are sharing a full-time opportunity for an experienced Director of Infrastructure Engineering with deep expertise in AWS, GCP, infrastructure as code, CI/CD, platform engineering, reliability, security, and technical leadership to build and scale infrastructure supporting production AI systems. The role combines hands-on infrastructure engineering with strategic leadership across cloud architecture, developer platforms, observability, reliability, security, and engineering operations., Cloud Infrastructure & Platform Strategy

  • Own multi-cloud architecture across AWS and GCP
  • Define infrastructure strategy around scalability, reliability, security, and cost
  • Build and maintain infrastructure as code using Terraform or comparable tooling
  • Develop reusable platform abstractions, automation, and internal infrastructure tooling
  • Improve developer productivity while maintaining strong operational standards

Reliability, Delivery & Observability

  • Design and improve CI/CD systems for fast, reliable, and secure software delivery
  • Establish SLOs, error budgets, incident-response processes, and on-call practices
  • Lead disaster-recovery and resilience initiatives
  • Build observability across metrics, logs, traces, alerting, and operational signals
  • Use production data and postmortems to improve reliability and reduce deployment risk

Security, Operations & Leadership

  • Embed security into cloud architecture, platform tooling, and software-delivery workflows
  • Support compliance with frameworks such as ISO 27001, SOC 2, and CMMC
  • Lead and develop a high-performing Infrastructure or Platform Engineering team
  • Mentor engineers and influence infrastructure strategy across technical and executive stakeholders
  • Balance long-term platform strategy with hands-on production and incident-management responsibilities

Requirements

  • 8+ years of experience in production infrastructure, platform engineering, DevOps, or SRE
  • 3+ years of engineering leadership experience
  • Deep expertise with AWS, GCP, or multi-cloud production environments
  • Strong Terraform or comparable infrastructure-as-code experience
  • Strong knowledge of Kubernetes and containerised infrastructure
  • Experience designing and operating modern CI/CD platforms
  • Demonstrated success building highly available, observable, and resilient systems
  • Strong understanding of infrastructure security, compliance, and operational risk
  • Experience scaling infrastructure and engineering teams in fast-moving environments
  • Excellent written and verbal communication and ability to influence technical strategy
  • AI/ML infrastructure or large-scale data-platform experience is highly valuable
  • Familiarity with model training, inference, evaluation, or data-pipeline infrastructure is advantageous
  • Experience with FedRAMP, GovCloud, CMMC Level 2, or comparable regulated environments is beneficial

Benefits & conditions

Engagement Details

  • Full-time engagement
  • Fully remote
  • Base compensation: $350,000-$500,000/year
  • Work will involve AWS, GCP, infrastructure as code, CI/CD, observability, reliability engineering, platform tooling, security, and technical leadership
  • Responsibilities will span strategic architecture, production operations, developer experience, team leadership, and incident management
  • The role may support AI/ML infrastructure, large-scale data platforms, or regulated cloud environments
  • Infrastructure priorities, compliance requirements, and platform architecture may evolve as production systems scale

About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all