Lead Cloud Platform Engineer (Data & Execution Platform)

Fmr LLC
Westlake, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Remote
Westlake, United States of America

Tech stack

Java
Artificial Intelligence
Airflow
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Architectural Patterns
Big Data
Cloud Engineering
Software Quality
Computer Networks
Cron
Directed Acyclic Graph (Directed Graphs)
ETL
Software Debugging
Distributed Data Store
Distributed Systems
DNS
Fault Tolerance
Identity and Access Management
Subnetting
Python
Network Security
Network Troubleshooting
Networking Basics
Performance Tuning
Productivity Software
Systems Integration
TCP/IP
Data Logging
Transport Layer Security
Load Balancing
Cloud Platform System
GitHub Copilot
Autoscaling
Large Language Models
Grafana
Spark
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Kubernetes
Information Technology
Deployment Automation
Data Management
Functional Programming
Data Pipelines
Docker

Job description

We are seeking a hands-on Lead Cloud Platform Engineer to implement, scale, and operate cloud-native infrastructure and services that power large-scale data processing systems. This role focuses on translating defined architectures into production-grade platforms that are reliable, observable, secure, and performant. You will lead the implementation and operation of a modern execution platform built on Apache Spark for distributed compute and an Airflow orchestration layer and DAG execution environment. The ideal candidate brings deep production experience in Spark and Airflow, and excels at troubleshooting, tuning, and operationalizing distributed systems in AWS environments, while leveraging modern developer productivity tools such as AI-assisted coding and LLM-based workflows.

The Expertise and Skills You Bring

  • Implement and operate cloud-native platform services for distributed data systems
  • Scale fault-tolerant, high-throughput systems aligned with architectural patterns
  • Own Spark data pipelines and Airflow orchestration layer and DAG execution
  • Tune Spark workloads (partitioning, memory, execution plans, shuffle optimization)
  • Troubleshoot Spark jobs and Airflow DAGs across performance and failures
  • Operate and optimize Kubernetes-based execution environments, including node group scaling, workload placement, and resource utilization
  • Troubleshoot Kubernetes infrastructure and workload issues, including scheduling, networking, and runtime performance
  • Leverage developer productivity tools (e.g., GitHub Copilot, LLMs) to accelerate development, debugging, and operational workflows.
  • Drive operational excellence including monitoring, incident response, and RCA
  • Implement observability (metrics, logging, tracing, dashboards, alerting)
  • Define and manage SLIs/SLOs for platform reliability
  • Deploy solutions using AWS services (EKS, EC2, S3, Lambda, RDS, etc.) (Implement secure networking (VPCs, IAM, subnets, load balancing)
  • Maintain CI/CD pipelines and deployment automation
  • Lead execution across planning, delivery, and cross-team coordination
  • Mentor engineers and promote reliability and scalability best practices
  • Strong understanding of distributed systems (fault tolerance, scalability, consistency
  • Expertise in Apache Spark (tuning, debugging, optimization)
  • Expertise in Apache Airflow (DAG execution, orchestration, troubleshooting)

Requirements

  • Strong experience operating Kubernetes (EKS preferred) including cluster scaling and lifecycle management
  • Hands-on management of node groups, autoscaling, and capacity planning
  • Deep understanding of Kubernetes networking and security (security groups, network policies, ingress/egress)
  • Experience with Kubernetes resources (Deployments, StatefulSets, Jobs, CronJobs)
  • Familiarity with Custom Resources (CRDs) and advanced configuration via annotations and labels
  • Experience monitoring Kubernetes clusters (metrics, logs, events) and integrating with observability tools
  • Troubleshooting Kubernetes workloads (scheduling failures, resource contention, networking issues)
  • Experience with AWS services and cloud-native design patterns
  • Proficiency in Python, Java, or Go
  • Experience with Docker and Kubernetes
  • Hands-on observability (metrics, logging, tracing)
  • Experience with SLI/SLO-based reliability models
  • Practical experience using AI-assisted development tools (e.g., GitHub Copilot, LLMs) to improve code quality, debugging, and productivity
  • Networking fundamentals (DNS, TCP/IP, TLS, VPC design)
  • Strong troubleshooting and performance tuning skills
  • Strong communication and leadership skills
  • Bachelor's or Master's degree in Computer Science or related field (or equivalent experience)
  • 8 plus years in software, platform, or cloud engineering roles
  • Experience operating large-scale distributed systems in production
  • Strong experience with AWS cloud platforms
  • Mandatory hands-on experience with Apache Spark and Apache Airflow in production
  • Experience supporting ETL, data platforms, or workflow execution systems at scale

Apply for this position