ML Ops & Data Platform Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+34 more
Job description
internal data science, engineering, and broader Science Office teams, this person will drive key pieces of the AWS architecture, automation, and operational reliability of both the ML pipelines and the underlying data platform, sharing accountability for these systems with the rest of the team. Deliverables include working infrastructure-as-code, CI/CD pipelines, data ingestion pipelines, catalog/governance set up, observability, and documentation/knowledge transfer to internal teams as engagements close.
Job Duties Includes, but is not limited to, the following:
ML Operations Design, build, and operate AWS-native ML infrastructure, as part of a team, supporting batch and real-time inference workloads, including: Amazon SageMaker (training jobs, endpoints, pipelines, model registry) ECS/EKS + Docker for containerized services
AWS Lambda and Step Functions for orchestration and event-driven workflows S3 as the primary data/artifact layer, with lifecycle and access policies CloudWatch (and related tooling) for logging, metrics, and alerting IAM & VPC design following least-privilege and network security best practices Infrastructure as Code via AWS CDK or CloudFormation Implement CI/CD pipelines for ML (data, model, and code), including automated testing, packaging, and promotion of models across dev/staging/production environments. Help establish model and data versioning, experiment tracking, and lineage for ML pipelines to support reproducibility and auditability (builds on the shared platform’’s catalog and lineage rather than duplicating it). Build monitoring, logging, and alerting for model performance, model-input data drift, and system health; help define SLOs/SLAs for critical ML services and build the automation needed to meet them. Collaborate with data scientists and software/platform engineering teams to translate experimental workflows into production-grade services.
Data Platform Engineering Contribute to the architecture and hands-on build of an S3-based data lake / lakehouse for the Science Office, including zone design (raw/curated/consumption), open table formats (e.g., Apache Iceberg), partitioning, and lifecycle policies. Build batch and streaming ingestion/ETL/ELT pipelines using AWS Glue (jobs, crawlers) Implement data cataloging, governance, and access control, in collaboration with governance/security stakeholders, using the Glue Data Catalog and Lake Formation (finegrained, cross-team permissions), and optionally DataZone for data product publishing/discovery across science teams Enable self-service analytics and data access for scientists across the Science Office via Athena including query patterns, cost controls, and onboarding documentation. Help establish data quality, schema/metadata management, and lineage frameworks (e.g., Glue Data Quality, schema registries) for shared datasets - distinct from ML-side model-input drift monitoring, this workstream covers dataset-level quality for assets consumed across multiple teams.
Shared Duties Identify and implement cost optimizations across both ML and data platform workloads (right-sizing, spot/scheduling strategies, storage tiering, Athena/Redshift query cost controls) without compromising reliability. Collaborate with data scientists, analysts, and scientists across the broader Science Office to translate experimental workflows and analytical needs into production-grade services. Document architecture, runbooks, and operational procedures for both the ML infrastructure and the data platform, and provide knowledge transfer to internal teams as part of engagement close-out.
Requirements
Required Experience 5+ years of hands-on experience in MLOps, data engineering, DevOps, or cloud infrastructure engineering roles, with a strong recent focus on AWS.
Demonstrated production experience with a substantial subset of ML infrastructure services: SageMaker, ECS/EKS, Lambda, Step Functions, S3, CloudWatch, IAM, VPC, and CDK or CloudFormation for infrastructure as code. Demonstrated production experience with a substantial subset of AWS data platform services: Glue, Lake Formation, Athena.
Experience designing and operating data lake/lakehouse architectures on S3, including ETL/ELT pipeline development and open table formats (e.g., Iceberg). Experience with data governance, cataloging, and multi-team access control (Lake Formation, Glue Data Catalog, or equivalent) for shared, regulated data assets.
Strong Python and SQL skills, with working familiarity with at least one ML framework (e.g., TensorFlow, PyTorch, scikit-learn) sufficient to integrate with and support data science workflows. Experience with containerization and orchestration (Docker, Kubernetes) in production settings. Experience building and operating CI/CD pipelines for ML and data systems (data, model, and code promotion).
Must be authorized to work in the country where work will be performed.
Nice to Have AWS certifications (Solutions Architect - Professional, Machine Learning - Specialty, Data Analytics/Data Engineer) or equivalent demonstrated expertise. Experience with multi-account AWS architectures and GPU workload management on AWS. Experience with MLflow, SageMaker Model Registry, or similar model governance tooling. Experience with DataZone or data-mesh/data-product patterns. Experience with dbt, Airflow/MWAA, or similar transformation/orchestration tooling.
Experience standing up self-service data platforms consumed by non-engineering scientific/analytical users. Terraform experience as an alternative/complement to CDK. Prior experience delivering as a contractor/consultant with clean documentation and handoff practices. Prior experience in healthcare, life sciences, or other regulated environments.
Success Criteria / Definition of Done Success is measured by this contractor’’s contributions to the following team outcomes: ML Operations Contractor-delivered ML pipeline components running in production with automated CI/CD. Monitoring, logging, and alerting live and validated against SLOs defined with the team. Data Platforming Data platform foundation live on AWS: S3 lakehouse zones defined, Glue Data Catalog populated, and Lake Formation permissions implemented for at least the initial Science Office teams/use cases. Ingestion pipelines operational for the agreed initial data sources, with data quality checks in place. Self-service query access (Athena/Redshift) validated by at least one non-ML science team, with onboarding documentation. Shared All infrastructure delivered as part of this engagement codified (CDK/CloudFormation) and checked into version control. Documentation and knowledge transfer completed with the internal team, covering both workstreams
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps – What’s the deal behind it?
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
MLOps And AI Driven Development