Senior Platform Engineer for AI Model Training

Johns Hopkins Applied Physics Laboratory
Washington, DC, United States
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$124,800.0 - $270,400.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Systems Engineering Automation of Tests Microsoft Azure Cloud Computing Software Debugging DevOps Disaster Recovery Distributed Systems Identity and Access Management Machine Learning
+10 more
Reliability Engineering Reinforcement Learning Google Cloud Autoscaling Multi-Cloud Kubernetes Infrastructure Automation Frameworks Restful APIs Terraform Programming Languages

Job description

Apply deep cloud infrastructure and platform engineering expertise to design realistic, production-grade scenarios that train and evaluate next-generation AI systems. You will author Reinforcement Learning environments that test models on designing, deploying, troubleshooting, securing, scaling, and recovering distributed cloud systems. No prior AI domain experience is required, the role prioritizes hands-on, production ownership and engineering judgment. Key Responsibilities

  • Create realistic cloud infrastructure tasks covering distributed systems, networking, security, scalability, and reliability.
  • Build reproducible, containerized environments with valid golden reference solutions and intentionally defective variants for failure and resilience testing.
  • Define measurable requirements across infrastructure configuration, deployed topology, and runtime behavior.
  • Develop deterministic integration, load, security, failure-injection, deployment, and recovery tests.
  • Debug environments, document technical decisions, and review tasks produced by other experts to ensure quality and correctness., * Work is task based, with experts responsible for completing reproducible tasks that meet project specifications.

Requirements

  • Senior-level experience in cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE, including personal ownership of a production platform.
  • Core skills: Cloud Benchmark Task Authoring, Cloud and Distributed Systems Architecture, Production Infrastructure Ownership.
  • Strong understanding of distributed systems, scalable APIs, queues, autoscaling, durable storage, and partial-failure scenarios.
  • Practical experience with IAM, private networking, least-privilege access, and service-to-service security.
  • Experience with observability, measurable SLOs, rolling deployments, rollback strategies, and disaster recovery.
  • Ability to write infrastructure automation or testing tools and to debug containerized environments using a relevant programming language.

Preferred

  • Experience with Terraform or OpenTofu.
  • Experience on AWS, Azure, GCP, Kubernetes, or multi-cloud deployments.
  • Experience building internal developer platforms, edge infrastructure, or shared platform services.
  • Experience with chaos engineering, fault injection, local cloud emulators, or resilience testing.
  • Experience creating technical evaluations, automated grading systems, or AI training/evaluation environments is helpful but not required.

Benefits & conditions

  • Listed pay range: $60 to $130 per hour.
  • Compensation model is output based, experts are paid per task that meets the project specifications.
  • Minimum submission requirements apply, and payment depends on meeting the task acceptance criteria.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all