Senior Platform Engineer for AI Model Training
Johns Hopkins Applied Physics Laboratory
Washington, DC, United States
8 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$124,800.0 - $270,400.0
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Amazon Web Services
Systems Engineering
Automation of Tests
Microsoft Azure
Cloud Computing
Software Debugging
DevOps
Disaster Recovery
Distributed Systems
Identity and Access Management
Machine Learning
+10 more
Reliability Engineering
Reinforcement Learning
Google Cloud
Autoscaling
Multi-Cloud
Kubernetes
Infrastructure Automation Frameworks
Restful APIs
Terraform
Programming Languages
Job description
Apply deep cloud infrastructure and platform engineering expertise to design realistic, production-grade scenarios that train and evaluate next-generation AI systems. You will author Reinforcement Learning environments that test models on designing, deploying, troubleshooting, securing, scaling, and recovering distributed cloud systems. No prior AI domain experience is required, the role prioritizes hands-on, production ownership and engineering judgment. Key Responsibilities
- Create realistic cloud infrastructure tasks covering distributed systems, networking, security, scalability, and reliability.
- Build reproducible, containerized environments with valid golden reference solutions and intentionally defective variants for failure and resilience testing.
- Define measurable requirements across infrastructure configuration, deployed topology, and runtime behavior.
- Develop deterministic integration, load, security, failure-injection, deployment, and recovery tests.
- Debug environments, document technical decisions, and review tasks produced by other experts to ensure quality and correctness., * Work is task based, with experts responsible for completing reproducible tasks that meet project specifications.
Requirements
- Senior-level experience in cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE, including personal ownership of a production platform.
- Core skills: Cloud Benchmark Task Authoring, Cloud and Distributed Systems Architecture, Production Infrastructure Ownership.
- Strong understanding of distributed systems, scalable APIs, queues, autoscaling, durable storage, and partial-failure scenarios.
- Practical experience with IAM, private networking, least-privilege access, and service-to-service security.
- Experience with observability, measurable SLOs, rolling deployments, rollback strategies, and disaster recovery.
- Ability to write infrastructure automation or testing tools and to debug containerized environments using a relevant programming language.
Preferred
- Experience with Terraform or OpenTofu.
- Experience on AWS, Azure, GCP, Kubernetes, or multi-cloud deployments.
- Experience building internal developer platforms, edge infrastructure, or shared platform services.
- Experience with chaos engineering, fault injection, local cloud emulators, or resilience testing.
- Experience creating technical evaluations, automated grading systems, or AI training/evaluation environments is helpful but not required.
Benefits & conditions
- Listed pay range: $60 to $130 per hour.
- Compensation model is output based, experts are paid per task that meets the project specifications.
- Minimum submission requirements apply, and payment depends on meeting the task acceptance criteria.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
ER
Erin Rifkin
about 1 year ago
LM
Luis Minvielle
How to Become an AI Engineer
over 2 years ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
almost 2 years ago
EF
Elizabeth Fuentes Leone, AWS Developer Advocate, GenAI
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
9 months ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago