> Markdown version of [/jobs/ext/1919779-senior-platform-engineer-for-ai-model-training](https://www.wearedevelopers.com/jobs/ext/1919779-senior-platform-engineer-for-ai-model-training). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform Engineer for AI Model Training - **Company:** Johns Hopkins Applied Physics Laboratory - **Location:** Washington, DC, United States - **Experience:** Expert - **Salary:** $124,800.0 - $270,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Automation of Tests, Microsoft Azure, Cloud Computing, Software Debugging, DevOps, Disaster Recovery, Distributed Systems, Identity and Access Management, Machine Learning, Reliability Engineering, Reinforcement Learning, Google Cloud, Autoscaling, Multi-Cloud, Kubernetes, Infrastructure Automation Frameworks, Restful APIs, Terraform, Programming Languages - **Published:** August 4, 2026 - **Apply:** https://www.careerjet.com/job/register/us6f85b084868f83662f482cbd160e9178 ## About the Role * Senior-level experience in cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE, including personal ownership of a production platform. * Core skills: Cloud Benchmark Task Authoring, Cloud and Distributed Systems Architecture, Production Infrastructure Ownership. * Strong understanding of distributed systems, scalable APIs, queues, autoscaling, durable storage, and partial-failure scenarios. * Practical experience with IAM, private networking, least-privilege access, and service-to-service security. * Experience with observability, measurable SLOs, rolling deployments, rollback strategies, and disaster recovery. * Ability to write infrastructure automation or testing tools and to debug containerized environments using a relevant programming language. Preferred * Experience with Terraform or OpenTofu. * Experience on AWS, Azure, GCP, Kubernetes, or multi-cloud deployments. * Experience building internal developer platforms, edge infrastructure, or shared platform services. * Experience with chaos engineering, fault injection, local cloud emulators, or resilience testing. * Experience creating technical evaluations, automated grading systems, or AI training/evaluation environments is helpful but not required. ## Description Apply deep cloud infrastructure and platform engineering expertise to design realistic, production-grade scenarios that train and evaluate next-generation AI systems. You will author Reinforcement Learning environments that test models on designing, deploying, troubleshooting, securing, scaling, and recovering distributed cloud systems. No prior AI domain experience is required, the role prioritizes hands-on, production ownership and engineering judgment. Key Responsibilities * Create realistic cloud infrastructure tasks covering distributed systems, networking, security, scalability, and reliability. * Build reproducible, containerized environments with valid golden reference solutions and intentionally defective variants for failure and resilience testing. * Define measurable requirements across infrastructure configuration, deployed topology, and runtime behavior. * Develop deterministic integration, load, security, failure-injection, deployment, and recovery tests. * Debug environments, document technical decisions, and review tasks produced by other experts to ensure quality and correctness., * Work is task based, with experts responsible for completing reproducible tasks that meet project specifications. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)