Senior Technical Architect for AI Model Training
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
Apply senior platform engineering and production operations expertise to design realistic cloud infrastructure challenges that train and evaluate next-generation AI systems. You will create Reinforcement Learning environments that test an AI model’’s ability to design, deploy, secure, scale, troubleshoot, and recover production-grade cloud systems. No prior AI experience is required, the role prioritizes hands-on production ownership and domain expertise. Key Responsibilities
- Create realistic cloud infrastructure tasks that cover distributed systems, networking, security, scalability, and reliability.
- Design and build reproducible, containerized environments, including valid golden reference solutions and intentionally defective variants for evaluation.
- Define measurable requirements across infrastructure configuration, deployed topology, and runtime behavior.
- Develop deterministic integration, load, security, failure-injection, deployment, and recovery tests to validate model behavior.
- Build Reinforcement Learning environments that exercise IAM, queues, durable storage, observability, rolling deployments, and disaster recovery scenarios.
-
Debug environments, document technical decisions, and review and improve tasks created by other experts., + $105,000-245,000 per year Description Are you eager to architect and deliver robust, scalable, and secure software systems that enable real operational decisions for warfighters? Are you excited to dire…
-
2 days ago, + $105,000-245,000 per year Description Are you eager to architect and deliver robust, scalable, and secure systems that enable real operational decisions for warfighters? Are you excited to directly impa…
- 2 days ago +
Requirements
- Senior-level experience in technical architecture, cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE, including personal ownership of a production platform.
- Strong knowledge of distributed systems, scalable APIs, queues, autoscaling, durable storage, and partial-failure scenarios.
- Practical experience with IAM, private networking, least-privilege access, and service-to-service security.
- Experience with observability, measurable SLOs, rolling deployments, rollback strategies, and disaster recovery.
- Ability to write infrastructure automation or testing tools and to debug containerized environments using a relevant programming language.
Preferred Qualifications
- Experience with Terraform or OpenTofu.
- Experience with AWS, Azure, GCP, Kubernetes, or multi-cloud infrastructure.
- Experience building internal developer platforms, edge infrastructure, or shared platform services.
- Familiarity with chaos engineering, fault injection, local cloud emulators, or resilience testing.
- Experience creating technical evaluations, automated grading systems, or AI training/evaluation environments is helpful but not required.
Benefits & conditions
- Output-based compensation, experts are paid per task that meets project specifications.
- Time required to complete work varies by expert experience and workflow, minimum submission requirements apply.
- Experts must submit a minimum number of tasks per week.
- Assignments and volume depend on project availability and may vary over time.
Compensation
- Pay range: 60 to 130 hourly.
- Compensation is paid per completed task that meets the project specifications, rather than a fixed hourly guarantee.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
How to Become an AI Engineer
MLOps And AI Driven Development