Software Development Engineer II, AWS SageMaker AI

Amazon.com, Inc.
Bellevue, WA, United States
7 days ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$143,700.0 - $194,400.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Software Applications C Sharp (Programming Language) C++ (Programming Language) Cloud Computing Code Review Computer Programming Data Cleansing Software Design Patterns
+21 more
Distributed Systems Perl (Programming Language) Fault Tolerance Machine Learning Object-Oriented Software Development Service-Oriented Architecture Software Engineering AWS Cdk Data Logging Multithreading Graphics Processing Unit (GPU) Pytorch Backend Cloudformation Kubernetes Information Technology Build Process Slurm Software Coding Software Version Control Programming Languages

Job description

We’re looking for a Software Development Engineer to help design, build, and operate the distributed systems at the core of this platform - the orchestration engine, compute integrations, SDK and contract layer, and the infrastructure that runs large-scale training and customization jobs. You’ll own components end-to-end, from design through delivery and on-call operations, and work closely with the ML scientists and platform teams who depend on Model Factory every day. You’ll turn requirements into robust, scalable, supportable services that fit cleanly into the overall architecture, uphold a high engineering bar, and grow your scope and technical leadership as you go., As a Software Development Engineer on the SageMaker AI team, you will:

  • Design, build, test, and operate services that orchestrate foundation-model data preparation, training, evaluation, and deployment as reliable, contract-validated workflows.
  • Own delivery of individual components end-to-end - from design and implementation through deployment, monitoring, and on-call operations.
  • Build and extend compute-backend integrations and job launchers - submitting, monitoring, and recovering large-scale training jobs across SageMaker (Training/Processing/HyperPod), EMR, AWS Batch, and Kubernetes/EKS.
  • Improve the platform’s resiliency and operability for long-running distributed jobs - checkpoint/resume, fault detection and recovery, retries, and observability (metrics, logging, experiment tracking).
  • Contribute to the SDK, workflow orchestration, and schema/contract layer that teams use to declare and run jobs, and to the CDK infrastructure that deploys the platform.
  • Integrate containerized training and evaluation frameworks (e.g., PyTorch/FSDP, verl, NeMo/Megatron) into the platform’s task and recipe model.
  • Contribute to design and architecture discussions, write clear technical designs, and uphold engineering best practices (code review, testing, operational readiness).
  • Collaborate with ML scientists and internal customers to translate training requirements into reliable, self-service platform capabilities, and help onboard and mentor interns and new engineers as you grow.

Requirements

A successful candidate has a strong software-engineering foundation, writes high-quality distributed-systems and services code, communicates clearly, and is motivated to deliver results in a fast-paced, ambiguous environment., 3+ years of non-internship professional software development experience

  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 1+ years of software development engineer or related occupational experience
  • 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience
  • 1+ years of Object Oriented Design experience
  • Bachelor’s degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
  • Experience programming with at least one software programming language, 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor’s degree in computer science or equivalent
    • Experience with workflow/pipeline orchestration (Airflow, Step Functions, or similar) and event-driven or service-oriented architectures.
    • Experience with container and cluster compute (Kubernetes/EKS, Ray, Slurm, AWS Batch) and cloud infrastructure-as-code (AWS CDK/CloudFormation).
    • Experience building and operating fully-managed cloud services at scale, including resiliency, checkpointing, and fault tolerance for long-running jobs.
    • Familiarity with machine-learning / deep-learning training workflows, GPU/accelerator compute (SageMaker HyperPod, AWS Trainium, P5-class GPUs), or distributed-training frameworks (PyTorch FSDP, Megatron-LM, DeepSpeed, verl).

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, WA, BELLEVUE - 143,700.00 - 194,400.00 USD annually

About the company

At AWS SageMaker AI, we’re making it easy to build state-of-the-art foundation models on the cloud. Model Factory is our platform for building, training, customizing, and evaluating foundation models at scale. Instead of hand-chaining data prep, distributed training, evaluation, and deployment across thousands of GPU and AWS Trainium devices, Model Factory lets teams express the whole lifecycle as a single, contract-validated workflow - orchestrated, reproducible, and fully managed. As LLMs and Generative AI scale, Model Factory is the platform that turns frontier training research into a reliable, repeatable pipeline for our internal teams and customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.amazon.jobs

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · WWC 2024

Videos

See all

Related articles

See all