Sr. DevOps AWS - United States
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
- Design, build, and maintain scalable and reliable cloud infrastructure primarily on AWS.
- Develop and manage production-grade containerized workloads using Docker and OCI standards.
- Design and maintain secure container environments and implement container security best practices.
- Build and manage infrastructure using Infrastructure as Code (IaC) practices.
- Develop and maintain CI/CD pipelines to automate application, infrastructure, and model deployments.
- Design and manage cloud-native workloads using services such as Amazon EKS, ECS, SageMaker, and AWS Batch.
- Build and manage infrastructure and workflows supporting AI/ML inference workloads.
- Develop reusable platform capabilities that can support multiple applications, services, and workloads.
- Support GPU-based workloads, including GPU/CUDA compatibility, resource allocation, and scheduling.
- Implement robust observability solutions across infrastructure and applications, including monitoring, logging, alerting, and performance metrics.
- Troubleshoot complex infrastructure, deployment, networking, and performance issues in production environments.
- Collaborate with engineering and ML teams to improve reliability, scalability, deployment processes, and operational efficiency.
- Support the lifecycle of ML artifacts and production model deployments.
- Implement security controls for infrastructure and workloads, particularly when handling sensitive or regulated data.
- Conduct performance testing and optimization of infrastructure and distributed workloads.
- Identify opportunities to improve automation, reliability, scalability, and developer experience.
- Communicate technical decisions, risks, incidents, and recommendations clearly with both internal teams and customers.
- Operate effectively in ambiguous and fast-paced environments, taking ownership of problems from identification through resolution.
- Contribute to technical standards, documentation, reusable tooling, and best practices across the platform.
- Stay current with emerging cloud, DevOps, MLOps, and AI infrastructure technologies.
Requirements
We are seeking a highly hands-on and self-driven Senior DevOps Engineer with strong AWS experience to join our team and help build and scale modern cloud infrastructure and AI/ML platforms.
In this role, you will take ownership of infrastructure, deployment, observability, and platform capabilities supporting production workloads in a fast-paced startup environment. You will work closely with engineering and technical teams to design reusable, secure, and scalable solutions rather than one-off scripts or manual processes.
The ideal candidate has strong experience with AWS, containers, Infrastructure as Code, CI/CD, observability, and distributed systems, along with the ability to work effectively in ambiguous environments. Experience supporting AI/ML workloads, inference containers, GPU-based workloads, or MLOps platforms is highly valuable.
This role requires someone who is comfortable taking ownership, communicating proactively with both technical and non-technical stakeholders, learning quickly, and operating as part of a single, collaborative team., * Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience.
- 5+ years of experience in DevOps, Cloud Infrastructure, MLOps, SRE, Platform Engineering, or related disciplines.
- Strong hands-on experience with AWS cloud services and cloud-native architecture.
- Experience working in fast-paced startup environments.
- Hands-on experience with Docker and production-grade containerization using OCI standards.
- Experience with Kubernetes and/or container orchestration platforms such as Amazon EKS or ECS.
- Experience with AWS services such as EKS, ECS, SageMaker, and/or AWS Batch.
- Strong experience implementing Infrastructure as Code using tools such as Terraform, CloudFormation, or similar technologies.
- Strong experience designing and maintaining CI/CD pipelines.
- Experience building and operating production infrastructure with a strong focus on reliability and scalability.
- Experience implementing observability solutions, including monitoring, logging, metrics, tracing, and alerting.
- Strong Python experience and the ability to develop automation and infrastructure tooling.
- Experience working with distributed systems and production workloads.
- Experience with cloud security and secure infrastructure practices.
- Experience troubleshooting complex production issues and working effectively under pressure.
- Strong communication skills and the ability to communicate proactively with both technical and non-technical stakeholders.
- Highly self-driven, collaborative, and comfortable taking ownership in environments with ambiguity and rapidly changing priorities.
- Ability to learn new technologies and concepts quickly., * Experience with MLOps, AI platforms, or ML infrastructure.
- Experience building and managing inference container workflows.
- Experience supporting GPU workloads and CUDA compatibility and scheduling.
- Familiarity with ML artifact management and model lifecycle processes.
- Experience with AI/ML evaluation infrastructure, evaluation harnesses, or similar platforms.
- Experience with high-performance inference runtimes such as vLLM, Triton, TensorRT, ONNX, or similar technologies.
- Experience with performance testing and optimization of AI/ML workloads.
- Experience working with sensitive, regulated, or compliance-driven data environments.
- Experience in healthcare, med-tech, or HIPAA-regulated environments.
- Experience building reusable internal platforms and developer tooling.
- Experience with distributed processing frameworks and large-scale data workloads.
- Strong familiarity with Python-based automation and cloud-native development practices.
Benefits & conditions
Pulled from the full job description
- Flexible schedule, Join a global team committed to Distillery’s core values: Unyielding Commitment, Relentless Pursuit, Courageous Ambition, and Authentic Connection.
- 100% Remote Work: Enjoy the freedom to work from anywhere while collaborating with a diverse, multinational team.
- Competitive Compensation: Generous and competitive package in USD, along with a comprehensive benefits plan.
- Flexible Hours: Create a schedule that aligns with your life and priorities.
- Home Office Setup: Receive all the hardware and software needed to succeed from home.
- Innovative Workplace: Collaborate with the global Top 1% of talent in a multicultural and dynamic environment.
- Focus on Growth: Pursue professional and personal development while contributing your unique talents to a team where you can truly shine!
About the company
Distillery is a software development partner that helps organizations design, build, and scale modern technology solutions. Leading companies work with us to accelerate digital initiatives, solve complex technical challenges, and bring innovative products to market faster. Working as an extension of our clients’ teams, we deliver lasting business value.
Distillery is committed to diversity and inclusion. We embrace a dynamic and inclusive culture where you have the opportunity to continuously learn and grow.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 121 - AI goes offline
Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters
DevOps Engineer Salary [2023]
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again