Sr. DevOps AWS - United States

Distil Networks, Inc.
United States
26 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Shift work
Job source

Tech stack

Artificial Intelligence Amazon Web Services Big Data Cloud Computing Cloud Computing Security Nvidia CUDA Continuous Integration Data Infrastructure DevOps Programming Tools Distributed Computing Environment Distributed Systems
+18 more
Python (Programming Language) Reliability Engineering Cloud Services Azure Machine Learning AI Infrastructure Data Logging Performance Testing Infrastructure as Code (IaC) Cloudformation AI Platforms Kubernetes Information Technology ONNX (Open Neural Network Exchange) Format Machine Learning Operations TensorRT Terraform Oracle Cloud Infrastructure Docker

Job description

  • Design, build, and maintain scalable and reliable cloud infrastructure primarily on AWS.
  • Develop and manage production-grade containerized workloads using Docker and OCI standards.
  • Design and maintain secure container environments and implement container security best practices.
  • Build and manage infrastructure using Infrastructure as Code (IaC) practices.
  • Develop and maintain CI/CD pipelines to automate application, infrastructure, and model deployments.
  • Design and manage cloud-native workloads using services such as Amazon EKS, ECS, SageMaker, and AWS Batch.
  • Build and manage infrastructure and workflows supporting AI/ML inference workloads.
  • Develop reusable platform capabilities that can support multiple applications, services, and workloads.
  • Support GPU-based workloads, including GPU/CUDA compatibility, resource allocation, and scheduling.
  • Implement robust observability solutions across infrastructure and applications, including monitoring, logging, alerting, and performance metrics.
  • Troubleshoot complex infrastructure, deployment, networking, and performance issues in production environments.
  • Collaborate with engineering and ML teams to improve reliability, scalability, deployment processes, and operational efficiency.
  • Support the lifecycle of ML artifacts and production model deployments.
  • Implement security controls for infrastructure and workloads, particularly when handling sensitive or regulated data.
  • Conduct performance testing and optimization of infrastructure and distributed workloads.
  • Identify opportunities to improve automation, reliability, scalability, and developer experience.
  • Communicate technical decisions, risks, incidents, and recommendations clearly with both internal teams and customers.
  • Operate effectively in ambiguous and fast-paced environments, taking ownership of problems from identification through resolution.
  • Contribute to technical standards, documentation, reusable tooling, and best practices across the platform.
  • Stay current with emerging cloud, DevOps, MLOps, and AI infrastructure technologies.

Requirements

We are seeking a highly hands-on and self-driven Senior DevOps Engineer with strong AWS experience to join our team and help build and scale modern cloud infrastructure and AI/ML platforms.

In this role, you will take ownership of infrastructure, deployment, observability, and platform capabilities supporting production workloads in a fast-paced startup environment. You will work closely with engineering and technical teams to design reusable, secure, and scalable solutions rather than one-off scripts or manual processes.

The ideal candidate has strong experience with AWS, containers, Infrastructure as Code, CI/CD, observability, and distributed systems, along with the ability to work effectively in ambiguous environments. Experience supporting AI/ML workloads, inference containers, GPU-based workloads, or MLOps platforms is highly valuable.

This role requires someone who is comfortable taking ownership, communicating proactively with both technical and non-technical stakeholders, learning quickly, and operating as part of a single, collaborative team., * Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related technical field, or equivalent practical experience.

  • 5+ years of experience in DevOps, Cloud Infrastructure, MLOps, SRE, Platform Engineering, or related disciplines.
  • Strong hands-on experience with AWS cloud services and cloud-native architecture.
  • Experience working in fast-paced startup environments.
  • Hands-on experience with Docker and production-grade containerization using OCI standards.
  • Experience with Kubernetes and/or container orchestration platforms such as Amazon EKS or ECS.
  • Experience with AWS services such as EKS, ECS, SageMaker, and/or AWS Batch.
  • Strong experience implementing Infrastructure as Code using tools such as Terraform, CloudFormation, or similar technologies.
  • Strong experience designing and maintaining CI/CD pipelines.
  • Experience building and operating production infrastructure with a strong focus on reliability and scalability.
  • Experience implementing observability solutions, including monitoring, logging, metrics, tracing, and alerting.
  • Strong Python experience and the ability to develop automation and infrastructure tooling.
  • Experience working with distributed systems and production workloads.
  • Experience with cloud security and secure infrastructure practices.
  • Experience troubleshooting complex production issues and working effectively under pressure.
  • Strong communication skills and the ability to communicate proactively with both technical and non-technical stakeholders.
  • Highly self-driven, collaborative, and comfortable taking ownership in environments with ambiguity and rapidly changing priorities.
  • Ability to learn new technologies and concepts quickly., * Experience with MLOps, AI platforms, or ML infrastructure.
  • Experience building and managing inference container workflows.
  • Experience supporting GPU workloads and CUDA compatibility and scheduling.
  • Familiarity with ML artifact management and model lifecycle processes.
  • Experience with AI/ML evaluation infrastructure, evaluation harnesses, or similar platforms.
  • Experience with high-performance inference runtimes such as vLLM, Triton, TensorRT, ONNX, or similar technologies.
  • Experience with performance testing and optimization of AI/ML workloads.
  • Experience working with sensitive, regulated, or compliance-driven data environments.
  • Experience in healthcare, med-tech, or HIPAA-regulated environments.
  • Experience building reusable internal platforms and developer tooling.
  • Experience with distributed processing frameworks and large-scale data workloads.
  • Strong familiarity with Python-based automation and cloud-native development practices.

Benefits & conditions

Pulled from the full job description

  • Flexible schedule, Join a global team committed to Distillery’s core values: Unyielding Commitment, Relentless Pursuit, Courageous Ambition, and Authentic Connection.
  • 100% Remote Work: Enjoy the freedom to work from anywhere while collaborating with a diverse, multinational team.
  • Competitive Compensation: Generous and competitive package in USD, along with a comprehensive benefits plan.
  • Flexible Hours: Create a schedule that aligns with your life and priorities.
  • Home Office Setup: Receive all the hardware and software needed to succeed from home.
  • Innovative Workplace: Collaborate with the global Top 1% of talent in a multicultural and dynamic environment.
  • Focus on Growth: Pursue professional and personal development while contributing your unique talents to a team where you can truly shine!

About the company

Distillery is a software development partner that helps organizations design, build, and scale modern technology solutions. Leading companies work with us to accelerate digital initiatives, solve complex technical challenges, and bring innovative products to market faster. Working as an extension of our clients’ teams, we deliver lasting business value.

Distillery is committed to diversity and inclusion. We embrace a dynamic and inclusive culture where you have the opportunity to continuously learn and grow.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:35 min

Accessing software containers and developer training platforms

Paul Graham Paul Graham · World Congress 2024

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all