Senior Cloud Operations Engineer

The Linux Foundation
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$95,000.0 - $133,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon S3 Unit Testing Microsoft Azure Bash Shell Cloud Computing Continuous Integration Software Debugging Linux DevOps Github
+18 more
Python (Programming Language) Open Source Technology Performance Tuning Runbook TypeScript Datadog Scripting Google Cloud Pytorch Multi-Cloud Containerization Kubernetes Information Technology Cloudflare Free and Open-Source Software Cloudwatch Terraform Docker

Job description

The Linux Foundation is a driving force in fostering open source collaboration and supporting communities across a range of projects, including PyTorch. We are dedicated to enhancing and expanding our infrastructure to meet the growing demands of PyTorch and related AI projects. We are seeking a Senior Cloud Operations Engineer who will focus on the infrastructure operations of the PyTorch project, automating processes, optimizing cloud-native tools, and ensuring a robust and scalable cloud environment., * Contribute to architectural exercises with open source community and technical leads to validate new cloud infrastructure

  • Implement and maintain infrastructure-as-code using Terraform via pytorch/ci-infra and pytorch/test-infra
  • Optimize cloud resource utilization and implement FinOps practices for cost management and reporting

CI/CD and DevOps

  • Design, implement, and maintain CI/CD pipelines using GitHub Actions and ARC, including runner configurations and other elements of the CI ecosystem
  • Debug and triage issues in build and test pipelines, including experience with unit testing
  • Develop monitoring and alerting solutions for CI/CD workflows and critical infrastructure

Performance Optimization and Security

  • Manage and optimize Cloudflare CDN deployments for PyTorch assets (R2/S3)
  • Implement best practices for CDN and overall infrastructure security

Monitoring and Incident Response

  • Develop comprehensive monitoring and observability solutions using Datadog, AWS CloudWatch, and other telemetry data collection and processing tools
  • Review and recommend monitoring solutions as project and community needs evolve
  • Participate in on-call rotations supporting operations and incident response using incident.io
  • Establish and maintain escalation procedures and resolution processes

Community Collaboration and Project Management

  • Participate in ci-infra and multi-cloud working groups and support architecture decisions
  • Collaborate with external contributors and promote DevOps best practices
  • Manage GitHub repositories, including user onboarding and access control
  • Attend and contribute to technical meetings, including Infrastructure, CI Workflow, and Technical Advisory Council sessions

Documentation and Best Practices

  • Develop and maintain technical documentation for infrastructure and processes
  • Provide guidance on developer best practices and tooling
  • Create and update runbooks for common operational tasks and incident response

Requirements

Do you have experience in TypeScript?, * Ability to work with communities made up of industry specialists and collaborate outside of the Linux Foundation

  • Bachelor’s degree in Computer Science, Engineering, or related field
  • 7+ years of experience in cloud operations with significant AWS expertise
  • Strong knowledge of infrastructure-as-code principles and tools, particularly Terraform
  • Proficiency in scripting languages (Python, TypeScript, Bash) and containerization technologies (Docker, Kubernetes)
  • Experience with Cloudflare CDN management and optimization
  • Expertise in implementing and managing monitoring solutions, specifically Datadog and AWS CloudWatch
  • Familiarity with incident management tools and processes, particularly incident.io
  • Demonstrated experience in CI/CD pipeline design and implementation
  • Strong problem-solving skills and ability to troubleshoot complex systems
  • Excellent communication skills and experience collaborating with open source communities, * Experience with PyTorch or other open source communities
  • Multi-cloud expertise across AWS, GCP, and Azure
  • GitHub ARC experience
  • Knowledge of FinOps principles and cloud cost optimization strategies
  • Contributions to open source projects, especially in infrastructure management roles
  • Familiarity with the Linux Foundation or similar open source foundations
  • Experience mentoring other engineers and fostering a collaborative team environment, The Linux Foundation is unable to provide visa sponsorship for this position. Candidates must be authorized to work in their country of residence without employer sponsorship, now or in the future.

Benefits & conditions

3.33.3 out of 5 stars Remote $95,000 - $133,000 a year - Full-time

About the company

The Senior Cloud Operations Engineer will play a pivotal role in the PyTorch Foundation, leading cloud infrastructure and DevOps initiatives. This position is crucial for maintaining and optimizing the technical operations that support PyTorch, one of the world’s leading open source machine learning frameworks. The ideal candidate will blend expertise in cloud technologies, DevOps practices, and open source collaboration to ensure PyTorch’s infrastructure remains robust, secure, and efficient., The Linux Foundation maintains a predominantly remote workforce and is committed to hiring top-notch talent. We are as passionate about providing a flexible and supportive work culture as we are about open source software. Collaboration is embedded in our DNA, and we take pride in our ability to work closely together while not being confined to a traditional office space.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all