Site Reliability Engineer - SRE Fleet
Role details
Job location
Tech stack
Job description
The SRE Fleet team is responsible for maintaining the stability, scalability, and efficiency of the infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US, Canada, and the UK, we combine deep infrastructure expertise with a strong focus on automation, reliability, and operational excellence. We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide.
The team operates with a high degree of autonomy, giving engineers the opportunity to drive both critical initiatives and grassroots improvements that solve real operational challenges. One of our most exciting areas of focus is expanding our ability to build highly automated regional, sovereign, and isolated cloud environments that support new markets and evolving regulatory requirements.
Everyone on the team has a voice, and engineers are encouraged to identify problems, propose solutions, and take ownership of improvements that make the platform more reliable and easier to operate.
Your Impact
Develop and maintain automation solutions that improve the reliability, scalability, and operational efficiency of infrastructure spanning more than 2,000 machines across global cloud environments. Design and enhance deployment pipelines, testing frameworks, and operational tooling to support the continued growth of a platform serving millions of managed devices worldwide. Troubleshoot complex infrastructure and distributed systems issues to ensure high availability while helping teams identify and address performance and scalability challenges.
Contribute to critical projects such as cluster build out by building automation that enables the rapid and repeatable deployment of new sovereign, regional, and purpose-built cloud environments. Partner with other engineering teams, product management, and business partners across multiple teams and time zones to understand platform dependencies, seek opportunities for improvement, and deliver solutions that enhance reliability and reduce operational overhead.
Requirements
- 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based production environments.
- Experience developing and maintaining infrastructure automation using Ansible.
- Experience programming in Ruby and developing automated tests using RSpec or comparable testing frameworks.
- Experience administering and troubleshooting Linux-based systems and distributed infrastructure environments.
- Experience designing, implementing, and maintaining CI/CD pipelines, including GitLab CI.
- Experience supporting large-scale infrastructure environments consisting of hundreds or thousands of systems., * Familiarity with AWS or other public cloud platforms and hybrid infrastructure environments.
- Knowledge of monitoring, observability, and reliability engineering practices and tooling.
- Familiarity with Kubernetes concepts and containerized application platforms.
- Experience leveraging AI-assisted development tools to improve software development, automation, operational analysis, and engineering productivity.