Site Reliability / Gitops Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Job description
Join a global Information Systems team responsible for operating and evolving critical production services at significant scale. In this role, youâll use automation, Infrastructure as Code, and GitOps practices to make cloud operations more reliable, consistent, and efficient. Youâll work across private and public cloud environments, strengthening infrastructure resilience, scalability, observability, and performance. Your expertise will also influence the evolution of open-source infrastructure technologies through hands-on feedback, bug reporting, and collaboration. Youâll troubleshoot complex systems, improve operational processes, and help eliminate repetitive manual work through thoughtful automation. Working with a distributed team of experienced SREs, youâll have opportunities to share knowledge, mentor colleagues, and contribute to major engineering initiatives. This is an ideal opportunity for an automation-first technologist who is passionate about Linux, open source, and building robust systems at scale. Accountabilities:
- Apply Infrastructure as Code expertise to continuously improve automation practices, processes, and operational consistency.
- Automate software operations across private and public clouds while accounting for the complexities of distributed systems.
- Develop new capabilities and improve the resilience, scalability, and reliability of cloud and container infrastructure.
- Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliable day-to-day performance.
- Troubleshoot complex systems, perform capacity planning and performance investigations, and develop strong operational expertise.
- Set up, maintain, and use observability and monitoring solutions such as Prometheus, Grafana, and Elasticsearch.
- Design and maintain monitoring and alerting for critical systems and services.
- Collaborate with development teams on service architecture, documentation, playbooks, policies, and operational procedures.
- Work closely with globally distributed engineering, operations, and support teams to resolve issues and improve services.
- Dedicate focused development time to larger engineering projects and the automation of repetitive manual processes.
- Share technical knowledge and best practices through design sessions, mentoring, and collaborative problem-solving.
- Take final responsibility for resolving time-critical operational escalations., * Exposure to private and public cloud environments, Infrastructure as Code, GitOps, CI/CD, observability, and open-source technologies.
- Dedicated development time for impactful automation and larger engineering projects.
- Collaboration with a highly experienced, globally distributed SRE and engineering community.
- Opportunities for mentoring, knowledge sharing, and cross-functional technical collaboration.
- Remote work flexibility, with the role available across time zones.
- Opportunities to meet colleagues in person 2-4 times per year at internal events, typically lasting 1-2 weeks.
- International exposure through collaboration with distributed teams and participation in global company events.
- Compensation and benefits are determined according to the role, location, experience, and applicable company policies.
Requirements
- Deep experience defining IT operations through code, using version control, peer review, and CI/CD to deploy application and infrastructure changes.
- Strong modern software engineering practices, including peer review, unit testing, source control management, CI/CD, and Agile methodologies.
- Significant Python development experience, including work on large or complex projects.
- Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
- Familiarity with Linux storage technologies, ranging from Ceph to database systems.
- Hands-on experience administering enterprise Linux servers.
- Extensive understanding of cloud computing concepts, architectures, and technologies.
- Bachelorâs degree or higher, preferably in computer science, software engineering, or a related technical discipline.
- Strong English communication skills across written and spoken channels, including email, chat, video calls, voice calls, and in-person collaboration.
- Strong troubleshooting abilities, with the curiosity and persistence to investigate issues from the Linux kernel through to the web layer.
- Ability to collaborate effectively while knowing when to seek input from teammates and subject-matter experts.
- Adaptability, willingness to learn quickly, and comfort working in fast-changing technical environments.
- Ability to thrive within globally distributed teams and collaborate across different locations and time zones.
- Strong interest in open-source technologies, particularly Ubuntu or Debian.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where To Find Software Engineering Jobs
Fully Remote Software Engineer Jobs
Where to Find Entry-Level Software Engineering Jobs
Finding Jobs in Germany