Site Reliability / Gitops Engineer

Jobgether
Germany
1 day ago
Apply on www.adzuna.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Unit Testing Ubuntu (Operating System) Cloud Computing Information Systems Databases Continuous Integration Debian Linux Linux Distributed Systems Elasticsearch Information Technology Operations
+15 more
Python (Programming Language) Linux Kernel Routing Open Source Technology Red Hat Enterprise Linux Reliability Engineering Prometheus Software Engineering Ceph (Software) Grafana Firewalls (Computer Science) Storage Technologies Bug Reporting Information Technology Software Version Control

Job description

Join a global Information Systems team responsible for operating and evolving critical production services at significant scale. In this role, you’ll use automation, Infrastructure as Code, and GitOps practices to make cloud operations more reliable, consistent, and efficient. You’ll work across private and public cloud environments, strengthening infrastructure resilience, scalability, observability, and performance. Your expertise will also influence the evolution of open-source infrastructure technologies through hands-on feedback, bug reporting, and collaboration. You’ll troubleshoot complex systems, improve operational processes, and help eliminate repetitive manual work through thoughtful automation. Working with a distributed team of experienced SREs, you’ll have opportunities to share knowledge, mentor colleagues, and contribute to major engineering initiatives. This is an ideal opportunity for an automation-first technologist who is passionate about Linux, open source, and building robust systems at scale. Accountabilities:

  • Apply Infrastructure as Code expertise to continuously improve automation practices, processes, and operational consistency.
  • Automate software operations across private and public clouds while accounting for the complexities of distributed systems.
  • Develop new capabilities and improve the resilience, scalability, and reliability of cloud and container infrastructure.
  • Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliable day-to-day performance.
  • Troubleshoot complex systems, perform capacity planning and performance investigations, and develop strong operational expertise.
  • Set up, maintain, and use observability and monitoring solutions such as Prometheus, Grafana, and Elasticsearch.
  • Design and maintain monitoring and alerting for critical systems and services.
  • Collaborate with development teams on service architecture, documentation, playbooks, policies, and operational procedures.
  • Work closely with globally distributed engineering, operations, and support teams to resolve issues and improve services.
  • Dedicate focused development time to larger engineering projects and the automation of repetitive manual processes.
  • Share technical knowledge and best practices through design sessions, mentoring, and collaborative problem-solving.
  • Take final responsibility for resolving time-critical operational escalations., * Exposure to private and public cloud environments, Infrastructure as Code, GitOps, CI/CD, observability, and open-source technologies.
  • Dedicated development time for impactful automation and larger engineering projects.
  • Collaboration with a highly experienced, globally distributed SRE and engineering community.
  • Opportunities for mentoring, knowledge sharing, and cross-functional technical collaboration.
  • Remote work flexibility, with the role available across time zones.
  • Opportunities to meet colleagues in person 2-4 times per year at internal events, typically lasting 1-2 weeks.
  • International exposure through collaboration with distributed teams and participation in global company events.
  • Compensation and benefits are determined according to the role, location, experience, and applicable company policies.

Requirements

  • Deep experience defining IT operations through code, using version control, peer review, and CI/CD to deploy application and infrastructure changes.
  • Strong modern software engineering practices, including peer review, unit testing, source control management, CI/CD, and Agile methodologies.
  • Significant Python development experience, including work on large or complex projects.
  • Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
  • Familiarity with Linux storage technologies, ranging from Ceph to database systems.
  • Hands-on experience administering enterprise Linux servers.
  • Extensive understanding of cloud computing concepts, architectures, and technologies.
  • Bachelor’s degree or higher, preferably in computer science, software engineering, or a related technical discipline.
  • Strong English communication skills across written and spoken channels, including email, chat, video calls, voice calls, and in-person collaboration.
  • Strong troubleshooting abilities, with the curiosity and persistence to investigate issues from the Linux kernel through to the web layer.
  • Ability to collaborate effectively while knowing when to seek input from teammates and subject-matter experts.
  • Adaptability, willingness to learn quickly, and comfort working in fast-changing technical environments.
  • Ability to thrive within globally distributed teams and collaborate across different locations and time zones.
  • Strong interest in open-source technologies, particularly Ubuntu or Debian.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Structuring agile and interdisciplinary engineering teams

Oliver Zimmert ¡ LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard ¡ World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet ¡ LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos ¡ LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri ¡ World Congress 2023

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani ¡ Europe 2026 Virtual

Videos

See all

Related articles

See all