Site Reliability Engineer in Network Infrastructure

Jobgether
Germany
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Cloud Computing Complex Networks Network Congestion Continuous Integration Software Debugging Python (Programming Language) Linux Kernel Linux System Administration Network Architecture Network Control Networking Basics Network Service
+9 more
Performance Tuning Reliability Engineering Cloud Services Software Engineering Computer Networking Systems Load Balancing Containerization AI Platforms Low Latency

Job description

This role offers the opportunity to strengthen the reliability and scalability of critical network infrastructure supporting advanced cloud and AI platforms. As a Site Reliability Engineer, you will help design, automate, and operate the systems that enable high-performance digital services at scale. You will focus on building resilient environments, improving observability, and creating safer operational processes across complex network architectures. Working closely with network and platform engineering teams, you will transform operational challenges into long-term solutions through automation and engineering excellence. This position combines software development, infrastructure expertise, and reliability engineering in a fast-moving international environment. You will have significant ownership and the chance to influence the future of large-scale cloud infrastructure. Accountabilities:

As a Site Reliability Engineer in Network Infrastructure, you will be responsible for improving the reliability, performance, and operational maturity of large-scale network systems. You will combine engineering practices with operational expertise to ensure critical infrastructure remains secure, scalable, and efficient.

  • Define and manage reliability objectives for network services, including SLIs, SLOs, availability targets, and error budgets where applicable.
  • Drive reliability improvements across network infrastructure, including services, site readiness, inter-site connectivity, and operational processes.
  • Own incident response activities, lead technical investigations, conduct postmortems, and implement long-term solutions to prevent recurring issues.
  • Build and improve observability systems through meaningful metrics, logs, traces, alerting, and faster troubleshooting workflows.
  • Design safer infrastructure change processes through automation, CI/CD workflows, testing environments, staged deployments, rollback strategies, and auditability.
  • Collaborate closely with network engineers and platform teams to improve system operability and integrate reliability practices into technical designs.
  • Automate operational workflows and continuously improve infrastructure management processes.
  • Contribute to building scalable and resilient network environments that support growing cloud workloads.

Requirements

The ideal candidate has strong experience in site reliability engineering, infrastructure operations, and network systems. You are comfortable working with complex production environments, debugging challenging technical issues, and developing automation to improve reliability and efficiency.

  • Strong knowledge of production Linux environments and a structured approach to troubleshooting complex systems.
  • Solid understanding of networking fundamentals, including control plane and data plane concepts, latency, packet loss, and failure domains.
  • Experience operating high-availability systems and continuously improving their reliability over time.
  • Ability to write and maintain automation and infrastructure software, with experience in Go preferred and Python welcomed.
  • Experience with modern infrastructure tooling, including Infrastructure as Code, CI/CD systems, and container platforms.
  • Strong engineering mindset with the ability to balance operational stability, automation, and continuous improvement.
  • Experience with high-throughput traffic systems such as load balancers, tunneling, decapsulation, NAT64, or similar technologies is a plus.
  • Knowledge of low-level networking performance optimization, including eBPF/XDP, DPDK, perf/ftrace, or Linux kernel networking internals is advantageous.
  • Experience building safe network delivery pipelines, including testing environments, staged rollouts, automated validation, and drift detection is beneficial.
  • Familiarity with large-scale network observability and telemetry solutions is considered a plus.

Benefits & conditions

  • Competitive compensation package.
  • Career growth and continuous learning opportunities.
  • Flexible working environment with a strong focus on ownership and autonomy.
  • Opportunity to work on impactful cloud and AI infrastructure projects.
  • Collaborative culture with experienced international engineering teams.
  • Exposure to cutting-edge technologies in networking, reliability engineering, and cloud platforms.
  • Opportunity to contribute to systems shaping the future of AI infrastructure.
  • Fast-paced environment focused on innovation, meaningful impact, trust, and professional growth.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.de

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:58 min

Building engineering communities and finding technical inspiration

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

2:36 min

Choosing between managed AI platforms and custom governance

Péter Farkas Péter Farkas · Europe 2026 Virtual

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem · WWC 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:37 min

Accessing API documentation and testing remote driving latency

Alexandru Ciinaru Alexandru Ciinaru +3 · WWC 2025

Videos

See all

Related articles

See all