> Markdown version of [/jobs/ext/2150689-infrastructure-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2150689-infrastructure-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Site Reliability Engineer - **Company:** Nebius Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $180,000.0 - $224,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Continuous Integration, Software Debugging, Linux, Python (Programming Language), Network Control, Networking Basics, Software Deployment, Load Balancing, Delivery Pipeline, Perf (Linux), Containerization, Low Latency - **Published:** August 20, 2026 - **Apply:** https://careers.nebius.com/?gh_jid=4906462101 ## About the Role * Strong production Linux fundamentals and a structured approach to debugging complex systems * Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.) * Hands-on experience operating high-availability systems and improving them over time (not just "keeping lights on") * Ability to write and maintain software/automation (Go is common for us; Python is also welcome) * Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows, * Experience with high-throughput traffic processing: load balancers, tunneling/decap, NAT64, or similar datapath-heavy systems * Low-level networking performance/debug background (eBPF/XDP, DPDK, perf/ftrace, kernel networking internals) * Experience building network-safe delivery pipelines (testing labs, staged rollouts, automated verification, drift detection) * Background with large-scale network observability/telemetry (e.g., routing/flow telemetry, regression detection at scale), Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. ## Description We're looking for a Network Site Reliability Engineer (NetSRE) to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you'll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly. Your responsibilities will include: * Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense) * Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards * Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting) * Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents * Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes * Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)