Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+20 more
Job description
-
Experience supporting cloud-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices in a service-provider environment. \n
-
Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management. \n
-
Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes. \n
-
Understanding of access technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks interact. \n
-
Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices. \n
-
Experience building auto-remediation, safe self-service operations or internal reliability platforms.
Requirements
-
5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role. \n
-
Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code. \n
-
Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers. \n
-
Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments. \n
-
Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code. \n
-
Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis. \n
-
Strong troubleshooting & debugging skills in Kubernetes platforms. \n
-
Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring. \n
-
Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems. \n
-
Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices. \n
-
Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents. \n
-
Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning. \n
-
Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting. \n
-
Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams. \n
-
Bachelor’s degree in computer science, engineering or equivalent practical experience. \n, * Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
What’s the Difference between a Junior, Mid, and Senior Developer?
Best Countries for Software Engineers
Fully Remote Software Engineer Jobs