> Markdown version of [/jobs/ext/3541611-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3541611-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** NexGen Tech Solutions - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Apache HTTP Server, Bash Shell, Cloud Computing, Databases, Dynamic Host Configuration Protocol, Software Debugging, Linux, Network Address Translation, DevOps, Domain Name System (DNS), Ethernet, Firmware, Python (Programming Language), Routing, Prometheus, Data Streaming, TCP/IP, Policy as Code, Network Routers, Transport Layer Security, Google Cloud, Load Balancing, Grafana, Git, Kubernetes, Information Technology, Apache Kafka, Terraform, Oracle Cloud Infrastructure, Golang - **Published:** September 30, 2026 - **Apply:** https://www.thejobnetwork.com/job/d7a44606-ad08-4f66-bda7-dedb2dde59e2/site-reliability-engineer ## About the Role * 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role. \n * Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code. \n * Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers. \n * Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments. \n * Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code. \n * Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis. \n * Strong troubleshooting & debugging skills in Kubernetes platforms. \n * Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring. \n * Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems. \n * Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices. \n * Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents. \n * Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning. \n * Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting. \n * Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams. \n * Bachelor's degree in computer science, engineering or equivalent practical experience. \n, * Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment. ## Description * Experience supporting cloud-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices in a service-provider environment. \n * Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management. \n * Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes. \n * Understanding of access technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks interact. \n * Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices. \n * Experience building auto-remediation, safe self-service operations or internal reliability platforms. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)