> Markdown version of [/jobs/ext/538629-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/538629-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** NSCALE, LLC - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $100,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Computer Programming, Computer Networks, Data Centers, Distributed Systems, InfiniBand, Python (Programming Language), Remote Direct Memory Access, Reliability Engineering, Software Engineering, High Performance Computing, Grafana, Kubernetes, Bare Metal, Operational Systems - **Published:** June 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=22f90bf683ea3800 ## About the Role Do you have experience in System troubleshooting?, * 2-5 years of experience in Site Reliability Engineering, Systems Engineering, or Software Engineering in Data Center Environment * 2+ years programming skills (e.g., Python, Go, or similar) with interest in automation and tooling * Working knowledge of Linux systems, networking concepts, and distributed systems * Experience troubleshooting system or application issues in production environments * Familiarity with monitoring or observability tools (e.g., logs, metrics, dashboards) * Strong willingness to learn and improve reliability and operational practices * Ability to work in fast-paced environments and collaborate across teams Preferred Experience * Exposure to cloud platforms, Kubernetes, or virtualized/bare-metal environments * Experience in AI, GPU workloads, or high-performance computing (HPC) * Basic understanding of high-performance networking concepts (e.g., InfiniBand, RDMA) * Exposure to production monitoring or alerting systems at small or medium scale ## Description * Help build and improve automation, tooling, and infrastructure that supports AI workloads * Support the development of operational systems and platform services * Assist in defining and maintaining basic SLOs/SLIs and monitoring dashboards * Participate in incident response, troubleshooting, and post-incident reviews * Investigate and help resolve performance and reliability issues across systems * Collaborate with Engineering, Networking, and Infrastructure teams to improve system stability * Contribute to improving availability, scalability, and operational efficiency * Learn from senior engineers and grow your expertise in reliability engineering ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)