> Markdown version of [/jobs/ext/1315345-senior-site-reliability-engineer-ai-infrastructure-operations](https://www.wearedevelopers.com/jobs/ext/1315345-senior-site-reliability-engineer-ai-infrastructure-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer -AI Infrastructure Operations - **Company:** NSCALE, LLC - **Location:** Houston, TX, United States - **Experience:** Expert - **Salary:** $170,000.0 - $265,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Cloud Computing, Data Centers, Linux, Distributed Systems, InfiniBand, Python (Programming Language), Remote Direct Memory Access, Reliability Engineering, Software Engineering, AI Infrastructure, High Performance Computing, Kubernetes, Bare Metal, Build Tools - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=47352ccbb0ebd086 ## About the Role * 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production at scale in a data center or cloud environment. * Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt, not scripts that run once and rot. * Deep command of Linux, networking, and distributed systems, plus the judgment to know where the real failure modes hide. * Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the metal, not just the cloud console. * Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to get there fast. * Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale, incident process, on-call that people can actually live with. * A track record as the senior voice in incidents and design reviews, trusted to make the call under pressure. * A habit of raising the people around you without being asked to. Nice to Have * Familiarity with high-performance networking (InfiniBand, RDMA). ## Description This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up to it. You'll own the hardest problems on the platform: the automation other engineers build on, the services that can't go down, and the design decisions that determine whether either holds up at scale. You'll still carry a pager, but the real job is making sure it fires less, for everyone, over time. What You'll Do * Own reliability for critical production services end to end; set the direction, not just respond to what breaks. * Grow the team, not just the systems; mentor other SREs through design review, pairing, and incident debriefs, and hold the bar that pulls everyone up to it. * Set the standards the rest of the team works to: the SLO framework, the incident process, and the on-call practices that keep it sustainable. * Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted on after the first outage. * Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change that keeps it from coming back. * Build the tooling and automation that removes toil for the whole team, not just your own surface ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)