> Markdown version of [/jobs/ext/1556563-site-reliability-engineers](https://www.wearedevelopers.com/jobs/ext/1556563-site-reliability-engineers). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineers - **Company:** GitLab - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $126,000.0 - $314,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Build Automation, Continuous Integration, Software Debugging, Ruby, Software Engineering, Systems Architecture, Gitlab, Kubernetes, Terraform - **Published:** July 11, 2026 - **Apply:** https://www.builtincolorado.com/auth/login?destination=/job/site-reliability-engineer-intermediate-senior-staff-infrastructure-platforms/10145126 ## About the Role We don't expect every candidate to have experience with every technology in our environment. We're looking for engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly. We'll support you in becoming successful with GitLab's tools, systems, and ways of working., * Experience keeping production systems reliable, combining an operations mindset with real software engineering practice * Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch * The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes * Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level * Hands-on experience with at least one major cloud provider (GCP or AWS) * Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions * Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure * Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment * A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work * Alignment with GitLab's values and a commitment to working in accordance with them ## Description Because this is a single application for SRE roles across Infrastructure Platforms, our process is built to evaluate you once and match you well, rather than interviewing separately for every team. * Recruiter Screen: A conversation about your background, what you're looking for, and the level and teams that fit, so we can point your process in the right direction. * Core Technical: The shared assessment every SRE candidate takes, regardless of eventual team. A low-stress, collaborative discussion covering source code, system architecture, and incident review. * Peer Technical: Team-specific depth, run by SREs from the team you're most likely to join, focused on the problems that team actually works on. * Hiring Manager Interview: A conversation about ownership, judgment, execution, collaboration, and growth, the non-technical signals that make an SRE effective at GitLab. * Skip-Level Interview: A conversation with a senior leader on values alignment, and how you'll work across teams. After your interviews, we consider your performance alongside our current hiring needs to confirm the level and team where you'll do your best work. Interview results are a major factor, and final placement also reflects our active hiring priorities at the time. What level am I? We calibrate your level during the process, but here is roughly what each looks like so you know where you might land. Intermediate * You make meaningful contributions to reliability, automation, and operational efficiency, working independently within a scoped area * You diagnose issues on your own, understand system dependencies, and can explain the tradeoffs you made * You prioritize well, break work into manageable steps, and use automation to reduce toil * You document your work clearly and keep yourself moving without needing check-ins Senior * You drive reliability improvements across multiple projects or services and prioritize them based on real system needs * You lead investigations, anticipate cascading failures, and coordinate incident response * You own delivery end to end, unblock others, and improve the patterns your team works by * You communicate complex ideas clearly, influence how work gets done, and enable coordination across teams Staff * You shape reliability strategy across teams and services and define patterns that others reuse * You introduce prevention strategies, identify systemic weaknesses, and influence incident response practices beyond your immediate area * You design execution and automation approaches that work at organizational scale * You connect reliability work to platform and business needs Senior Staff * You set technical direction for reliability across a sub-department, not just a team * You drive the hardest, most ambiguous systems problems and establish standards and guardrails that multiple teams adopt * You mentor Staff and Senior engineers * You align reliability strategy with long-range platform direction and represent Infrastructure's interests across the wider Engineering organization What you'll do * Keep user-facing services and production systems reliable, scalable, and efficient * Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows * Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling * Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps * Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately * Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages * Take part in incident response and post-incident reviews, turning learnings into changes in automation and process * Document runbooks, architecture decisions, and reviews so your findings become repeatable practices, Lead and grow a team of Solutions Architects serving public sector customers; drive technical pre-sales strategy, guide evaluations, collaborate with sales and cross-functional teams, improve team practices and execution, track KPIs, and act as a hands-on architect for strategic opportunities. ## Related Videos - [GitOps for the people](https://www.wearedevelopers.com/videos/815-gitops-for-the-people) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [GitOps keeps focus on apps, not on infrastructure](https://www.wearedevelopers.com/videos/182-gitops-keeps-focus-on-apps-not-on-infrastructure) ## Related Articles - [Stop Googling Git Commands. Start Actually Learning Git.](https://www.wearedevelopers.com/magazine/730-stop-googling-git-commands-start-actually-learning-git) - [Why SmartGit Is More Than a Git Client](https://www.wearedevelopers.com/magazine/689-why-smartgit-is-more-than-a-git-client) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Top 10 GitHub Alternatives (2025)](https://www.wearedevelopers.com/magazine/298-the-top-10-github-alternatives-2025) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care)