> Markdown version of [/jobs/ext/2872624-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2872624-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** TransUnion LLC - **Location:** Remote, OR, United States - **Experience:** Expert - **Salary:** $112,500.0 - $187,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Engineering, Linux, Load Testing, Performance Tuning, Reliability Engineering, Prometheus, Systems Integration, Datadog, Google Cloud, Cloud Platform System, Large Language Models, Grafana, Kubernetes, Pagerduty - **Published:** September 13, 2026 - **Apply:** https://www.careerjet.com/job/usabbca99e57cc6ceea4e03b7c3cbb020b/eaa ## About the Role * 5+ years of experience in Cloud Architecture, Site Reliability Engineering, Platform Engineering, or related fields - with a proven track record of designing and delivering at enterprise scale. * Deep, hands-on expertise with Google Cloud Platform (GCP) and Kubernetes (K8s) - running high-volume, high-availability workloads with 99.999% reliability targets. * Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog, Prometheus, Grafana, PagerDuty) - you define what good looks like. * Deep Linux expertise - from kernel internals and system performance tuning to hardening and troubleshooting at the OS level in production environments.. * Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms - including familiarity with LLM orchestration, vector databases, model serving infrastructure, and AI observability - with the ability to evaluate emerging tools and translate them into reliable, production-grade capabilities. ## Description * Recognized expert across multiple systems; actively contributes to architectural and strategic decisions around major platform components. * Leads research, testing, implementation, and continuous improvement for new systems and tooling. * Performs complex, high-impact work including capacity planning, load testing, and security improvements. * Fully participates in the team's on-call rotation; models calm, effective, and blameless incident response. * Serves as a significant technical contributor during major incidents and problem resolution. * Plans and leads high-risk maintenance events with minimal to no customer impact. * Elevates team standards through new tooling, processes, procedures, and effective communication. * Capable of stepping in to lead and represent the team - a trusted resource during transitions or coverage gaps. * Sets new professional benchmarks in technical quality, engineering culture, and cross-functional collaboration. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)