> Markdown version of [/jobs/ext/1544470-sr-engineer-site-reliability](https://www.wearedevelopers.com/jobs/ext/1544470-sr-engineer-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr Engineer - Site Reliability - **Company:** World Wide Technology - **Location:** St. Louis, MO, United States (Remote available) - **Experience:** Expert - **Salary:** $108,400.0 - $135,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Programming, Databases, Continuous Integration, Data Sharing, Linux, DevOps, Distributed Systems, Github, Python (Programming Language), PostgreSQL, Enterprise Messaging Systems, MongoDB, OpenShift, RabbitMQ, Redis, Reliability Engineering, Prometheus, Runbook, Software Engineering, Datadog, Scripting, Istio, Grafana, Caching, Build Management, Kubernetes, Low Latency, Jenkins - **Published:** July 7, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27828043/Sr-Engineer-Site-Reliability-Missouri-St-Louis-7449 ## About the Role * 5+ years of experience in Site Reliability Engineering, Platform Engineering, or related roles supporting production systems * Strong experience with Kubernetes-based platforms; OpenShift experience preferred * Proven experience designing and operating distributed systems at scale * Experience implementing monitoring, alerting, and incident response practices * Strong scripting or programming skills (Python, Go, or similar) with a focus on automation * Experience with CI/CD pipelines and GitOps workflows (GitHub, Jenkins, or similar) * Experience with observability tooling (Prometheus, Grafana, ThousandEyes, or similar) * Strong Linux and troubleshooting skills across complex systems * Excellent communication skills with the ability to collaborate across technical and non-technical teams, * Experience supporting internal developer platforms or platform engineering organizations * Familiarity with shared services such as databases, messaging systems, and caching layers (MongoDB, PostgreSQL, Redis, RabbitMQ) * Experience implementing SLOs and error budget frameworks * Exposure to service mesh, traffic management, or advanced Kubernetes networking * Experience mentoring engineers and influencing technical decision-making Preferred Location: MO ## Description The INFOPS Application Development team focuses on enabling developers at scale. We build and evolve internal platforms, automation, and self-service capabilities that allow application teams to deploy and operate software safely and efficiently. Our work emphasizes: * Platform Engineering best practices * Kubernetes and OpenShift-based platforms * GitOps-driven CI/CD automation * Shared data and messaging platforms (MongoDB, PostgreSQL, RabbitMQ, etc.) * AI-assisted and automation-driven engineering workflows As the organization matures, we are evolving toward a model that blends Platform Engineering with Site Reliability Engineering (SRE) to ensure our platforms are not only scalable-but also highly reliable and resilient., * Improve platform availability, latency, and scalability across Kubernetes and supporting services, establishing and evolving reliability standards across shared platform capabilities, and measuring outage impact as a health signal to track progress in preventing, detecting, and recovering from failures * Design and implement monitoring, alerting, and observability frameworks across platform services, standardizing telemetry (metrics, logs, traces) for consistent visibility across environments, and proactively identifying risks and performance bottlenecks before they impact users * Lead triage and resolution of platform-related incidents, driving root cause analysis, long-term fixes, and blameless postmortem practices * Collaborate with Support Engineers on the incident queue to surface recurring failure patterns, develop long-term remediation plans, reduce operational toil through automation, and drive sustainable resolution of systemic issues * Design and build automation to reduce manual intervention, improve incident response and recovery time, and enable self-healing platform behaviors; develop reusable tools, runbooks, and operational patterns that scale across teams * Collaborate with INFOPS Engineering and the Support Automation team to hand off support work where appropriate * Partner with internal customers, application teams, and platform/DevOps engineering functions to deliver scalable, repeatable solutions, address shared reliability challenges, and promote a culture of operational excellence and continuous improvement, We strive to create an environment where all employees are empowered to succeed based on their skills, performance, and dedication. Our goal is to cultivate a culture of belonging that encourages innovation, collaboration, and respect for all team members, ensuring that WWT remains a great place to work for All! ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)