> Markdown version of [/jobs/ext/2545552-observability-sre-engineer](https://www.wearedevelopers.com/jobs/ext/2545552-observability-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Observability & SRE Engineer - **Company:** World Wide Technology - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $90,200.0 - $112,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Linux, Distributed Systems, Fault Tolerance, Github, Information Technology Operations, Python (Programming Language), Machine Learning, Regression Testing, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Load Balancing, Test-Driven Development (TDD), Grafana, Caching, Containerization, Kubernetes, Information Technology, Data Analytics, Terraform, Splunk, Dynatrace, Docker, Programming Languages - **Published:** August 15, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27938267/Observability-Sre-Engineer-Any-Remote-Nationwide-7449 ## About the Role The ideal candidate is a passionate technologist who enjoys instrumenting, measuring, and improving systems to strengthen observability, reliability, and operational decision-making. They bring an AI-first mindset, using AI, automation, and data-driven practices responsibly to improve reliability, accelerate response, and reduce manual effort., * 5+ years of professional experience in IT Operations * Experience in Metrics, Monitoring and Alerting, including tools like Prometheus and Grafana, Big Panda, Splunk, etc. * Knowledge of programming languages Python, Go, or equivalent * Knowledge of system frameworks including Git and GitHub * Critical thinker with excellent written and verbal communication skills * Team-oriented individual with very strong work ethic * Familiarity with Linux, preferably administrative knowledge * Understanding of container technologies (Docker, Podman, etc.) * Understanding of SRE concepts such as SLIs, SLOs, incident response, root cause analysis, capacity planning, and reliability automation * Practical understanding of AI-enabled tools, automation patterns, and responsible AI use to improve operational efficiency * Willingness to participate in an on-call rotation and drive incidents through triage, mitigation, and resolution * Experience with Infrastructure as Code (Terraform, Ansible, or similar) and CI/CD pipelines * Understanding of distributed systems fundamentals: fault tolerance, redundancy, load balancing, and caching, * Kubernetes experience * Experience working with Agile methodology * Test-First development mindset with functional, end-to-end and regression testing experience * Experience with AI-assisted engineering, AIOps, or machine learning techniques for monitoring, alerting, or incident response * Hands-on SRE experience with SLOs, error budgets, blameless reviews, toil reduction, and production readiness * Experience with chaos engineering or resilience/failure-injection testing * Experience authoring and maintaining runbooks, playbooks, and production readiness review checklists * Public cloud platform experience (AWS, Azure, or GCP) * Familiarity with distributed tracing and log aggregation tooling (e.g., OpenTelemetry, ELK/Splunk) ## Description The Observability & SRE team brings together infrastructure, operations, automation, and reliability engineering skills to build consumable services and platforms that improve visibility, confidence, and resilience across WWT's IT systems. This is an opportunity for someone looking to grow technically while helping modernize IT through AI-enabled operations and reliability-focused engineering., * Collection and strategic application of metrics to drive organizational decisions * Providing a holistic view of system health using observability practices * APM, RUM, and Synthetic Transaction monitoring * Driving reliability through monitoring, alerting, observability, and SRE practices * Applying AI-first thinking to automate workflows, correlate alerts, surface insights, and speed incident response * Improving reliability through SLOs, SLIs, error budgets, incident reviews, and continuous improvement * Defining and maintaining on-call rotations, escalation paths, and incident response processes, including participating in on-call coverage * Reducing operational toil through automation, self-healing systems, and infrastructure as code * Capacity planning and performance engineering to ensure systems scale reliably under load * Partnering with development teams on production readiness reviews, architecture reviews, and resilience/chaos testing to prevent incidents before they happen ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship)