> Markdown version of [/jobs/ext/202465-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/202465-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Avaya Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $143,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Application Performance Management, Microsoft Azure, Cloud Computing, Continuous Integration, DevOps, Distributed Systems, Github, Log Analysis, Reliability Engineering, Ansible, Prometheus, Datadog, Grafana, Mttr, Multi-Cloud, Terraform, Dynatrace, Jenkins - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=93cc248964e1485d ## About the Role * 5+ years in Site Reliability, DevOps, Cloud Operations, or Customer support roles. * Demonstrated experience in application-level troubleshooting by analyzing logs and traces to identify bugs, performance bottlenecks, and error conditions. * Expertise in Azure and GCP cloud operations and distributed system reliability. * Understanding of Terraform, Ansible, and CI/CD pipelines (Jenkins, GitHub Actions). * Experience with observability and AI-Ops tools (Azure Monitor, GCP Operations Suite, Grafana, Prometheus, Datadog, etc.). * Solid grasp of incident management frameworks (P1-P3 handling, RCA, PIRs, on-call rotations). * Excellent analytical, troubleshooting, and communication skills. Desired Behaviours * Proactive Prevention: Identifies and resolves risks before they escalate into incidents. * AI-Driven Mindset: Applies AI and automation to improve reliability and reduce human intervention. * Accountability: Owns service reliability and communicates with clarity. * Collaboration: Works seamlessly with platform, DevOps, and product teams. * Efficiency: Focuses on automation to reduce manual effort and improve MTTR. * Continuous Improvement: Learns from failures, iterates processes, and enhances documentation., Applicants must be currently authorized to work in the United States without the need for visa sponsorship now or in the future. ## Description We are seeking a Site Reliability Engineer (SRE) who will drive stability, reliability, and performance across our Azure and GCP-based platforms. This role blends operational excellence, proactive incident management, and strong collaboration with DevOps, Cloud, and Security teams. The ideal candidate will have hands-on experience with multi-cloud environments (Azure and GCP), IaC (Terraform/Ansible), CI/CD (Jenkins/GitHub Actions), and modern observability and AI-Ops systems. The engineer will also contribute to governance, cost optimization, and automation strategies that reduce toil and prevent issues before they occur. A key aspect of this role is the ability to perform deep-dive troubleshooting of application performance and errors by analyzing logs and traces in platforms like Grafana and Datadog. This position includes 24×7 support coverage (rotational) and requires strong ownership in managing major incidents, RCA processes, and continuous service improvements., Reliability & Incident Management * Serve as a key member of the 24×7 on-call rotation, responding to and managing incidents across production and pre-production environments. * Lead incident bridges, coordinate root cause analysis (RCA), and ensure post-incident reviews drive systemic improvements. * Maintain clear communication with cross-functional teams and leadership during major incidents. Monitoring, AI-Ops, Alerts & Prevention * Build, tune, and maintain observability dashboards (Azure Monitor, GCP Operations Suite, Prometheus, Grafana, Datadog, Log Analytics). * Perform deep-dive troubleshooting of application and service-level issues using distributed tracing and log analysis (Grafana, Datadog) to pinpoint root causes beyond infrastructure. * Define SLOs, SLIs, and error budgets to proactively identify and mitigate reliability risks before customer impact. * Integrate AI-Ops tools for anomaly detection, predictive alerting, and automated incident correlation. * Continuously enhance alert quality, reduce false positives, and automate runbooks for faster recovery. * Analyze trends to prevent recurring issues and support teams in resilience engineering. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)