> Markdown version of [/jobs/ext/2707943-senior-software-engineer-sre-aiops](https://www.wearedevelopers.com/jobs/ext/2707943-senior-software-engineer-sre-aiops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - SRE & AIOps - **Company:** ServiceNow - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $143,200.0 - $243,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Microsoft Azure, Bash Shell, Cloud Computing, Computer Networks, Computer Engineering, Continuous Integration, Data Centers, DevOps, Distributed Systems, Fault Tolerance, Python (Programming Language), Network Security, Linux System Administration, Machine Learning, Reliability Engineering, Runbook, Software Engineering, Scripting, Application Enhancement Tool, Cloud Platform System, Istio, Mttr, Software Troubleshooting, Multi-Cloud, HybridCloud, Cloudformation, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Hardware Infrastructure, Cloud Migration, Terraform, Servicenow, Golang, Microservices - **Published:** September 4, 2026 - **Apply:** https://jobs.smartrecruiters.com/ServiceNow/744000147282649-senior-software-engineer-sre-aiops ## About the Role * Kubernetes Proficiency: Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues. * Incident Remediation Experience: Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self-healing mechanisms. * Cloud Platform Knowledge: Strong hands-on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE-related services. * DevOps & IaC Skills: Solid experience with Infrastructure-as-Code tools (Terraform, CloudFormation) and GitOps practices. * SRE Tooling Familiarity: Working knowledge of observability platforms, incident management systems, and log aggregation tools. * Distributed Systems Understanding: Understanding of distributed system challenges, fault tolerance, and resilience patterns. * On-Call Operations: Experience participating in on-call rotations and understanding 24/7 operational models, runbook development, and escalation procedures. * Cloud & Hybrid Operations: Hands-on experience working with cloud infrastructure and understanding hybrid cloud concepts. * Systems Administration: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, or Bash). * Collaborative Mindset: Ability to work effectively with infrastructure and application teams, contribute to technical discussions, and help drive reliability improvements., * Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry. * 5+ years in software engineering or infrastructure operations, with 3+ years in SRE, DevOps, or cloud platform engineering roles with a Bachelor's degree; or 3 years and a Master's degree; or a PhD without experience; or equivalent work experience. * 2+ years of hands-on experience working with production Kubernetes clusters. * Proficiency in at least one Infrastructure-as-Code tool: Terraform, CloudFormation, or equivalent. * Demonstrable hands-on experience with at least one major cloud platform: AWS, Azure, or GCP. * Experience operating in on-call environments and participating in incident response. * Experience implementing or improving automated remediation and alert systems. * Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash). * Demonstrated commitment to reliability engineering and continuous improvement through hands-on contributions. * Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience). Preferred: * Kubernetes certification (CKA, CKAD, or equivalent). * Experience with service mesh technologies or advanced Kubernetes networking. * Background in cloud migration or infrastructure modernization projects. * Experience with cost optimization in cloud environments. * Track record of implementing automation solutions that significantly reduced operational toil. ## Description ServiceNow is seeking a Senior Software Engineer - SRE & AIOps to contribute to infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale. This role combines solid hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with growing technical leadership capabilities. You will contribute to SRE tooling design, develop auto-remediation capabilities, and help establish patterns that maintain ServiceNow's cloud platform reliability while minimizing operational toil across follow-the-sun global teams. What you get to do in this role: * Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments. * Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden. * Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations. * Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue. * Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails. * Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments. * Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls. * Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews that drive continuous improvement. * Share knowledge and mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices. * Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation. * Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity., We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)