> Markdown version of [/jobs/ext/2689560-usa-senior-software-engineer](https://www.wearedevelopers.com/jobs/ext/2689560-usa-senior-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # (USA) Senior, Software Engineer - **Company:** Wal-Mart Stores, Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $117,000.0 - $234,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Cloud Computing, Computer Engineering, Continuous Availability, Linux, DevOps, Intrusion Detection and Prevention, Python (Programming Language), Reliability Engineering, Prometheus, Shell Script, Software Engineering, SQL Databases, Web Platforms, Google Cloud, Grafana, Reliability of Systems, Infrastructure as Code (IaC), Kubernetes, Information Technology, Performance Monitor, Restful APIs, Splunk, Dynatrace, Api Management, Docker, Microservices - **Published:** September 3, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3375520092&tx=HT7266TYT&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Proven experience supporting large-scale Omnichannel eCommerce platforms, with expertise in Major Incident Management, production operations, incident triage, escalation management, and rapid service restoration. * Demonstrated success leading cross-functional incident response teams and driving operational excellence across highly available, customer-facing systems. * Strong technical knowledge of Linux/Unix, Microsoft Azure, Google Cloud Platform (GCP), Kubernetes, Docker, networking, RESTful APIs, and distributed microservices architectures. * Hands-on experience with enterprise observability and monitoring solutions, including Grafana, Prometheus, Splunk, Dynatrace, and OpenTelemetry. * Proficient in Python, Java, JavaScript, Shell scripting, SQL, REST API development and integration, and workflow automation. * Experience implementing Site Reliability Engineering (SRE) best practices, CI/CD pipelines, Infrastructure as Code (IaC), self-healing capabilities, and AI-assisted operational solutions. * Excellent analytical, problem-solving, communication, collaboration, and technical leadership skills, with the ability to influence and drive outcomes across cross-functional teams., Outlined below are the required minimum qualifications for this position. If none are listed, there are no minimum qualifications. Option 1: Bachelor's degree in computer science, computer engineering, computer information systems, software engineering, or related area and 3 years' experience in software engineering or related area. Option 2: 5 years' experience in software engineering or related area. Preferred Qualifications... Outlined below are the optional preferred qualifications for this position. If none are listed, there are no preferred qualifications. Master's degree in Computer Science or related field and 2 years' experience in software engineering or related field ## Description Walmart is seeking a Senior Software Engineer - Site Reliability Operations Lead to drive the reliability and operational excellence of its global Omnichannel eCommerce platforms. In this role, you will provide technical leadership for 24x7 production operations, proactive monitoring, major incident management, API integrations, observability, and automation initiatives. Partnering closely with Engineering, Platform, Cloud, and Product teams, you will design and optimize intelligent operational workflows, enhance system reliability, automate critical processes, and improve customer experience through modern Site Reliability Engineering (SRE) practices and AI-driven operational solutions. About the team: Our team serves as Walmart's Command & Control Center for Omnichannel eCommerce, ensuring continuous availability, reliability, and performance of customer-facing digital platforms. We lead proactive monitoring, incident detection, triage, and Major Incident Management to minimize disruptions and support seamless shopping experiences. By advancing enterprise observability, automation, and AI-assisted capabilities, we enhance incident response and platform resilience. Collaborating closely with engineering teams, we perform root cause analysis and drive continuous improvements to strengthen operational stability and improve the reliability of Walmart's global eCommerce ecosystem. What you'll do: * Lead the monitoring, detection, triage, and resolution of major production incidents across Walmart's Omnichannel eCommerce platforms. * Serve as the incident commander during critical outages, leading bridge calls and coordinating cross-functional teams to restore services as quickly as possible. * Design, develop, and maintain RESTful API integrations, operational tools, and workflow automation solutions. * Continuously enhance monitoring, observability, and alerting capabilities using enterprise platforms and Site Reliability Engineering (SRE) best practices. * Diagnose and resolve complex issues across Linux/Unix systems, cloud infrastructure, networking, Kubernetes, APIs, and microservices-based environments. * Implement automation and self-healing capabilities through DevOps methodologies and Infrastructure as Code (IaC) practices. * Leverage AI-powered operational solutions to improve incident response, accelerate issue resolution, and increase operational efficiency. * Lead root cause analysis (RCA) initiatives and drive continuous improvements in system reliability, automation, performance, and operational excellence. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Navigating the Corporate Jungle: Life as a Developer in a large Company](https://www.wearedevelopers.com/videos/621-navigating-the-corporate-jungle-life-as-a-developer-in-a-large-company) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology)