> Markdown version of [/jobs/ext/2586278-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2586278-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Cognizant Technology Solutions Corporation - **Location:** Bentonville, AR, United States - **Experience:** Experienced - **Salary:** $70,000.0 - $80,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Microsoft Azure, Software Debugging, Fault Tolerance, Java Database Connectivity, Model View Controller (MVC), Reliability Engineering, Software Engineering, Web Application Frameworks, Real Time Systems, Software Troubleshooting, Restful APIs, Microservices - **Published:** August 16, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3355968508&tx=DT11202UYD&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * 3+ years of experience in software engineering, with a focus on reliability, infrastructure, and platform engineering. * Proficiency in Java, with a solid understanding of its ecosystem and Microservice Architecture. * Basic understanding of MVC (Model-View-Controller) patterns, JDBC (Java Database Connectivity), and RESTful web services. * Experience working with web application frameworks. * Ability to write clean, readable, and maintainable Java code. * Strong troubleshooting and debugging skills, with a track record of resolving production issues efficiently. * Working knowledge of cloud platforms (Azure). * Solid grounding in Site Reliability Engineering (SRE) principles and best practices. * Willingness and availability to support off-hours and weekend on-call rotations. ## Description We are seeking a Site Reliability Engineer to support platform stability, reliability, automation, and production operations. This role will partner with cross-functional teams to improve resilient systems, troubleshoot production issues, strengthen observability, and apply site reliability engineering best practices. In This Role, You Will * Collaborate with cross-functional teams to gather and understand the environment and build and plan for system and platform stability. * Design and lead the implementation of fault-tolerant platforms, employing redundancy, self-healing mechanisms, and graceful degradation techniques to maintain service continuity despite unexpected failures. * Troubleshoot and debug software issues, providing timely resolutions. * Collaborate with the team to continuously improve platform stability and implement SRE best practices. * Drive automation initiatives across infrastructure and operational workflows. * Build and maintain observability dashboards for real-time system visibility. * Drive incident reduction efforts and lead resolution of critical (P1/P2) issues. * Self-motivate to support incident resolution and incident prevention efforts. * Provide off-hours and weekend on-call support for critical production issues. Work Model We strive to provide flexibility wherever possible. Based on this role's business requirements, this is a onsite position open to qualified applicants able to work in Bentonville AR 72712. Regardless of your working arrangement, we are here to support a healthy work-life balance through our various wellbeing programs. The working arrangements for this role are accurate as of the date of posting and may change based on business and client requirements. ## Related Videos - [Developer Tools for Microsoft Azure](https://www.wearedevelopers.com/videos/450-developer-tools-for-microsoft-azure) - [Rest API Antipatterns](https://www.wearedevelopers.com/videos/100208-rest-api-antipatterns) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [REST In Peace: Why LLMs Can't CRUD](https://www.wearedevelopers.com/videos/100272-rest-in-peace-why-llms-can-t-crud) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)