Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
We are seeking a Site Reliability Engineer to support platform stability, reliability, automation, and production operations. This role will partner with cross-functional teams to improve resilient systems, troubleshoot production issues, strengthen observability, and apply site reliability engineering best practices.
In This Role, You Will
- Collaborate with cross-functional teams to gather and understand the environment and build and plan for system and platform stability.
- Design and lead the implementation of fault-tolerant platforms, employing redundancy, self-healing mechanisms, and graceful degradation techniques to maintain service continuity despite unexpected failures.
- Troubleshoot and debug software issues, providing timely resolutions.
- Collaborate with the team to continuously improve platform stability and implement SRE best practices.
- Drive automation initiatives across infrastructure and operational workflows.
- Build and maintain observability dashboards for real-time system visibility.
- Drive incident reduction efforts and lead resolution of critical (P1/P2) issues.
- Self-motivate to support incident resolution and incident prevention efforts.
- Provide off-hours and weekend on-call support for critical production issues.
Work Model
We strive to provide flexibility wherever possible. Based on this role’s business requirements, this is a onsite position open to qualified applicants able to work in Bentonville AR 72712. Regardless of your working arrangement, we are here to support a healthy work-life balance through our various wellbeing programs. The working arrangements for this role are accurate as of the date of posting and may change based on business and client requirements.
Requirements
- 3+ years of experience in software engineering, with a focus on reliability, infrastructure, and platform engineering.
- Proficiency in Java, with a solid understanding of its ecosystem and Microservice Architecture.
- Basic understanding of MVC (Model-View-Controller) patterns, JDBC (Java Database Connectivity), and RESTful web services.
- Experience working with web application frameworks.
- Ability to write clean, readable, and maintainable Java code.
- Strong troubleshooting and debugging skills, with a track record of resolving production issues efficiently.
- Working knowledge of cloud platforms (Azure).
- Solid grounding in Site Reliability Engineering (SRE) principles and best practices.
- Willingness and availability to support off-hours and weekend on-call rotations.
Benefits & conditions
The annual salary for this position is between $70,000 - $80,000 depending on experience and other qualifications of the successful candidate.
This position is also eligible for Cognizant’s discretionary annual incentive program, based on performance and subject to the terms of Cognizant’s applicable plans., Cognizant offers the following benefits for this position, subject to applicable eligibility requirements:
Medical/Dental/Vision/Life Insurance
Paid holidays plus Paid Time Off
401(k) plan and contributions
Long-term/Short-term Disability
Paid Parental Leave
Employee Stock Purchase Plan
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
What is Software Engineering?
The Best Software Developer Blogs to Read
Fully Remote Software Engineer Jobs