> Markdown version of [/jobs/ext/2093779-site-reliability-engineer-ii](https://www.wearedevelopers.com/jobs/ext/2093779-site-reliability-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II - **Company:** Nationsbenefits - **Location:** Spain - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Bash Shell, C Sharp (Programming Language), Continuous Integration, DevOps, Monitoring of Systems, Python (Programming Language), Log Analysis, PCI Data Security Standards, Windows PowerShell, Reliability Engineering, Spring Cloud, System Availability, Kubernetes, Devsecops, Docker - **Published:** August 16, 2026 - **Apply:** https://es.trabajo.org/oferta-9000-eaba23f1187cbf7fd9af00c3d9791bc2 ## About the Role 2, ISO 27001, and HITRUST. Required Qualifications - 3-5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support. - Hands-on experience with production incident response, troubleshooting, and escalation. - ## Description internal and external customers alike. Our goal is to transform the healthcare industry for the better We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India. Location: Remote (US-Based Candidates Only) Site Reliability Engineer II (SRE) Position Overview We are seeking a Site Reliability Engineer II (SRE) to join our growing Site Reliability Engineering team. In this role, you will help ensure the availability, reliability, and performance of our production platforms by monitoring systems, responding to incidents, troubleshooting infrastructure issues, and driving automation initiatives. You will collaborate closely with Development, DevSecOps, and Engineering teams to maintain highly available cloud-native applications while supporting mission-critical healthcare and fintech services. This position is ideal for someone who enjoys solving production challenges, improving operational efficiency, and working in a fast-paced environment. Key Responsibilities Incident Management - Serve as the first responder for production incidents by identifying, triaging, and resolving issues. - Monitor and respond to alerts generated by Datadog and other monitoring platforms. - Perform initial root cause analysis and escalate incidents according to defined SLAs. - Communicate incident status and resolution updates to internal stakeholders. - Partner with senior engineers to resolve complex production issues. Monitoring & Platform Reliability - Continuously monitor application health, infrastructure performance, and system availability. - Configure and optimize monitoring dashboards and alert thresholds. - Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis. - Support containerized applications running in Kubernetes and Docker environments. Production Support - Participate in a weekday "Follow-the-Sun" production support model with global engineering teams. - Participate in an on-call rotation for critical production systems as needed. - Help maintain high availability and system uptime. Automation & Continuous Improvement - Develop automation scripts and operational tools using one or more of the following: - Python - PowerShell - Bash - C# - Java - Support CI/CD pipeline monitoring and deployment reliability. - Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks. Collaboration - Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams. - Recommend improvements to monitoring, tooling, and operational processes. - Collaborate effectively with globally distributed engineering teams. Documentation & Compliance - Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews. - Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Security Pitfalls for Software Engineers](https://www.wearedevelopers.com/videos/726-security-pitfalls-for-software-engineers) - [DevSecOps culture](https://www.wearedevelopers.com/videos/783-devsecops-culture) ## Related Articles - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs)