> Markdown version of [/jobs/ext/2710884-site-reliability-engineer-ii](https://www.wearedevelopers.com/jobs/ext/2710884-site-reliability-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II - **Company:** WAYSTAR HEALTHCARE, LLC - **Location:** Louisville, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Microsoft Azure, Bash Shell, Configuration Management, Databases, Software Debugging, Linux, DevOps, Distributed Systems, Github, Python (Programming Language), Networking Basics, NoSQL, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Ruby, Runbook, Software Deployment, Datadog, Scripting, Google Cloud, Grafana, Reliability of Systems, Cloudformation, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Terraform, Splunk, Docker, Jenkins, Golang - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-ii-waystar-inc-8504734 ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. * 5+ years of experience in a Site Reliability Engineering, DevOps, or highly related infrastructure engineering role. * Strong proficiency in at least one scripting/programming language (e.g., Python, Go, Java, Ruby, Bash). * Extensive experience with cloud platforms (AWS, Azure, GCP) including services related to compute, networking, storage, and databases. * Deep understanding of Linux operating systems and networking fundamentals. * Proven experience with infrastructure as code tools (e.g., Terraform, CloudFormation, Ansible). * Solid experience with CI/CD pipelines and related tools (e.g., Jenkins, GitLab CI, GitHub Actions). * Demonstrable expertise in monitoring and alerting systems (e.g., Prometheus, Grafana, Datadog, Splunk). * Strong problem-solving skills with a methodical approach to debugging complex distributed systems. * Excellent communication and collaboration skills, with the ability to work effectively across cross-functional teams. * Experience with containerization technologies (Docker, Kubernetes) is highly desirable. * Familiarity with database technologies (relational and NoSQL) and their operational challenges. ## Description * Design, implement, and maintain automation for infrastructure provisioning, configuration management, and application deployments across various environments (on-premise and cloud). * Proactively monitor system health, performance, and availability, utilizing a range of observability tools and defining key performance indicators (KPIs) and service level objectives (SLOs). * Lead the investigation and resolution of complex production incidents, perform root cause analysis, and implement preventative measures to minimize future occurrences. * Collaborate with development teams to ensure software is designed for reliability, scalability, and operational efficiency, participating in architectural reviews and providing expert guidance. * Develop and maintain robust incident response procedures, runbooks, and disaster recovery plans. * Contribute to the evolution of our SRE practices, tooling, and best standards, driving continuous improvement and knowledge sharing within the team. * Participate in an on-call rotation to provide 24/7 support for critical production systems. * Mentor junior SREs and contribute to the growth and development of the team. * Evaluate and implement new technologies and solutions to enhance system reliability and operational efficiency. ## Related Videos - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)