> Markdown version of [/jobs/ext/2963131-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2963131-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** HCA Healthcare Inc. - **Location:** Nashville, TN, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Microsoft Windows, Microsoft Azure, Health Informatics, C Sharp (Programming Language), Client Server Models, Cloud Computing, Cloud Engineering, Continuous Integration, Linux, DevOps, Monitoring of Systems, Python (Programming Language), Uptime, Open Source Technology, Performance Tuning, Reliability Engineering, Site Reliability Engineering Practices, Cloud Services, Software Engineering, Software Systems, Systems Architecture, TypeScript, Swift (Programming Language), Git, Kotlin, Information Technology, Web Technologies, Terraform, Software Version Control, Programming Languages - **Published:** September 17, 2026 - **Apply:** https://dejobs.org/x/x/B35726A5CAB048DDB9CA9C2AC1AF84E8/job/ ## About the Role * Bachelor's degree Computer Science or related field preferred * 8+ years of experience a Software development or engineering roles required Or equivalent combination of education and/or experience Microsoft Certified: Azure Solutions Architect Expert preferred Microsoft Certified: DevOps Engineer Expert preferred Knowledge, Skills, Abilities, Behaviors: * Knowledge of infrastructure, frameworks, and software/cloud design patterns for implementing applications in the cloud preferred * Experience in the use and implementation of relevant tools and platforms (e.g., cloud platforms (IaaS and PaaS), web technologies, client-server technologies, continuous integration, and deployment) required * Experience with version control (Git) and open-source practices preferred * Experience in one or more coding languages. (JavaScript/Typescript, C#, Python, Java, Swift or Kotlin) preferred * Experience with automation of CI/CD pipelines preferred * Experience with IaC such as Terraform preferred * A proactive approach to spotting problems, areas for improvement, and performance bottlenecks required * Be a creative thinker, not bound by "the way things have always been done". What you know is less important than how well you learn and innovate. We don't need engineers who know all the answers; we need engineers who can invent the answers no one has thought of yet, to the questions yet to be asked required * Experienced in helping define SLIs, SLOs & SLOs, and the experience to build observability to report on operating against those objectives required * Strong ability to communicate complex technical information in a condensed manner to various stakeholders verbally and in writing required * Ability to build and maintain strong cross-functional partnerships at all levels of the organization required * Strong: Learning and teaching other team members and others external to the team preferred * Ability to work, make aligned decisions, plan, and accomplish goals without explicit direction/guidance from leadership required * Experience with system architectures, how software systems interact, and integrate required * Ability to evaluate new technologies to assist senior leadership align it to the HCA Healthcare strategic roadmap required * Strong understanding of SRE practices and implementations required * Expertise in knowledge of Linux and Windows Systems Administration and how to manage through code required * Ability to determine best practices and articulate authoritative direction required * Ability to help establish and grow the SRE principles with the team required * Growth mindset and a willingness to learn new skills, technologies, and frameworks required ## Description What makes HCA Healthcare Information Technology Group (ITG) unique as a technology company is that our solutions ultimately impact the care of patients. Although our skills are needed in many industries, we in ITG apply them specifically to the noble cause of healthcare. We are "Healthcare Inspired." This guiding vision pervades and positively influences every level of our organization. It shapes our mission, defines our values, and brings our leaders and employees together in a shared enthusiasm for their work, setting ITG apart as a uniquely purpose-driven company in the IT industry. As a part of that, we exist to raise the bar, unlock possibilities, and care like family. As a Staff Site Reliability Engineer (SRE), you will provide SRE best practices for mission-critical applications across the enterprise. When these applications fail, you'll have the skills and decision-making capabilities to quickly restore services, investigate the root cause, and develop a plan that mitigates future failures. You will spend time analyzing system performance and identifying ways to enhance the reliability of our environments, from developing dashboards, performing configuration changes, building robust monitoring systems, and learning how to leverage automation to drive efficiencies. You will help drive uptime and reliability across the enterprise. We are on a mission to change the face of the healthcare industry through value-driven products. These products will create innovation for all healthcare users across HCA's nationwide ecosystem. To do this, we are building both curious and quick teams to adapt to new technologies. Major Responsibilities: * Practices and adheres to the "Code of Conduct" philosophy and "Mission and Value Statement." * Promote a collaborative team environment and work closely with colleagues to achieve business objectives. * Collaborate with stakeholders (e.g., business stakeholders, product owners, project managers, and end users) to understand functional and non-functional requirements. * Lead Investigations and solution proposals to development and design problems. * Participate with team members in scope of work estimation and forecasting. * Improve performance of existing software by diagnosing and resolving critical issues. * Prepare technical documentation, including software & architectural design evaluation plans, data flow diagrams, test results, and technical manuals. * Adhere to and influence established development practices and processes. * Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault finding. * Ongoing review of technology, infrastructure, and code to enhance and build resiliency into the applications. * Create sustainable systems and services through automation and uplifts. * Balance feature development & deployments with speed, reliability, and well-defined service-level objectives. * Partner with development teams and vendors of 3rd party applications to improve services through rigorous testing and release procedures. * Build/Develop automations to "self-heal" applications and reduce the toil of manual operational tasks. Pursuit of operational excellence, uptime, and reliability of our applications * Participate, lead, and drive in creating postmortem analysis of why services broke or degraded, including recommendations for long-term fixes. It may require going across multiple teams and organizations within the enterprise. Determine root-cause for all production-level incidents and write corresponding high-quality RCA reports. ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Important Tips That Every Software Developer Should Know](https://www.wearedevelopers.com/magazine/101-7-important-tips-that-every-software-developer-should-know) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [What’s the Difference between a Junior, Mid, and Senior Developer?](https://www.wearedevelopers.com/magazine/238-what-s-the-difference-between-a-junior-mid-and-senior-developer) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)