> Markdown version of [/jobs/ext/2539399-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2539399-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Synapse Health, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $133,600.0 - $183,700.0 - **Contract:** Permanent contract - **Skills:** .NET Framework, Artificial Intelligence, Microsoft Azure, Bash Shell, C Sharp (Programming Language), Cloud Computing, Continuous Integration, Federated Identity Management, Github, Virtual Private Networks (VPN), Python (Programming Language), Network Security, Networking Basics, Performance Tuning, Reliability Engineering, Azure Active Directory, Prometheus, Datadog, Data Logging, Scripting, Grafana, Firewalls (Computer Science), Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Terraform, Microservices - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a7b59b173e3606a2 ## About the Role * 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering * Hands-on experience working in cloud environments (Azure preferred, AWS or GCP acceptable) * Strong experience with Kubernetes in production environments * Experience deploying and managing applications on Kubernetes using Helm * Proficiency with Infrastructure-as-Code tools such as Terraform * Strong scripting skills (Python, Bash, or similar) used to automate and solve infrastructure challenges * Experience with observability, monitoring, and incident response in production environments * Experience building or supporting CI/CD pipelines (GitHub Actions and/or GitLab CI/CD; experience with CI/CD migrations is a plus) * Solid understanding of networking fundamentals, system design, and cloud infrastructure components * Familiarity with Azure Entra ID, app registrations, federated identity credentials, and workload identity * Proven ability to take loosely defined problems and drive them to practical, scalable solutions * Strong communication skills and ability to collaborate across engineering and non-technical stakeholders * Comfort operating in fast-paced environments with evolving priorities * Understanding of PHI handling requirements, access control patterns, and audit controls in a healthcare environment. What Sets You Apart: * Experience working on platform or architectural transformations (e.g., monolith to microservices, functions to containers) * Familiarity with .NET / C# application environments * Deeper networking expertise (firewalls, gateways, and low-level infrastructure components) * Builder's mindset with a focus on automation, scalability, and long-term system design * Ability to balance speed and pragmatism with long-term reliability and maintainability * Curiosity around emerging technologies, including AI/ML as applied to infrastructure and engineering workflows ## Description We're seeking a Sr. Site Reliability Engineer (SRE) to help lead that transformation. This role will play a critical part in evolving our platform from legacy Azure-based services toward a Kubernetes-driven, microservices-oriented environment. As a senior member of the team, you will take ownership of complex, ambiguous infrastructure challenges and drive them through to practical, scalable solutions. You'll partner closely with engineering and data teams to ensure reliability, performance, and scalability are built into our systems from the ground up. This is an ideal opportunity for an engineer who thrives in fast-paced environments, enjoys solving real infrastructure problems, and wants to have a direct impact on the technical direction of a growing healthcare platform. What You Will Do: Platform & Infrastructure Evolution * Contribute to the migration from legacy Azure services and function-based architectures to containerized, microservices-based systems * Help design, build, and scale Kubernetes-based infrastructure and supporting tooling * Partner with engineering teams to ensure new systems are designed for reliability, scalability, and operational efficiency from day one * Drive standardization across infrastructure to reduce silos and enable broader team ownership * Cost Optimization: Monitor cloud usage and spending, identify inefficiencies, and recommend and implement cost optimization strategies * Networking & Connectivity: Design, deploy, and manage secure networking, including public and private endpoints, environment segmentation, site-to-site and point-to-site VPNs, and inter-environment connectivity. Reliability & Observability * Design and maintain highly available, resilient systems in a cloud-native environment * Implement and evolve observability practices including monitoring, alerting, and logging (e.g., Datadog, Prometheus, Grafana) * Define and manage SLIs, SLOs, and SLAs aligned to system performance and user experience * Lead incident response efforts and drive root cause analysis and long-term improvements Automation & Developer Enablement * Build and optimize CI/CD pipelines to support fast, safe, and repeatable deployments * Champion Infrastructure-as-Code practices using tools such as Terraform to eliminate manual processes * Leverage scripting (Python, Bash, or similar) to solve problems, automate workflows, and reduce operational toil Performance, Scalability & Strategy * Drive capacity planning, performance tuning, and infrastructure improvements to support rapid growth * Proactively identify system risks and scalability bottlenecks before they impact customers * Contribute to infrastructure strategy and help shape how the platform evolves as the business scales Knowledge Sharing & Team Enablement * Document systems, processes, and best practices to improve team-wide reliability and reduce single points of failure * Contribute to cross-training efforts as the team moves toward broader ownership and standardization * Share knowledge and elevate the team through mentorship and collaboration Note: These responsibilities reflect the general nature and scope of the role but are not exhaustive. Responsibilities may evolve to meet changing business needs. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)