> Markdown version of [/jobs/ext/3649298-site-reliability-engineer-ii](https://www.wearedevelopers.com/jobs/ext/3649298-site-reliability-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II - **Company:** Blink Health - **Location:** Pittsburgh, PA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Proxy Servers, Microsoft Azure, Bash Shell, Command-Line Interface, Cloud Computing, Computer Networks, Linux, Network Address Translation, Domain Name System (DNS), Monitoring of Systems, Python (Programming Language), Key Management, Reliability Engineering, Ansible, TCP/IP, Data Logging, Pulumi, Load Balancing, ReactJS, Cloudformation, Containerization, Information Technology, Terraform, Vulnerability Analysis, Golang, Microservices - **Published:** October 9, 2026 - **Apply:** https://startup.jobs/senior-site-reliability-engineer-ii-blink-health-10354981 ## About the Role * Bachelor's or Master's degree in Computer Science or equivalent practical experience. * 5+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles, with demonstrated impact at scale., * Expert-level, methodical troubleshooting across the entire stack, from application to kernel to network. * Strong command-line proficiency and deep expertise in Linux systems and operating system fundamentals. * Advanced understanding of networking concepts including load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication., * Experience working across multiple languages (e.g., Python, Go, Bash, and familiarity troubleshooting application stacks such as React or similar). Strong proficiency in at least one. * Strong track record of automating repetitive and complex operational work to reduce toil and increase reliability. * Ability to design and build internal tools (Python or Go) that standardize and scale IT practices. * Comfortable operating in an agile environment, with disciplined testing and quality practices., * Experience with cloud platforms (AWS preferred, GCP/Azure acceptable), particularly managed services and production-grade architectures. * Expertise in Kubernetes and container orchestration (EKS, Helm), including lifecycle management and operational best practices. * Proven experience designing and implementing observability systems, including metrics, logging, tracing, dashboards, and alerting. * Deep understanding of container technologies, security scanning, secrets management, dynamic configuration, and microservices architectures. * Familiarity with service meshes and advanced traffic management concepts., * Experience designing and maintaining company-wide IaC codebases using tools such as Terraform, Pulumi, CloudFormation, or Ansible. * Ability to think holistically about infrastructure design, cost, reliability, security, and long-term maintainability. ## Description * Define and drive observability strategy for IT system and process health, performance, and reliability, including alerting quality, dashboards, and service health indicators. * Design and implement software-driven solutions within the IT domain, automating manual processes and eliminating operational complexity and toil. * Act as a technical leader and force multiplier, helping set priorities and influencing decision-making across the entire IT team - support, systems and networks. * Take ownership of large, ambiguous initiatives, driving them from concept to delivery while aligning stakeholders across IT, engineering and partner teams. * Proactively identify systemic risks and reliability gaps in both tools and processes, recommending and leading platform upgrades and architectural improvements before they become incidents. * Provide technical mentorship, architecture guidance, and high-quality design and code reviews for engineers across IT teams. * Lead by example in documentation and knowledge sharing, ensuring systems and processes are well-understood and not dependent on individual ownership. * Participate in and help mature incident response, escalation practices, and post-incident learning across the organization.