Site Reliability Engineer II
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Job description
- Define and drive observability strategy for IT system and process health, performance, and reliability, including alerting quality, dashboards, and service health indicators.
- Design and implement software-driven solutions within the IT domain, automating manual processes and eliminating operational complexity and toil.
- Act as a technical leader and force multiplier, helping set priorities and influencing decision-making across the entire IT team - support, systems and networks.
- Take ownership of large, ambiguous initiatives, driving them from concept to delivery while aligning stakeholders across IT, engineering and partner teams.
- Proactively identify systemic risks and reliability gaps in both tools and processes, recommending and leading platform upgrades and architectural improvements before they become incidents.
- Provide technical mentorship, architecture guidance, and high-quality design and code reviews for engineers across IT teams.
- Lead by example in documentation and knowledge sharing, ensuring systems and processes are well-understood and not dependent on individual ownership.
- Participate in and help mature incident response, escalation practices, and post-incident learning across the organization.
Requirements
- Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.
- 5+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles, with demonstrated impact at scale., * Expert-level, methodical troubleshooting across the entire stack, from application to kernel to network.
- Strong command-line proficiency and deep expertise in Linux systems and operating system fundamentals.
- Advanced understanding of networking concepts including load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication., * Experience working across multiple languages (e.g., Python, Go, Bash, and familiarity troubleshooting application stacks such as React or similar). Strong proficiency in at least one.
- Strong track record of automating repetitive and complex operational work to reduce toil and increase reliability.
- Ability to design and build internal tools (Python or Go) that standardize and scale IT practices.
- Comfortable operating in an agile environment, with disciplined testing and quality practices., * Experience with cloud platforms (AWS preferred, GCP/Azure acceptable), particularly managed services and production-grade architectures.
- Expertise in Kubernetes and container orchestration (EKS, Helm), including lifecycle management and operational best practices.
- Proven experience designing and implementing observability systems, including metrics, logging, tracing, dashboards, and alerting.
- Deep understanding of container technologies, security scanning, secrets management, dynamic configuration, and microservices architectures.
- Familiarity with service meshes and advanced traffic management concepts., * Experience designing and maintaining company-wide IaC codebases using tools such as Terraform, Pulumi, CloudFormation, or Ansible.
- Ability to think holistically about infrastructure design, cost, reliability, security, and long-term maintainability.
About the company
It is rare to have a company that both deeply impacts its customers and is able to provide its services across a massive population. At Blink, we have a huge impact on people when they are most vulnerable: at the intersection of their healthcare and finances. We are also the fastest growing healthcare company in the country and are driving that impact across millions of new patients every year. Our business model not only helps people, but drives economics that allow us to build a generational company. We are a relentlessly learning, constantly curious, and aggressively collaborative cross-functional team dedicated to inventing new ways to improve the lives of our customers.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…