Site Reliability Engineering Manager in United

Energy Jobline
United, WV, United States
8 days ago
Apply on www.energyjobline.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Agile Methodology Amazon Web Services Microsoft Azure Bash Shell C Sharp (Programming Language) Computer Programming DevOps Python (Programming Language) MySQL NoSQL Windows PowerShell
+8 more
Reliability Engineering SQL Databases Datadog Scripting Kubernetes Deployment Automation Docker Pagerduty

Job description

You will also collaborate closely with SRE leadership in India as part of our global follow-the-sun support model.

Requirements

  • 5-8+ years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering. \n

  • 1-2+ years of experience leading, mentoring, or managing engineers. \n

  • Experience operating in a player-coach leadership model. \n

  • Strong hands-on experience with production incident management and escalation. \n

  • Experience with Datadog or similar observability platforms. \n

  • Production experience with Kubernetes and Docker. \n

  • Strong scripting/programming skills using Python, Bash, PowerShell, Java, or C#. \n

  • Experience with Helm, CI/CD pipelines, and deployment automation. \n

  • Working knowledge of ITIL processes and Agile methodologies. \n

  • Experience with SQL, MySQL, or NoSQL databases. \n

  • Strong communication and stakeholder-management skills. \n

  • Willingness to participate in PagerDuty/on-call escalation and a global follow-the-sun operating model., * Experience with AWS, Azure, or GCP. \n

  • Experience building or scaling SRE teams and on-call programs. \n

  • Experience defining and managing SLIs, SLOs, SLAs, and error budgets. \n

  • Prior experience in healthcare, fintech, or another regulated industry. \n

  • Knowledge of security and compliance frameworks used in regulated environments. \n

Benefits & conditions

  • Improve system reliability, resilience, and observability using Datadog or similar tools. \n

  • Drive automation, self-healing capabilities, and runbook maturity. \n

  • Partner with Development, DevOps, DevSecOps, and Engineering teams to embed reliability into the SDLC. \n

  • Contribute hands-on to technical reviews, tooling, scripting, and automation. \n

\n

Global Collaboration

\n \n

  • Partner closely with SRE leadership in India to support follow-the-sun operations. \n

  • Represent the US SRE team in cross-functional planning and operational reviews. \n

  • Communicate effectively with technical and non-technical stakeholders. \n

\n

Documentation & Compliance

\n \n

  • Maintain documentation for incidents, postmortems, runbooks, and operational procedures. \n

  • Support adherence to regulated-environment standards including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST. \n

\n

\n

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.energyjobline.com
Prepare application

Good distractions

Loading talks and stories from around this role…