Lead Site Reliability Engineer

McGraw-Hill
United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$124,000.0 - $155,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Application Performance Management Systems Engineering Cloud Engineering Code Review Cyber Security DevOps Github Identity and Access Management Uptime Scrum Methodology Reliability Engineering
+8 more
Software Engineering Datadog Enterprise Software Applications Grafana Performance Monitor Cloudwatch Terraform New Relic (SaaS)

Job description

  • Lead a 6 member SRE team supporting production infrastructure and services
  • Manage backlog, sprint planning, and team velocity
  • Own reliability, uptime, security, cost, and performance of services
  • Define and monitor SLOs for application workloads
  • Plan on-call rotations and work to reduce alert fatigue
  • Forecast seasonal growth and capacity planning
  • Mentor engineers and foster professional growth
  • Report status and issues to leadership monthly
  • Partner with development teams
  • Collaborate with CyberSecurity on risk mitigation
  • Collaborate with FinOps on cost reduction
  • Design and troubleshoot highly-distributed, cloud-based production systems
  • Maintain infrastructure-as-code and monitoring-as-code practices
  • Improve system resiliency through failure injection and chaos testing
  • Participate in on-call rotation and resolve operational issues
  • Optimize existing systems for performance and cost
  • Ensure telemetry provides visibility to application performance
  • Support agile development practices and code reviews

Requirements

  • 5+ years of experience in SRE, DevOps, or Software Engineering roles supporting enterprise applications.
  • Strong problem-solving, triage, and root cause analysis skills with a systems engineering mindset
  • Deep expertise in the AWS ecosystem, with hands-on experience across core services including primarily EKS, RDS, EKS, IAM, CloudWatch, and networking configurations.
  • Expertise with Terraform for managing and automating scalable cloud infrastructure
  • Skilled in CI/CD pipelines (e.g., GitHub Actions) and managing end-to-end software delivery lifecycles.
  • Strong familiarity with telemetry and observability tools (e.g., New Relic, Datadog), including querying logs and metrics for performance monitoring.

Benefits & conditions

The work you do at McGraw Hill will be work that matters. We are collectively designing content that will build the future of education. Play your part and experience a sense of fulfillment that will inspire you to even greater heights.

The pay range for this position is between $124,000 - $155,000 annually. However, base pay offered may vary depending on job-related knowledge, skills, experience, and location. An annual bonus plan may be provided as part of the compensation package, in addition to a full range of medical and/or other benefits, depending on the position offered. Click here to learn more about our benefit offerings.

McGraw Hill recruiters always use a “@mheducation.com” or “@careers.mheducation.com” email addresses and/or from our Applicant Tracking System, iCIMS. Any variation of this email domain should be considered suspicious. Additionally, McGraw Hill recruiters and authorized representatives will never request sensitive information in email.

50935

About the company

McGraw Hill, a leading provider of digital educational resources and content, is seeking a Lead Site Reliability Engineer to lead a team of 6 Engineers for our Digital Platform Group in supporting our K-12 learning platforms. These platforms serve millions of students and educators nationwide, and you’ll play a key role in ensuring their reliability, scalability, and performance. Working closely with engineering and product teams, you’ll leverage your expertise in AWS, Terraform, and observability tools to drive automation, enhance resiliency, and maintain the health of our cloud-based infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:22 min

Leveraging unique cultural backgrounds in engineering design

Ixchel Ruiz · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

Videos

See all

Related articles

See all