Manager, Site Reliability Engineer

Greenhouse Software, Inc.
New York, NY, United States
6 days ago
Apply on job-boards.greenhouse.io
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$150,000.0 - $220,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Continuous Integration DevOps Disaster Recovery Distributed Systems Reliability Engineering Ansible Software Engineering Datadog Reliability of Systems
+6 more
Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Cloudwatch Terraform

Job description

As an engineering organization, we pride ourselves on engineering as a creative activity. Engineering managers enable engineers to do their best work by maintaining a culture and environment where engineers can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems highly available for customers, while partnering closely with Platform, Engineering, Security, Compliance, and Product teams to improve reliability, observability, incident response, and operational maturity. This is an opportunity for a hands-on technical leader who can coach engineers, improve production operations, and help Forge build and run secure, scalable, and highly reliable products., * Manage Forge’s Site Reliability Engineering team responsible for keeping Forge systems highly available for customers.

  • Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning.
  • Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics.
  • Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation.
  • Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards.
  • Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
  • Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability.
  • Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
  • Partner with Security, Compliance, and Risk partners to ensure reliability and infrastructure practices meet the needs of a regulated business.

Requirements

  • 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
  • 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
  • Bachelor’s degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
  • Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
  • Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support.
  • Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
  • Strong technical judgment, communication skills, and ability to influence across engineering and non-engineering stakeholders., * Experience in FinTech, financial services, or another regulated industry.
  • Experience with AWS and/or Azure cloud platforms.
  • Familiarity with Kubernetes, container platforms, infrastructure-as-code, Terraform, Ansible, or similar automation tooling.
  • Experience with observability platforms such as Datadog, CloudWatch, or similar tools.
  • Experience improving developer experience through paved-road platforms, standardization, and self-service infrastructure capabilities.
  • Experience supporting growth-stage companies where speed, scale, reliability, and operational discipline must be balanced.

About the company

At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values of being bold, accountable, and humble. We experience the value that our vision brings to the world every day, helping the teams behind the greatest innovations of our generation, from space travel to artificial intelligence, and more.

With liquidity solutions, exclusive data and insights, a custody offering, and a vibrant marketplace, Forge’s goal is to build the best-in-class technology infrastructure to power a global private market that is transparent, accessible, and seamless for companies, their employees, and investors. Through Forge, employees can sell their private shares, employers can reward shareholders with pre-IPO liquidity and individual and institutional investors can participate in private unicorn growth.

Forge’s differentiated global marketplace addresses rising demand among individual and institutional investors for exposure to private company stocks and is building a growing network effect.

Our ability to offer these powerful financial solutions has generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies, teams, and innovators in this way.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on job-boards.greenhouse.io
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all