Cloud/Platform Operations Manager

Bain Capital
Boston, MA, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$180,000.0 - $210,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Cloud Computing Identity and Access Management Performance Tuning Prometheus Data Logging Grafana Containerization Kubernetes Cloudwatch

Job description

We are seeking a hands-on, operationally focused Cloud / Platform Operations Manager to lead the day-to-day management of our AWS and Kubernetes environments. This leader will bring deep technical expertise in operating, troubleshooting, and optimizing production Kubernetes platforms while ensuring the reliability, security, and governance of our cloud infrastructure.

This role is responsible for ensuring platform stability, enforcing cloud governance and guardrails, and delivering responsive, high-quality support to internal users. The ideal candidate is comfortable rolling up their sleeves to resolve complex Kubernetes and AWS operational issues while also leading and developing a high-performing operations team.

Unlike traditional platform engineering roles focused on feature development, this position is centered on operations, service delivery, and execution. The team supports a high volume of incoming requests and operational needs, and this leader will prioritize responsiveness, reliability, and consistency over agile delivery.

This role also requires a strong people leader who can manage escalations, support team development, and act as a first line of leadership for operational issues before they reach senior leadership.

Responsibilities

Operational Leadership & Service Delivery

  • Own the day-to-day operations of AWS and Kubernetes environments, ensuring reliability, availability, and performance.
  • Provide technical leadership and hands-on support for complex Kubernetes and cloud infrastructure issues.
  • Lead a ticket-driven support model, prioritizing and ensuring timely fulfillment of internal user requests.
  • Establish and enforce SLAs/SLOs for platform support and operational responsiveness.
  • Drive a culture of execution, accountability, and customer service within the team.

AWS Governance, Guardrails & Policy Enforcement

  • Define, implement, and enforce AWS guardrails, IAM policies, and governance frameworks.
  • Ensure proper account structure, access controls, cost controls, and compliance standards.
  • Partner with security teams to enforce best practices and reduce risk across the platform.
  • Continuously audit and improve cloud policy adherence.

Kubernetes Operations & Stability

  • Own the operational health and lifecycle management of production Kubernetes clusters.
  • Maintain and troubleshoot Kubernetes control planes, worker nodes, networking, storage, ingress, and containerized workloads.
  • Perform hands-on cluster administration, upgrades, patching, capacity planning, and performance tuning.
  • Ensure robust monitoring, logging, alerting, and incident response for Kubernetes environments.
  • Drive operational best practices for cluster reliability, security, and resiliency.
  • Partner with application teams to troubleshoot deployment, scaling, networking, and workload issues.

Incident & Escalation Management

  • Participate in a rotating after-hours on-call schedule to provide operational support for critical production incidents, ensuring timely response, issue resolution, and service continuity.
  • Act as the primary escalation point for AWS and Kubernetes operational issues.
  • Lead technical troubleshooting during production incidents and guide the team through resolution.
  • Handle and resolve escalations effectively before they reach senior leadership.
  • Lead incident response, root cause analysis, and post-incident improvements.
  • Build strong communication channels with stakeholders during incidents.

Team Leadership & Development

  • Lead, coach, and develop a team of cloud/platform engineers with a strong operational mindset.
  • Mentor engineers on Kubernetes operations, AWS best practices, and operational excellence.
  • Set clear expectations around ownership, responsiveness, and execution.
  • Manage performance, provide feedback, and address gaps in delivery or behavior.
  • Foster a culture of accountability, urgency, and continuous improvement.

Cross-Functional Support

  • Partner with engineering, product, and business teams to support platform needs and unblock users.
  • Balance operational workload with longer-term improvements without compromising service delivery.
  • Act as a bridge between users and platform capabilities, ensuring needs are met efficiently.

Requirements

  • 8+ years of experience in cloud infrastructure or platform operations, with at least 2 years in a leadership role.
  • Strong hands-on experience administering and supporting production Kubernetes environments, including cluster operations, troubleshooting, upgrades, networking, storage, ingress, and workload management.
  • Demonstrated ability to independently diagnose and resolve complex Kubernetes platform issues in production.
  • Strong hands-on experience with AWS, including IAM, networking, account structure, and governance controls.
  • Proven experience managing AWS guardrails, policies, and access controls at scale.
  • Experience with Kubernetes observability and operational tooling (such as Prometheus, Grafana, CloudWatch, Fluent Bit, or similar).
  • Deep operational experience running and supporting Kubernetes environments in production.
  • Experience leading ticket-based support or operations teams with high request volume.
  • Strong incident management and escalation handling experience.
  • Demonstrated ability to lead teams focused on execution, reliability, and service delivery while remaining technically engaged.
  • Excellent communication and stakeholder management skills.

Benefits & conditions

Expected Annual Base Salary $180k-$210k

Actual base salary will be determined by a wide range of factors including but not limited to role, function, level, experience, qualifications and geographic location. In addition to a competitive base salary, this position may be eligible for a discretionary annual bonus based upon factors such as individual impact, team and firm performance. Bain Capital offers a competitive benefits package designed to support employees’ health, financial security, family needs, and overall well-being.

About the company

With approximately $225 billion of assets under management, Bain Capital is one of the world’s leading private investment firms. We create lasting impact for our investors, teams, businesses, and the communities in which we live. Over four decades we have strategically grown our platform to focus on Private Equity, Growth & Venture, Capital Solutions, Credit, and Real Assets. Today, our team includes 1,985+ employees in 24 offices on four continents.

We partner differently to help people and companies embrace possibility and realize potential. Founded as a private partnership in 1984, we have fostered a culture of innovation, entrepreneurialism, and agility, empowering our people to define and own their career trajectories. Today, our partnership approach enables us to pursue strategic growth, build enduring relationships with a robust external network, and collaborate across our integrated platform to connect the deep and diverse expertise that unlocks breakthrough insights.

Our people are the heart of our advantage. Colleagues at all levels have a seat at the table as they tackle business challenges with a principal investor mindset. By asking incisive questions, respectfully challenging one another, and remaining intellectually agile, we work together to achieve exceptional outcomes.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

6:51 min

Audience questions on cloud security and operational capacity

Steffen Heilmann · WWC 2021

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all