> Markdown version of [/jobs/ext/2717833-platform-operations-manager](https://www.wearedevelopers.com/jobs/ext/2717833-platform-operations-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Operations Manager - **Company:** Huxley Associates - **Location:** Boston, MA, United States - **Experience:** Experienced - **Salary:** $180,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Identity and Access Management, Performance Tuning, System Availability, Kubernetes - **Published:** September 4, 2026 - **Apply:** https://www.huxley.com/en-gb/job/platform-operations-manager/4071560 ## About the Role * 8+ years of cloud infrastructure or platform operations experience, including 2+ years in a leadership role. ## Description This is a hands-on leadership role focused on the operations, stability, governance, and support of AWS and Kubernetes environments. Unlike platform engineering positions centered on building new features, this role is responsible for ensuring production systems remain reliable, secure, compliant, and highly available while leading a team of cloud/platform engineers. The manager will oversee day-to-day cloud operations, Kubernetes administration, incident response, AWS governance, and service delivery. They will serve as the primary escalation point for complex operational issues, drive platform reliability, enforce cloud policies and guardrails, and support a high-volume, ticket-driven environment. The role also includes after-hours on-call responsibilities and close collaboration with engineering, security, and business stakeholders., * Lead day-to-day AWS and Kubernetes platform operations, ensuring high availability, performance, and stability. * Oversee Kubernetes cluster health, monitoring, upgrades, patching, troubleshooting, and performance optimization. * Drive SLA/SLO adherence and continuous improvements in service delivery and operational excellence. Incident Management & Support * Serve as the primary escalation point for production incidents and complex operational issues. * Lead incident response, root cause analysis, and corrective action planning. * Manage a customer-focused, ticket-driven support environment with an emphasis on responsiveness and execution. AWS Governance & Security * Define and enforce AWS governance standards, IAM policies, security controls, and compliance requirements. * Ensure proper cloud account structure, access management, and cost optimization practices. * Partner with security teams to proactively mitigate risk and strengthen platform security. Leadership & Collaboration * Lead, mentor, and develop a team of Cloud/Platform Engineers. * Foster a culture of accountability, ownership, and operational excellence. * Collaborate with engineering, security, and business stakeholders to support platform needs and drive continuous improvement. ## Related Videos - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)