Manager, Web and Mobile Site Reliability Engineering
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
The Manager, Web and Mobile Site Reliability Engineering (SRE) leads the engineering team responsible for ensuring maximum uptime, high availability, performance, and resilience for enterprise web applications, mobile app backends, and public API endpoints. This role defines reliability standards, oversees 24/7 incident response, manages edge infrastructure and bot mitigation, and drives automated deployment and observability pipelines., * Team Leadership & SRE Operations: Lead and develop a high-performing team of SRE and DevOps engineers supporting 24/7 high-volume web and mobile systems. Manage on-call rotations, incident command protocols, and operational readiness.
- Reliability & Observability Governance: Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. Architect end-to-end monitoring, tracing, and alerting strategies using tools like Datadog, Dynatrace, or Grafana.
- Incident Management & Remediation: Lead major incident response efforts, drive blameless post-mortems, and collaborate with engineering teams to prioritize root-cause fixes and architectural resiliency improvements.
- Traffic, Edge & Security Management: Partner with IT Security (PCL IT Security) and CDN providers (Akamai) to implement bot mitigation strategies, DDoS defense, WAF rules, and edge caching for key APIs and digital endpoints.
- Administrative: Perform all other administrative and organizational duties as required (time keeping, training, travel, collaboration and correspondence, etc.)
Knowledge & Skills:
- Scope: Direct management of SRE and DevOps engineers. Operational oversight for consumer-facing web platforms, mobile backend APIs, edge routing networks, and cloud deployment pipelines.
- Problem Solving: Rapidly diagnoses and mitigates complex system outages, performance bottlenecks, traffic anomalies, bot campaigns, and infrastructure failures in high-volume production environments.Resolves highly complex, enterprise-scale operational challenges that impact guest operations, maritime services, revenue-generating systems, regulatory requirements, and technology service availability. Anticipates emerging operational risks, evaluates competing business priorities, establishes governance frameworks, and makes decisions where significant operational, financial, service, and reputational consequences may exist. Develops innovative solutions to improve enterprise resilience, scalability, and operational effectiveness.
- Impact: Directly ensures continuous operational availability, system security, optimal site performance, and guest trust across web and mobile touchpoints.
- Leadership: The role requires strong leadership skills. Requires strong incident command leadership, strategic operational decision-making, calm under pressure, and collaborative mentorship.
Requirements
- Knowledge: In-depth understanding of Site Reliability Engineering practices, cloud platforms (AWS/Azure), containerization (Kubernetes, Docker), Akamai/CDN edge routing, bot detection, and CI/CD pipelines (GitLab).
- Skills: Production incident management, automated infrastructure management (Terraform), performance tuning, distributed tracing, metrics-driven SLI/SLO establishment.
- Abilities: Ability to lead teams during critical production outages, drive cross-functional engineering accountability for reliability, and automate operational workflows., * Bachelor’s degree in Computer Science, Computer Engineering, System Administration, or equivalent experience.
- 6+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
- 2+ years of leadership or direct engineering management experience.
Travel: Less than 25% with shoreside travel likely
Work Conditions: Work primarily in a climate-controlled environment with minimal safety/health hazard potential.
Physical Demands: Remain in a stationary position at a desk and/or computer for extended periods of time; reasonable accommodations will be offered.
Benefits & conditions
**This position is classified as “hybrid.” As an in-office role, it requires employees to work from a designated Princess location Mondays through Thursdays. On Fridays you can work from home.
Princess provides comprehensive and innovative benefits to meet your needs, including:
What You Can Expect
- Cruise and Travel Privileges for You and Your Family
- Health Benefits
- 401(k)
- Employee Stock Purchase Plan
- Training & Professional Development
- Tuition & Professional Certification Reimbursement
- Rewards & Incentives
Our Culture… Stronger Together
Our highest responsibility and top priority is compliance, environmental protection and the health, safety and well-being of our guests, the people in the communities we touch and serve, and our shipboard and shoreside employees. Please visit our site to learn more about our Culture Essentials, Corporate Vision Statement and our Core Values at: princess.com/en-us/company-information, * Aida
- HAP Alaska-Yukon
- Carnival Corporation
- Holland America Line
- Carnival Cruise Line
- Carnival UK
- Costa
- Princess
- Seabourn
About the company
One of the best-known names in cruising, Princess is the world’s leading international premium cruise line and tour company, carrying millions of guests each year to hundreds of destinations around the globe. We give our guests the Medallion Class experience others simply can’t. The Love Boat promises something for everyone.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Best Companies in the Netherlands: Top 25 Companies in 2023
Find a Developer Job: 12 Best Job Sites For Developers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again