Manager, Web and Mobile Site Reliability Engineering

Holland, Inc
Fort Lauderdale, FL, United States
7 days ago
Apply on eicl.fa.em5.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Microsoft Azure Cyber Security Computer Engineering DevOps Performance Tuning Reliability Engineering Site Reliability Engineering Practices Akamai Web Platforms Datadog
+10 more
Delivery Pipeline Grafana Gitlab Containerization Kubernetes Information Technology Terraform Ddos Dynatrace Docker

Job description

The Manager, Web and Mobile Site Reliability Engineering (SRE) leads the engineering team responsible for ensuring maximum uptime, high availability, performance, and resilience for enterprise web applications, mobile app backends, and public API endpoints. This role defines reliability standards, oversees 24/7 incident response, manages edge infrastructure and bot mitigation, and drives automated deployment and observability pipelines., * Team Leadership & SRE Operations: Lead and develop a high-performing team of SRE and DevOps engineers supporting 24/7 high-volume web and mobile systems. Manage on-call rotations, incident command protocols, and operational readiness.

  • Reliability & Observability Governance: Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. Architect end-to-end monitoring, tracing, and alerting strategies using tools like Datadog, Dynatrace, or Grafana.
  • Incident Management & Remediation: Lead major incident response efforts, drive blameless post-mortems, and collaborate with engineering teams to prioritize root-cause fixes and architectural resiliency improvements.
  • Traffic, Edge & Security Management: Partner with IT Security (PCL IT Security) and CDN providers (Akamai) to implement bot mitigation strategies, DDoS defense, WAF rules, and edge caching for key APIs and digital endpoints.
  • Administrative: Perform all other administrative and organizational duties as required (time keeping, training, travel, collaboration and correspondence, etc.)

Knowledge & Skills:

  • Scope: Direct management of SRE and DevOps engineers. Operational oversight for consumer-facing web platforms, mobile backend APIs, edge routing networks, and cloud deployment pipelines.
  • Problem Solving: Rapidly diagnoses and mitigates complex system outages, performance bottlenecks, traffic anomalies, bot campaigns, and infrastructure failures in high-volume production environments.Resolves highly complex, enterprise-scale operational challenges that impact guest operations, maritime services, revenue-generating systems, regulatory requirements, and technology service availability. Anticipates emerging operational risks, evaluates competing business priorities, establishes governance frameworks, and makes decisions where significant operational, financial, service, and reputational consequences may exist. Develops innovative solutions to improve enterprise resilience, scalability, and operational effectiveness.
  • Impact: Directly ensures continuous operational availability, system security, optimal site performance, and guest trust across web and mobile touchpoints.
  • Leadership: The role requires strong leadership skills. Requires strong incident command leadership, strategic operational decision-making, calm under pressure, and collaborative mentorship.

Requirements

  • Knowledge: In-depth understanding of Site Reliability Engineering practices, cloud platforms (AWS/Azure), containerization (Kubernetes, Docker), Akamai/CDN edge routing, bot detection, and CI/CD pipelines (GitLab).
  • Skills: Production incident management, automated infrastructure management (Terraform), performance tuning, distributed tracing, metrics-driven SLI/SLO establishment.
  • Abilities: Ability to lead teams during critical production outages, drive cross-functional engineering accountability for reliability, and automate operational workflows., * Bachelor’s degree in Computer Science, Computer Engineering, System Administration, or equivalent experience.
  • 6+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • 2+ years of leadership or direct engineering management experience.

Travel: Less than 25% with shoreside travel likely

Work Conditions: Work primarily in a climate-controlled environment with minimal safety/health hazard potential.

Physical Demands: Remain in a stationary position at a desk and/or computer for extended periods of time; reasonable accommodations will be offered.

Benefits & conditions

**This position is classified as “hybrid.” As an in-office role, it requires employees to work from a designated Princess location Mondays through Thursdays. On Fridays you can work from home.

Princess provides comprehensive and innovative benefits to meet your needs, including:

What You Can Expect

  • Cruise and Travel Privileges for You and Your Family
  • Health Benefits
  • 401(k)
  • Employee Stock Purchase Plan
  • Training & Professional Development
  • Tuition & Professional Certification Reimbursement
  • Rewards & Incentives

Our Culture… Stronger Together

Our highest responsibility and top priority is compliance, environmental protection and the health, safety and well-being of our guests, the people in the communities we touch and serve, and our shipboard and shoreside employees. Please visit our site to learn more about our Culture Essentials, Corporate Vision Statement and our Core Values at: princess.com/en-us/company-information, * Aida

  • HAP Alaska-Yukon
  • Carnival Corporation
  • Holland America Line
  • Carnival Cruise Line
  • Carnival UK
  • Costa
  • Princess
  • Seabourn

About the company

One of the best-known names in cruising, Princess is the world’s leading international premium cruise line and tour company, carrying millions of guests each year to hundreds of destinations around the globe. We give our guests the Medallion Class experience others simply can’t. The Love Boat promises something for everyone.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eicl.fa.em5.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:51 min

Rising DDoS attacks and evaluating CDN mitigation strategies

Chris Heilmann +2 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

3:11 min

Surviving sudden scale events and malicious traffic

Justin Kitagawa · Coffee With Developers

2:48 min

Daily responsibilities and alignment practices for technical engineering leadership

Edoardo Dusi · LIVE

Videos

See all

Related articles

See all