Site Reliability Engineer

The Cloudbeds
United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$120,000.0 - $150,000.0
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Content Delivery Networks Databases Continuous Integration DevOps Middleware Github Monitoring of Systems PostgreSQL Machine Learning Memcached
+17 more
MySQL Nginx Performance Tuning Redis Prometheus Web Applications Datadog Load Balancing Grafana Kubernetes Helm Charts Amazon Virtual Private Cloud (VPC) Kubernetes Information Technology Cloudwatch Api Gateway Amazon Simple Queue Service (SQS) Terraform

Job description

As a Site Reliability Engineer, you’ll be the guardian of our platform’s reliability and performance, ensuring millions of hospitality transactions flow seamlessly across the globe. You’ll architect and implement scalable AWS cloud solutions that keep the most ambitious hotels running 24/7, while fostering a culture of automation, resilience, and continuous improvement across our engineering teams.

Our SRE Team:

We’re a bottom-up, collaborative team that thrives on healthy debate and shared ownership of our infrastructure. You’ll have endless opportunities to influence architecture decisions while working with cutting-edge cloud technologies at scale. We believe the best solutions come from engineers who are empowered to innovate, experiment, and challenge the status quo.

What You Bring to the Team:

  • Design and implement a reliable and scalable AWS architecture to meet the needs of the organization.
  • Maintain and support highly loaded Kubernetes (EKS) clusters and infrastructure-related components.
  • Support the CICD process with ArgoCD and GitOps.
  • Automate the platform deployments with Terraform infrastructure-as-code.
  • Develop and continuously improve product Observability and Monitoring systems based on the Grafana, Prometheus, DataDog, and Cloudwatch.
  • Respond and participate with Incident Management and Root Cause Analysis, ensuring minimal impact on services.
  • Optimize system performance and troubleshoot issues as they arise.
  • Collaborate with development teams to establish monitoring best practices and ensure systems meet reliability targets.
  • Collaborate with security teams to implement and maintain security best practices.
  • Infrastructure support rotation providing guidance to other engineering teams., * Overall 10 Best Places to Work HotelTechAwards (2026)
  • Overall Top 10 Hotelier’s Choice HotelTechAwards (2026)
  • Finalist - Property Management Systems (PMS) & Channel Managers HotelTechAwards (2026)
  • Most Loved Workplace® Certified (2024)
  • Deloitte Technology Fast 500 (2024)

Discover our Benefits:

  • Remote First, Remote Always
  • PTO in accordance with local labor requirements
  • Fully Paid Parental Leave
  • Home office stipend based on country of residency
  • Professional development courses in Cloudbeds University
  • Access to professional development, including manager training, upskilling, and knowledge transfer, To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Cloudbeds. Staffing, recruiting agencies, and individuals being represented by an agency are not authorized to use this site or to submit applications, and any such submissions will be considered unsolicited. Cloudbeds does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Cloudbeds employees, or any other company location. Cloudbeds is not responsible for any fees related to unsolicited resumes/applications.

Requirements

  • 5+ years of experience as a DevOps or SRE working within the AWS ecosystem.
  • 5+ years of experience with Kubernetes (EKS) and Helm charts.
  • Experience with designing, building, and supporting CI/CD pipelines with ArgoCD and GitHub actions.
  • Experience with infrastructure-as-code methodologies with Terraform.
  • Experience with Observability and Monitoring with Grafana, Prometheus, DataDog, and Cloudwatch.
  • Experience with Incident Management, full stack troubleshooting, performance analysis and root cause analysis (RCA).
  • Experience with Web application systems such as Nginx, Ingress controllers, load balancing and Content Delivery Networks.
  • Experience with Databases (MySQL, PostgreSQL, Aurora) and Middleware technologies (Redis, Memcached and SQS)
  • Good networking skills with VPC, Security Groups and Network ACLs.
  • Ability to work remotely and manage your own time in a global team.
  • Good written and verbal communication in English.
  • Bachelor’s degree in Computer Science or equivalent experience.

Bonus Skills to Stand Out:

  • Advanced experience with Database Administration (Aurora, MySQL, PostgreSQL).
  • Experience working in a PCI-compliant environment.
  • Experience working with Kong API Gateway.

Benefits & conditions

Compensation: Depending on your skills and experience, you can expect your annual compensation to be between $120,000 - $150,000

About the company

Behind Cloudbeds’ revolutionary technology is a team redefining what’s possible in hospitality. We’re 650+ team members across 40+ countries, bringing together elite engineers, AI architects, world-class designers, hoteliers, and hospitality veterans to solve challenges others haven’t dared to tackle. Our diverse team speaks 30+ languages, but we all share one language: a passion for innovation and travel. From pioneering breakthroughs in machine learning to revolutionizing how hotels operate, we’re not just watching the future of hospitality unfold - we’re coding it, designing it, writing it, and shipping it. If you’re ready to work alongside some of the brightest minds in tech who are obsessed with using AI to transform a trillion-dollar industry, this is your chance to be part of something extraordinary.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

Videos

See all

Related articles

See all