> Markdown version of [/jobs/ext/1889025-engineer-site-reliability](https://www.wearedevelopers.com/jobs/ext/1889025-engineer-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineer, Site Reliability - **Company:** Royal Caribbean International - **Location:** Miramar, FL, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automation of Tests, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, System Configuration, Distributed Systems, Monitoring of Systems, Python (Programming Language), Log Analysis, Windows PowerShell, Reliability Engineering, Ansible, Backend, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Cloudwatch, Terraform, Splunk, Appdynamics, Dynatrace, Cisco, Pagerduty, Servicenow - **Published:** August 1, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27898557/Engineer-Site-Reliability-Florida-Miramar-Miramar-1889 ## About the Role * Bachelor's degree in Computer Science, Information Technology, or related field; equivalent experience considered. * 3-5 years of hands-on experience in Observability, Site Reliability Engineering, or similar technical disciplines. * Demonstrated experience with observability platforms such as AppDynamics, Splunk, or equivalent APM and log analysis tools. * Skilled in Kubernetes, preferably EKS/AKS, and cloud observability across AWS and Azure. * Proficiency in configuring, troubleshooting, and optimizing complex monitoring systems. * Strong understanding of OpenTelemetry, distributed tracing, and related telemetry standards. * Familiarity with AIOps, incident response workflows, and ITSM integration (ServiceNow). * Scripting and automation experience in Python, Bash, PowerShell, and infrastructure tools (Terraform, Ansible). * Strong problem-solving, analytical, and collaboration skills. * Ability to translate operational insights into actionable improvements. * Excellent communication skills to support cross-functional teams in a dynamic environment. ## Description The Observability Engineer will operate and continually enhance the enterprise observability platform within one of the most complex hospitality and maritime technology ecosystems. This role ensures reliable system visibility across hundreds of applications and mission-critical services by delivering instrumentation, telemetry, robust monitoring practices, and operational insights to enable proactive performance management across cloud and hybrid environments., * Configure, maintain, and optimize observability platforms such as Cisco AppDynamics, Splunk, ThousandEyes, and PagerDuty AIOps across AWS and Azure environments. * Partner with application, cloud, and platform teams to implement standards for metrics, logs, traces, and event handling, ensuring full-stack visibility. * Design and maintain dashboards, health indicators, and alert policies to improve detection accuracy and reduce noise. * Implement OpenTelemetry instrumentation and telemetry pipelines for data quality, enrichment, and multi-backend exporting. * Collaborate with Site Reliability Engineering teams to define and manage SLIs, SLOs, and actionable alerting thresholds for critical services. * Support AI-driven incident detection, event correlation, and workflow automation across ITSM systems such as ServiceNow. * Enable observability for Kubernetes workloads (EKS, AKS) and cloud infrastructure with AWS CloudWatch and Azure Monitor integrations. * Contribute to observability standards, onboarding playbooks, automation scripts, and continuous improvement initiatives. * Assist in troubleshooting performance issues across application tiers, cloud-native workloads, and distributed systems environments. * Participate in platform upgrades, telemetry hygiene activities, and cost-optimization strategies for large-scale observability systems. ## Related Videos - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Crew Management System for Airlines: Plan duties for pilots & flight attendants worldwide](https://www.wearedevelopers.com/videos/1444-crew-management-system-for-airlines-plan-duties-for-pilots-flight-attendants-worldwide) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [The user in the eye of the Cargo1492 storm](https://www.wearedevelopers.com/videos/96-the-user-in-the-eye-of-the-cargo1492-storm) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)