System Engineer - Site Reliability Engineering (SRE)

Marriott International, Inc.
Bethesda, MD, United States
2 days ago
Apply on ejwl.fa.us2.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$96,000.0 - $152,000.0
Working hours
Regular working hours

Tech stack

Microsoft Windows Application Programming Interfaces (APIs) Agile Methodology IBM AIX Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Systems Engineering JIRA Bash Shell Ubuntu (Operating System) CentOS
+74 more
Cloud Computing Cloud Computing Security Cloud Engineering Configuration Management Computer Programming Databases Continuous Integration Couchbase Servers Software Debugging Desktop Computing Linux Network Address Translation Distributed Systems Domain Name System (DNS) VMware ESX Servers Monitoring of Systems Hypertext Transfer Protocols (HTTP) Hypervisor Identity and Access Management IP Routing Subnetting Python (Programming Language) Kernel-Based Virtual Machine Lightweight Directory Access Protocols (LDAP) PostgreSQL Linux System Administration Apache Maven Windows Servers MySQL Networking Basics NoSQL PCI Data Security Standards Performance Tuning Windows PowerShell Systems Development Life Cycle Red Hat Enterprise Linux Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus Security Assertion Markup Language (SAML) Shell Script Software Engineering Vault (Revision Control System) TCP/IP Virtualization Technology VMware VSphere Software Vulnerability Management Transport Layer Security Load Balancing Cloud Platform System Autoscaling System Availability Grafana Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Rate Limiting Git Cloudformation Amazon Relational Database Service Kubernetes Infrastructure Automation Frameworks Information Technology Low Latency Cassandra Vba Programming Language Hashicorp Functional Programming Terraform Docker Elk Stack Jenkins Artifactory Vmware

Job description

The Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission-critical cloud and on-prem services that support millions of Marriot customers globally. This role involves overseeing incident management, driving automation efforts, and working closely with cross-functional teams to ensure alignment between SRE strategy and business objectives. Partners closely with Product Teams, Applications teams, Infrastructure, and the broader Applications and Infrastructure Delivery teams to develop key metrics and KPIs to improve applications stability, availability and performance. The ideal candidate will bring strong communication skills, collaborating with key stakeholders across the company to optimize cloud infrastructure and uphold the highest standards of operational excellence in a dynamic, fast-paced environment, * Provide support for priority incidents as directed by SRE leader

  • Collaborates through the incident with key team members (network, application, etc.) to engage service providers and other stakeholders to identify problem root cause and drive service restoration

  • Provides closure to incidents or initiates a process-based, hand-off to next shift

  • Works with support vendors and providers to ensure proper global coverage, phone support, parts replacement and smart hands

  • Work to create and mature incident response processes

  • Provides oversight to service providers or less experienced engineers

  • Drive the objectives associated with Problem Management; such as customer communication and Root Cause Analysis reports

  • Trains and/or mentors other team members, and peers as appropriate

  • Identifies opportunities to enhance the service delivery, operations and continual service improvement processes

  • Identifies solutions that may contribute to greater stability and reliability of property infrastructure

  • Develop implementation plans, test plans, and timelines for projects and tasks

Delivering Technology

  • Create and enhance administrative, operational and technical policies and procedures, adopting best practice guidelines, standards and procedures for employees, contractors and vendor engagements

  • Maintains a proper balance between business and operational risk

  • Establishes cadence of communication with other organizations to stay abreast of deployment and production support activities

  • Works in a concerted effort with application development and engineering teams to resolve complex issues

  • Provides oversight, collaboration, provisioning, management and maintenance of technology products and service alternatives that improve the production services environment

  • Responsible for the establishment and continuous development of monitoring and alerting for all production environments

  • Contributes to continuous improvement of internal processes

  • Attends training to ensure skillset and tools support the production environments and deliver on project commitments

  • Performs quantitative and qualitative analyses for operational availability to promote a zero-defect environment

  • Facilitates achievement of expected deliverables and obligations of Services Providers

  • Assists operational teams in system updates & upgrades

  • Provides consultation for routine systems development

  • Ensures early warning to the business stakeholder executives regarding degraded or missed service levels

Service Provider Management

  • Actively coordinates with IT service providers and vendors to bring incidents to resolution

  • Manages Service Providers with a focus on continuous service improvement and service restoration

  • Monitors, manages and leads Service Provider outcomes required to ensure operational availability and a zero-defect production environment

  • Consults with internal Service Management & external Service Providers on performance, business reporting, analytics metrics and business value dashboards

Maintaining Goals

  • Submits reports in a timely manner, ensuring delivery deadlines are met.

  • Promotes the documenting of project progress accurately.

  • Provides input and assistance to other teams regarding projects.

Managing Work, Projects, and Policies

  • Manages and implements work and projects as assigned.

  • Generates and provides accurate and timely results in the form of reports, presentations, etc.

  • Analyzes information and evaluates results to choose the best solution and solve problems.

  • Provides timely, accurate, and detailed status reports as requested.

Demonstrating and Applying Discipline Knowledge

  • Provides technical expertise and support to people inside and outside of the department.

  • Demonstrates knowledge of job-relevant issues, products, systems, and processes.

  • Demonstrates knowledge of function-specific procedures.

  • Keeps up-to-date technically and applies new knowledge to job.

  • Uses computers and computer systems (including hardware and software) to enter data and/ or process information.

Delivering on the Needs of Key Stakeholders

  • Understands and meets the needs of key stakeholders.

  • Develops specific goals and plans to prioritize, organize, and accomplish work.

  • Determines priorities, schedules, plans and necessary resources to ensure completion of any projects on schedule.

  • Collaborates with internal partners and stakeholders to support business/initiative strategies

  • Communicates concepts in a clear and persuasive manner that is easy to understand.

  • Generates and provides accurate and prompt results in the form of reports, presentations, etc.

Requirements

  • Undergraduate degree in an engineering or computer science discipline and/or equivalent experience/certification

  • 5+ years of experience as a Site Reliability Engineer (SRE), building and managing highly available and mission critical systems

  • Expertise in AWS services including designing highly available, multi-AZ and multi-region architectures including:

  • Compute: EC2, Auto Scaling, Lambda

  • Containers: EKS (Mandatory), ECS (good to have)

  • Networking: VPC, subnets, route tables, NAT gateways, Transit Gateway

  • Security: IAM roles/Policies, KMS, Secret manager

  • Storage and Databases: S3, EBS, EFS, RDS, DocumentDB.

  • Deep understanding of SRE practices such as Service Level Objectives, Error Budgets, Toil Management, Observability & Monitoring, Blameless Postmortems, Incident Response Process, Capacity Planning

  • Strong understanding of Cloud Security best practices and responsibility model.

  • Experience driving cloud cost optimization initiatives (rightsizing, reserved instances, autoscaling strategies, cost observability)

  • Proven automation and programming experience in one or more of the following languages: Python, Bash, PowerShell

  • Strong working knowledge of modern, continuous development techniques and pipelines (Agile, Kanban, Jira, CI/CD, Helm, Harness, Jenkins, Git, Artifactory, Vault)

  • Production level expertise with containerization orchestration engines such as Kubernetes (EKS, AKS, ACK)

  • Hands-on experience with service mesh technologies to enable secure and resilient service communication, including mTLS, traffic shaping, and policy enforcement.

  • Strong experience troubleshooting API-related issues in distributed systems, including latency, authentication/authorization failures, rate limiting, and upstream/downstream dependency failures.

  • Ability to analyze API traffic and debug issues using logs, traces, and metrics.

  • Experience with Infrastructure as Code (Iac) tools like Terraform and CloudFormation.

  • Experience with configuration management and automation tools such as Ansible.

  • Deep expertise and hands-on experience with Linux administration (RHEL, Ubuntu, CentOS, AWS Linux)

  • Solid understanding of Virtualization Technologies (VMware vSphere, KVM etc)

  • Strong understanding of networking fundamentals such as Load Balancing, Firewalls, Security Groups, NACLs, TCP/IP, DNS, HTTP/HTTPS, SSL/TLS etc

  • Deep understanding and/or experience with Cloud Native, Relational and NoSQL databases like RDS, MySQL, PostgreSQL, Cassandra or Couchbase

  • Strong experience designing and implementing end-to-end observability solutions across metrics, logs, and traces using tools like Prometheus, Grafana, ELK Stack, and OpenTelemetry.

  • Proven ability to define SLIs/SLOs, build actionable alerting systems, and leverage telemetry data for incident response, root cause analysis, and performance optimization.

  • Experience with deploying, monitoring, and troubleshooting large-scale, distributed applications in cloud environments such as AWS

  • Experience in vulnerability management, OS hardening, patching, security compliance of infrastructure, applications and databases

  • Experience in implementing OS and cloud hardening guidelines and perform regular vulnerability remediation.

  • Familiarity with security frameworks such as ISO27001, SOCII, PCI-DSS, and/or HIPAA

  • 5+ years progressive technology experience

  • 3+ years’ experience in the operational support of critical solutions in large scale environments and organizations with specific experience in

  • Windows Servers

  • RHEL

  • VMWare (vCenter and ESXi Hypervisor)

  • Ability to work outside of normal business hours and days in a 24x7x365 environment.

  • Some travel required

Preferred:

  • CI/CD pipeline technologies such as Git, Docker Trusted Registry, Artifactory, Hashicorp Vault, Maven, etc.

  • Security Protocols like SSL, SAML, LDAP etc.

  • AIX Administration

  • Automation and Scripting languages (Ansible, PowerShell, VB, ShellScript)

  • Strong organizational, written and verbal communication skills

  • ITIL v4 certification

  • Windows and Redhat Certification

  • Experience delivering technology solutions in a fast-paced, deadline driven enterprise environment

  • Experience learning and applying new technologies to solve business needs

  • Excellent understanding of change management, testing requirements, techniques, and tools to ensure high availability of systems

  • Experience in researching emerging technologies and trends, standards, and products

Benefits & conditions

Full-time positions also offer coverage for medical, dental, vision, health care flexible spending account, dependent care flexible spending account, life insurance, disability insurance, accident insurance, adoption expense reimbursements, paid parental leave and educational assistance. Washington Applicants Only: Employees will accrue paid sick leave, 0.077 PTO balance for every hour worked and be eligible to receive a minimum of 9 holidays annually. Marriott Headquarters (HQ) supports a hybrid work environment that helps associates Be Connected. HQ based roles are generally hybrid and require regular in-office presence at the assigned location, including for candidates who live within commuting distance of HQ for a role listed as remote. Remote roles will be clearly identified in the job posting.

About the company

Marriott International is the world’s largest hotel company, with more brands, more hotels and more opportunities for associates to grow and succeed. Be where you can do your best work, begin your purpose, belong to an amazing global team, and become the best version of you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on ejwl.fa.us2.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all