> Markdown version of [/jobs/ext/2694812-system-engineer-site-reliability-engineering-sre](https://www.wearedevelopers.com/jobs/ext/2694812-system-engineer-site-reliability-engineering-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # System Engineer - Site Reliability Engineering (SRE) - **Company:** Marriott International, Inc. - **Location:** Bethesda, MD, United States (Remote available) - **Experience:** Experienced - **Salary:** $96,000.0 - $152,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Application Programming Interfaces (APIs), Agile Methodology, IBM AIX, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Systems Engineering, JIRA, Bash Shell, Ubuntu (Operating System), CentOS, Cloud Computing, Cloud Computing Security, Cloud Engineering, Configuration Management, Computer Programming, Databases, Continuous Integration, Couchbase Servers, Software Debugging, Desktop Computing, Linux, Network Address Translation, Distributed Systems, Domain Name System (DNS), VMware ESX Servers, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Hypervisor, Identity and Access Management, IP Routing, Subnetting, Python (Programming Language), Kernel-Based Virtual Machine, Lightweight Directory Access Protocols (LDAP), PostgreSQL, Linux System Administration, Apache Maven, Windows Servers, MySQL, Networking Basics, NoSQL, PCI Data Security Standards, Performance Tuning, Windows PowerShell, Systems Development Life Cycle, Red Hat Enterprise Linux, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Security Assertion Markup Language (SAML), Shell Script, Software Engineering, Vault (Revision Control System), TCP/IP, Virtualization Technology, VMware VSphere, Software Vulnerability Management, Transport Layer Security, Load Balancing, Cloud Platform System, Autoscaling, System Availability, Grafana, Firewalls (Computer Science), Amazon Virtual Private Cloud (VPC), Rate Limiting, Git, Cloudformation, Amazon Relational Database Service, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Cassandra, Vba Programming Language, Hashicorp, Functional Programming, Terraform, Docker, Elk Stack, Jenkins, Artifactory, Vmware - **Published:** September 3, 2026 - **Apply:** https://ejwl.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/MI_CS_1/job/26112266/apply/email ## About the Role * Undergraduate degree in an engineering or computer science discipline and/or equivalent experience/certification * 5+ years of experience as a Site Reliability Engineer (SRE), building and managing highly available and mission critical systems * Expertise in AWS services including designing highly available, multi-AZ and multi-region architectures including: * Compute: EC2, Auto Scaling, Lambda * Containers: EKS (Mandatory), ECS (good to have) * Networking: VPC, subnets, route tables, NAT gateways, Transit Gateway * Security: IAM roles/Policies, KMS, Secret manager * Storage and Databases: S3, EBS, EFS, RDS, DocumentDB. * Deep understanding of SRE practices such as Service Level Objectives, Error Budgets, Toil Management, Observability & Monitoring, Blameless Postmortems, Incident Response Process, Capacity Planning * Strong understanding of Cloud Security best practices and responsibility model. * Experience driving cloud cost optimization initiatives (rightsizing, reserved instances, autoscaling strategies, cost observability) * Proven automation and programming experience in one or more of the following languages: Python, Bash, PowerShell * Strong working knowledge of modern, continuous development techniques and pipelines (Agile, Kanban, Jira, CI/CD, Helm, Harness, Jenkins, Git, Artifactory, Vault) * Production level expertise with containerization orchestration engines such as Kubernetes (EKS, AKS, ACK) * Hands-on experience with service mesh technologies to enable secure and resilient service communication, including mTLS, traffic shaping, and policy enforcement. * Strong experience troubleshooting API-related issues in distributed systems, including latency, authentication/authorization failures, rate limiting, and upstream/downstream dependency failures. * Ability to analyze API traffic and debug issues using logs, traces, and metrics. * Experience with Infrastructure as Code (Iac) tools like Terraform and CloudFormation. * Experience with configuration management and automation tools such as Ansible. * Deep expertise and hands-on experience with Linux administration (RHEL, Ubuntu, CentOS, AWS Linux) * Solid understanding of Virtualization Technologies (VMware vSphere, KVM etc) * Strong understanding of networking fundamentals such as Load Balancing, Firewalls, Security Groups, NACLs, TCP/IP, DNS, HTTP/HTTPS, SSL/TLS etc * Deep understanding and/or experience with Cloud Native, Relational and NoSQL databases like RDS, MySQL, PostgreSQL, Cassandra or Couchbase * Strong experience designing and implementing end-to-end observability solutions across metrics, logs, and traces using tools like Prometheus, Grafana, ELK Stack, and OpenTelemetry. * Proven ability to define SLIs/SLOs, build actionable alerting systems, and leverage telemetry data for incident response, root cause analysis, and performance optimization. * Experience with deploying, monitoring, and troubleshooting large-scale, distributed applications in cloud environments such as AWS * Experience in vulnerability management, OS hardening, patching, security compliance of infrastructure, applications and databases * Experience in implementing OS and cloud hardening guidelines and perform regular vulnerability remediation. * Familiarity with security frameworks such as ISO27001, SOCII, PCI-DSS, and/or HIPAA * 5+ years progressive technology experience * 3+ years' experience in the operational support of critical solutions in large scale environments and organizations with specific experience in * Windows Servers * RHEL * VMWare (vCenter and ESXi Hypervisor) * Ability to work outside of normal business hours and days in a 24x7x365 environment. * Some travel required Preferred: * CI/CD pipeline technologies such as Git, Docker Trusted Registry, Artifactory, Hashicorp Vault, Maven, etc. * Security Protocols like SSL, SAML, LDAP etc. * AIX Administration * Automation and Scripting languages (Ansible, PowerShell, VB, ShellScript) * Strong organizational, written and verbal communication skills * ITIL v4 certification * Windows and Redhat Certification * Experience delivering technology solutions in a fast-paced, deadline driven enterprise environment * Experience learning and applying new technologies to solve business needs * Excellent understanding of change management, testing requirements, techniques, and tools to ensure high availability of systems * Experience in researching emerging technologies and trends, standards, and products ## Description The Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission-critical cloud and on-prem services that support millions of Marriot customers globally. This role involves overseeing incident management, driving automation efforts, and working closely with cross-functional teams to ensure alignment between SRE strategy and business objectives. Partners closely with Product Teams, Applications teams, Infrastructure, and the broader Applications and Infrastructure Delivery teams to develop key metrics and KPIs to improve applications stability, availability and performance. The ideal candidate will bring strong communication skills, collaborating with key stakeholders across the company to optimize cloud infrastructure and uphold the highest standards of operational excellence in a dynamic, fast-paced environment, * Provide support for priority incidents as directed by SRE leader * Collaborates through the incident with key team members (network, application, etc.) to engage service providers and other stakeholders to identify problem root cause and drive service restoration * Provides closure to incidents or initiates a process-based, hand-off to next shift * Works with support vendors and providers to ensure proper global coverage, phone support, parts replacement and smart hands * Work to create and mature incident response processes * Provides oversight to service providers or less experienced engineers * Drive the objectives associated with Problem Management; such as customer communication and Root Cause Analysis reports * Trains and/or mentors other team members, and peers as appropriate * Identifies opportunities to enhance the service delivery, operations and continual service improvement processes * Identifies solutions that may contribute to greater stability and reliability of property infrastructure * Develop implementation plans, test plans, and timelines for projects and tasks Delivering Technology * Create and enhance administrative, operational and technical policies and procedures, adopting best practice guidelines, standards and procedures for employees, contractors and vendor engagements * Maintains a proper balance between business and operational risk * Establishes cadence of communication with other organizations to stay abreast of deployment and production support activities * Works in a concerted effort with application development and engineering teams to resolve complex issues * Provides oversight, collaboration, provisioning, management and maintenance of technology products and service alternatives that improve the production services environment * Responsible for the establishment and continuous development of monitoring and alerting for all production environments * Contributes to continuous improvement of internal processes * Attends training to ensure skillset and tools support the production environments and deliver on project commitments * Performs quantitative and qualitative analyses for operational availability to promote a zero-defect environment * Facilitates achievement of expected deliverables and obligations of Services Providers * Assists operational teams in system updates & upgrades * Provides consultation for routine systems development * Ensures early warning to the business stakeholder executives regarding degraded or missed service levels Service Provider Management * Actively coordinates with IT service providers and vendors to bring incidents to resolution * Manages Service Providers with a focus on continuous service improvement and service restoration * Monitors, manages and leads Service Provider outcomes required to ensure operational availability and a zero-defect production environment * Consults with internal Service Management & external Service Providers on performance, business reporting, analytics metrics and business value dashboards Maintaining Goals * Submits reports in a timely manner, ensuring delivery deadlines are met. * Promotes the documenting of project progress accurately. * Provides input and assistance to other teams regarding projects. Managing Work, Projects, and Policies * Manages and implements work and projects as assigned. * Generates and provides accurate and timely results in the form of reports, presentations, etc. * Analyzes information and evaluates results to choose the best solution and solve problems. * Provides timely, accurate, and detailed status reports as requested. Demonstrating and Applying Discipline Knowledge * Provides technical expertise and support to people inside and outside of the department. * Demonstrates knowledge of job-relevant issues, products, systems, and processes. * Demonstrates knowledge of function-specific procedures. * Keeps up-to-date technically and applies new knowledge to job. * Uses computers and computer systems (including hardware and software) to enter data and/ or process information. Delivering on the Needs of Key Stakeholders * Understands and meets the needs of key stakeholders. * Develops specific goals and plans to prioritize, organize, and accomplish work. * Determines priorities, schedules, plans and necessary resources to ensure completion of any projects on schedule. * Collaborates with internal partners and stakeholders to support business/initiative strategies * Communicates concepts in a clear and persuasive manner that is easy to understand. * Generates and provides accurate and prompt results in the form of reports, presentations, etc. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)