Senior Site Reliability Engineer

Sysco Corporation
United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Airflow Amazon Web Services Bash Shell BigQuery Cloud Computing Cloud Computing Security Cloud Engineering Cloud Storage Databases Continuous Integration Information Engineering Data Infrastructure
+23 more
Data Integrity DevOps Disaster Recovery Identity and Access Management Python (Programming Language) Linux System Administration Reliability Engineering Prometheus Datadog Scripting Google Cloud Cloud Monitoring Grafana Containerization Kubernetes Information Technology Google Cloud Functions Performance Monitor Data Management Cloud Optimization Terraform Data Pipelines Docker

Job description

  • Design, build, and continuously improve the reliability, availability, scalability, and performance of enterprise cloud data platforms across Google Cloud Platform (GCP) and Amazon Web Services (AWS).
  • Deploy, automate, and manage cloud infrastructure using Infrastructure as Code (Terraform preferred) and modern DevOps practices.
  • Build and enhance observability across cloud infrastructure, databases, and data pipelines using Datadog and cloud-native monitoring solutions.
  • Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational health metrics.
  • Partner with Data Engineering teams to improve data platform reliability, data quality, disaster recovery, and operational resilience through proactive monitoring and early issue detection.
  • Develop automation and self-healing solutions to reduce operational toil, streamline repetitive tasks, and improve engineering efficiency.
  • Lead production incident response, root cause analysis (RCA), and post-incident reviews, driving permanent improvements to platform reliability.
  • Drive cloud governance and FinOps initiatives by optimizing resource utilization, cloud costs, and operational best practices across GCP and AWS.
  • Evaluate and introduce modern SRE, DevOps, and cloud technologies that improve platform reliability, operational maturity, and engineering productivity.
  • Create and maintain operational documentation, runbooks, and recovery procedures while mentoring engineers and promoting Site Reliability Engineering best practices.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or an equivalent qualification.
  • 3+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering, or a similar role supporting enterprise production environments.
  • Hands-on experience with Google Cloud Platform (preferred) and/or Amazon Web Services.
  • Experience supporting cloud-native data platforms and services such as BigQuery, Cloud Composer, Pub/Sub, Cloud Run, Cloud Functions, DataStream, and Cloud Storage.
  • Experience with Infrastructure as Code using Terraform (preferred) or similar technologies.
  • Experience with Linux Administration, CI/CD pipelines, container platforms (Kubernetes/Docker), and automation using Python, Bash, or similar scripting languages.
  • Hands-on experience with observability and monitoring platforms such as Datadog, Cloud Monitoring, Prometheus, or Grafana.
  • Strong understanding of cloud security, IAM, FinOps principles, and cloud cost optimization best practices, and data reliability principles.
  • Proven experience in incident management, root cause analysis, and driving reliability improvements in production environments.
  • Excellent communication, collaboration, and documentation skills, with the ability to mentor engineers and promote SRE best practices.

About the company

Sysco LABS is the Global In-House Center of Sysco Corporation (NYSE: SYY), the world’s largest foodservice company. Sysco ranks 56th in the Fortune 500 list and is the global leader in the trillion-dollar foodservice industry.

Sysco employs over 75,000 associates, has 337 smart distribution facilities worldwide and over 14,000 IoT-enabled trucks serving 730,000 customer locations. For fiscal year 2025 that ended June 29, 2025, the company generated sales of more than $81.4 billion.

Sysco LABS Sri Lanka delivers the technology that powers Sysco’s end-to-end operations.

Sysco LABS’ enterprise technology is present in the end-to-end foodservice journey, enabling the sourcing of food products, merchandising, storage and warehouse operations, order placement and pricing algorithms, the delivery of food and supplies to Sysco’s global network and the in-restaurant dining experience of the end-customer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on wd5.myworkdaysite.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all