> Markdown version of [/jobs/ext/1933581-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1933581-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Sysco Corporation - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Bash Shell, BigQuery, Cloud Computing, Cloud Computing Security, Cloud Engineering, Cloud Storage, Databases, Continuous Integration, Information Engineering, Data Infrastructure, Data Integrity, DevOps, Disaster Recovery, Identity and Access Management, Python (Programming Language), Linux System Administration, Reliability Engineering, Prometheus, Datadog, Scripting, Google Cloud, Cloud Monitoring, Grafana, Containerization, Kubernetes, Information Technology, Google Cloud Functions, Performance Monitor, Data Management, Cloud Optimization, Terraform, Data Pipelines, Docker - **Published:** August 5, 2026 - **Apply:** https://wd5.myworkdaysite.com/recruiting/sysco/syscocareers/job/Sysco-LABS-----Sri-Lanka/Senior-Site-Reliability-Engineer_R250904 ## About the Role * Bachelor's degree in Computer Science, Information Technology, Engineering, or an equivalent qualification. * 3+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering, or a similar role supporting enterprise production environments. * Hands-on experience with Google Cloud Platform (preferred) and/or Amazon Web Services. * Experience supporting cloud-native data platforms and services such as BigQuery, Cloud Composer, Pub/Sub, Cloud Run, Cloud Functions, DataStream, and Cloud Storage. * Experience with Infrastructure as Code using Terraform (preferred) or similar technologies. * Experience with Linux Administration, CI/CD pipelines, container platforms (Kubernetes/Docker), and automation using Python, Bash, or similar scripting languages. * Hands-on experience with observability and monitoring platforms such as Datadog, Cloud Monitoring, Prometheus, or Grafana. * Strong understanding of cloud security, IAM, FinOps principles, and cloud cost optimization best practices, and data reliability principles. * Proven experience in incident management, root cause analysis, and driving reliability improvements in production environments. * Excellent communication, collaboration, and documentation skills, with the ability to mentor engineers and promote SRE best practices. ## Description * Design, build, and continuously improve the reliability, availability, scalability, and performance of enterprise cloud data platforms across Google Cloud Platform (GCP) and Amazon Web Services (AWS). * Deploy, automate, and manage cloud infrastructure using Infrastructure as Code (Terraform preferred) and modern DevOps practices. * Build and enhance observability across cloud infrastructure, databases, and data pipelines using Datadog and cloud-native monitoring solutions. * Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational health metrics. * Partner with Data Engineering teams to improve data platform reliability, data quality, disaster recovery, and operational resilience through proactive monitoring and early issue detection. * Develop automation and self-healing solutions to reduce operational toil, streamline repetitive tasks, and improve engineering efficiency. * Lead production incident response, root cause analysis (RCA), and post-incident reviews, driving permanent improvements to platform reliability. * Drive cloud governance and FinOps initiatives by optimizing resource utilization, cloud costs, and operational best practices across GCP and AWS. * Evaluate and introduce modern SRE, DevOps, and cloud technologies that improve platform reliability, operational maturity, and engineering productivity. * Create and maintain operational documentation, runbooks, and recovery procedures while mentoring engineers and promoting Site Reliability Engineering best practices. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)