Senior Site Reliability Engineer

Sysco Corporation
United States
3 days ago
Apply on wd5.myworkdaysite.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Microsoft Azure Software as a Service Cloud Computing Cloud Engineering Databases Continuous Integration Software Debugging
+28 more
Disaster Recovery Distributed Systems Github Python (Programming Language) Enterprise Messaging Systems Octopus Deploy Reliability Engineering Prometheus Software Engineering TypeScript Datadog Istio Grafana Caching Reliability of Systems Build Management Containerization Kubernetes Infrastructure Automation Frameworks Production Code Build Process Api Gateway Terraform Splunk Jenkins Golang Programming Languages Microservices

Job description

Join the Sysco Commercial Technology (CT) Site Reliability Engineering team as a Senior Software Site Reliability Engineer, where you will improve the reliability, scalability, performance, security, and operational excellence of Sysco’s Commercial Technology ecosystem. Our team is responsible for delivering end-to-end reliability across the technology stack-from cloud infrastructure and platform services to APIs, applications, and customer-facing digital experiences-supporting enterprise API platforms, digital commerce, B2B integrations, and other business-critical systems across the CT landscape.

This is a Software SRE role focused on applying software engineering principles to solve reliability challenges at scale. You will design and build automation, develop internal tools and platforms, enhance observability, reduce operational toil, and deliver engineering solutions that strengthen the reliability and resilience of distributed systems.

As a Senior Software Site Reliability Engineer, you will combine software engineering, systems thinking, and production reliability expertise to identify and address reliability gaps, improve incident response, and drive long-term reliability initiatives. Working closely with product, platform, and infrastructure engineering teams, you will help build highly available, scalable, and resilient services while embedding reliability throughout the software development lifecycle.

What You’ll Do

  • Own the reliability, scalability, performance, and operational excellence of critical platforms and services across the Commercial Technology ecosystem.
  • Apply software engineering principles to design and build systems that improve reliability, resilience, and engineering productivity.
  • Design, develop, and maintain automation, internal platforms, self-service capabilities, and reliability engineering tools that eliminate operational toil.
  • Define and evolve service reliability through SLIs, SLOs, error budgets, observability standards, and actionable operational metrics.
  • Partner with product, platform, and infrastructure engineering teams to design reliable, scalable, and operable systems from inception through production.
  • Build engineering solutions that improve deployment safety, release automation, progressive delivery, and production readiness.
  • Investigate complex production issues using application code, distributed systems knowledge, logs, metrics, traces, and infrastructure telemetry to identify systemic failures and drive permanent solutions.
  • Lead incident reviews and postmortems, ensuring root causes are understood and translated into engineering improvements rather than operational workarounds.
  • Drive continuous reliability improvements by identifying recurring operational patterns and solving them through software, automation, or platform capabilities.
  • Improve the resilience of distributed systems through capacity planning, performance engineering, disaster recovery, and resilience testing.
  • Participate in architecture and design reviews to ensure systems are reliable, scalable, observable, secure, and operationally efficient.
  • Contribute to shared engineering libraries, frameworks, and platform capabilities that enable development teams to build and operate reliable services.
  • Participate in an on-call rotation, using operational experience to continuously improve system reliability, automate manual work, and reduce future operational burden.
  • Mentor engineers and champion Software SRE principles, engineering excellence, and a culture of reliability across the organization., * Build software that improves the reliability of business-critical platforms across Sysco’s Commercial Technology ecosystem.
  • Solve large-scale distributed systems challenges using software engineering, automation, and platform engineering.
  • Shape the future of Software SRE by driving automation, observability, and engineering best practices.
  • Work across the entire technology stack-from cloud infrastructure and Kubernetes to APIs and applications.
  • Collaborate with talented engineers while making a measurable impact on reliability, scalability, and operational excellence.

Requirements

  • 4+ years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Production Engineering, or a related engineering role.
  • Strong software engineering skills with experience designing, building, testing, and operating reliable production systems.
  • Proficiency in one or more programming languages such as Java, Go, Python, or JavaScript/TypeScript, with a focus on writing maintainable, production-quality code.
  • Solid understanding of distributed systems, cloud-native architectures, microservices, APIs, databases, messaging systems, caching, networking, and system design.
  • Experience applying Site Reliability Engineering principles, including SLIs, SLOs, error budgets, incident management, postmortems, toil reduction, and automation.
  • Experience building automation, internal tools, developer platforms, or operational tooling to improve reliability and engineering productivity.
  • Hands-on experience with observability practices, including metrics, logs, traces, profiling, and modern observability platforms such as Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, or ELK.
  • Experience debugging complex production issues across applications, distributed systems, APIs, infrastructure, and cloud platforms.
  • Experience with Kubernetes, containers, CI/CD, Infrastructure as Code, and public cloud platforms.
  • Ability to read, understand, and debug application code to identify systemic issues and implement long-term engineering solutions.
  • Strong analytical and systems thinking skills, with the ability to translate operational challenges into scalable engineering improvements.
  • Excellent collaboration and communication skills, with experience working across product, software, platform, and infrastructure engineering teams.
  • Demonstrated ownership, curiosity, and a continuous improvement mindset, with the ability to drive reliability initiatives independently., * Experience with large-scale distributed systems, enterprise SaaS, digital commerce, or high-traffic platforms.
  • Experience building internal developer platforms, automation frameworks, or reliability engineering tools.
  • Experience with AWS, GCP, or Azure and cloud-native technologies such as Kubernetes.
  • Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, or similar platform engineering tools.
  • Experience with service mesh, API gateways, messaging systems, caching, database reliability, or resilience engineering.
  • Experience with AIOps, AI-assisted operations, or modern observability platforms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on wd5.myworkdaysite.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all