Site Reliability Engineer

Cisco Systems, Inc.
Durham, NC, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services CentOS Software Design Documents Linux Python (Programming Language) Microsoft Software Network Administration Reliability Engineering Ansible System Availability Multi-Cloud Kubernetes
+7 more
Performance Monitor Enterprise Integration Terraform Splunk Cisco Servicenow Vmware

Job description

Experteer Overview As a Site Reliability Engineer on the CX engineering team, you will manage and optimize large-scale, multi-cloud environments to ensure high availability and security. You will work with cross-functional delivery teams to provide proactive monitoring, automation, and design guidance that improves service reliability. This role focuses on a world-class monitoring platform and network-centric infrastructure across compute, virtualization, and container environments. You will shape architecture and incident response while mentoring others and delivering meaningful business outcomes. Compensation / Benefits * Administer and automate tasks in tools like ServiceNow and Splunk within a network management architecture * Write and optimize SPL queries, configure alerts, and build dashboards for proactive monitoring and faster incident response * Work with compute, virtualization, and container environments (Cisco UCS, HyperFlex, VMware, Microsoft, Kubernetes) * Engage with multi-cloud environments, primarily AWS, for network integration * Apply automation and orchestration to improve customer stacks using Python, Ansible, Terraform, etc. * Develop and review High-Level and Low-Level Design, and Implementation/Change Management Plans * Diagnose and resolve outages across OS, hardware, network, and software, especially under high-priority conditions * Evolve application infrastructure architecture and recommend improvements * Lead sys-admin functions: patching, security configuration, reporting, and compliance monitoring Tasks * Hands-on experience with Splunk and ServiceNow * Bachelor’s degree and 6+ years of professional experience * Experience as a Linux Administrator or hands-on Linux (Alma, CentOS) * Experience with automation (Python, Ansible, Terraform) and virtualization/orchestration * Experience in ITIL-aligned customer support processes Key requirements * medical, dental and vision insurance * 401(k) with matching * paid parental leave * paid time off and holidays * sick time off * volunteer days

Requirements

multi-cloud environments, primarily AWS, for network integration * Apply automation and orchestration to improve customer stacks using Python, Ansible, Terraform, etc. * Develop and review High-Level and Low-Level Design, and Implementation/Change Management Plans * Diagnose and resolve outages across OS, hardware, network, and software, especially under high-priority conditions * Evolve application infrastructure architecture and recommend improvements * Lead sys-admin functions: patching, security configuration, reporting, and compliance monitoring Tasks * Hands-on experience with Splunk and ServiceNow * Bachelor’s degree and 6+ years of professional experience * Experience as a Linux Administrator or hands-on Linux (Alma, CentOS) * Experience with automation (Python, Ansible, Terraform) and virtualization/orchestration * Experience in ITIL-aligned customer support processes Key requirements * medical, dental and vision insurance * 401(k) with matching * paid parental leave * paid a and off and holidays * sick time off * volunteer days

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

1:29 min

Expanding practical knowledge with community sandboxes and resources

Stuart Clark · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · WWC 2024

Videos

See all

Related articles

See all