Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Experteer Overview As a Site Reliability Engineer on the CX engineering team, you will manage and optimize large-scale, multi-cloud environments to ensure high availability and security. You will work with cross-functional delivery teams to provide proactive monitoring, automation, and design guidance that improves service reliability. This role focuses on a world-class monitoring platform and network-centric infrastructure across compute, virtualization, and container environments. You will shape architecture and incident response while mentoring others and delivering meaningful business outcomes. Compensation / Benefits * Administer and automate tasks in tools like ServiceNow and Splunk within a network management architecture * Write and optimize SPL queries, configure alerts, and build dashboards for proactive monitoring and faster incident response * Work with compute, virtualization, and container environments (Cisco UCS, HyperFlex, VMware, Microsoft, Kubernetes) * Engage with multi-cloud environments, primarily AWS, for network integration * Apply automation and orchestration to improve customer stacks using Python, Ansible, Terraform, etc. * Develop and review High-Level and Low-Level Design, and Implementation/Change Management Plans * Diagnose and resolve outages across OS, hardware, network, and software, especially under high-priority conditions * Evolve application infrastructure architecture and recommend improvements * Lead sys-admin functions: patching, security configuration, reporting, and compliance monitoring Tasks * Hands-on experience with Splunk and ServiceNow * Bachelor’s degree and 6+ years of professional experience * Experience as a Linux Administrator or hands-on Linux (Alma, CentOS) * Experience with automation (Python, Ansible, Terraform) and virtualization/orchestration * Experience in ITIL-aligned customer support processes Key requirements * medical, dental and vision insurance * 401(k) with matching * paid parental leave * paid time off and holidays * sick time off * volunteer days
Requirements
multi-cloud environments, primarily AWS, for network integration * Apply automation and orchestration to improve customer stacks using Python, Ansible, Terraform, etc. * Develop and review High-Level and Low-Level Design, and Implementation/Change Management Plans * Diagnose and resolve outages across OS, hardware, network, and software, especially under high-priority conditions * Evolve application infrastructure architecture and recommend improvements * Lead sys-admin functions: patching, security configuration, reporting, and compliance monitoring Tasks * Hands-on experience with Splunk and ServiceNow * Bachelor’s degree and 6+ years of professional experience * Experience as a Linux Administrator or hands-on Linux (Alma, CentOS) * Experience with automation (Python, Ansible, Terraform) and virtualization/orchestration * Experience in ITIL-aligned customer support processes Key requirements * medical, dental and vision insurance * 401(k) with matching * paid parental leave * paid a and off and holidays * sick time off * volunteer days
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Best Countries for Software Engineers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence