> Markdown version of [/jobs/ext/1254379-customer-reliability-engineer-hypershield](https://www.wearedevelopers.com/jobs/ext/1254379-customer-reliability-engineer-hypershield). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Customer Reliability Engineer, Hypershield... - **Company:** Cisco Systems, Inc. - **Location:** Boise, ID, United States - **Experience:** Expert - **Salary:** $158,200.0 - $200,700.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Command-Line Interface, Data Centers, Linux, Network Troubleshooting, Netconf, NetFlow, Network Control, Network Segmentation, Packet Analyzer, Cisco Nexus Switches, Reliability Engineering, Site Reliability Engineering Practices, Ansible, VMware VSphere, Nx-os, Kubernetes, Cisco - **Published:** July 13, 2026 - **Apply:** https://www.juju.com/job/00000000gfyv7y ## About the Role + Bachelor's + 8 years of experience, Master's + 6 years, or equivalent industry experience + Experience supporting enterprise customers in an escalation capacity, including diagnosing and resolving complex production incidents under SLA pressure in unfamiliar environments + Experienceoperating and troubleshooting Cisco Nexus / NX-OS in production; equivalent depth on another major vendor accepted + Prior experienceto localize failures across a layered data-center architecture spanning switching/forwarding, services/enforcement, and control-plane domains + Linux operations experience at the command line, including production troubleshooting, with working exposure to containers or Kubernetes Preferred Qualifications + Direct experience using network troubleshooting tooling as a primary diagnostic method, including packet capture and flow-telemetry analysis (NetFlow/IPFIX) + Working knowledge of enterprise virtualization, sufficient to troubleshoot a VM-based appliance deployment; vSphere is the current deployment target + Operational Kubernetes and Helmproficiency, including diagnosing failures beyond the workload level: TLS certificates and service-account authentication, API-server connectivity, service exposure, persistent storage for stateful workloads, and custom resources and operators + Working knowledge of VXLAN EVPN fabrics, including the Smart Switch's placement within them and the ability to isolate faults across the fabric + Network segmentation andfirewallpolicy design (zone-based ormicrosegmentation), and thetradeoffsvs. traditional NGFWs + Familiarity with the NetOps/NetSecOpsoperating split in data-center security + Experience driving diagnosis and remediation through a customer's own team, working incidents in environments with no direct access where the customer accomplishes the steps + Experience with NX-OS automation and APIs (NX-API, NETCONF/RESTCONF,gNMI, or Ansible); familiarity with Cisco Nexus Dashboard a plus + Demonstrated ability to communicate incident status, root cause, and remediation clearly to both technical and executive audiences, verbally and in writing + CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert (a plus) ## Description The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield on the Cisco Nexus N9300 Series Smart Switches. The team owns the hardest break/fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the data-center fabric and the on-premises Kubernetes controller that manages the security policy enforced on it. The work demands methodical diagnosis, composure under incident pressure, and the ability to operate at the seam between customer environments and engineering. Your Impact The ideal candidate combines deep networking expertise with strong troubleshooting skills, customer-facing experience, and a passion for improving reliability across complex product environments. + Own Hypershieldcases escalated from Cisco TAC through to resolution, engaging customers directly as the incident requires + Diagnose complex production failures through the Hypershieldsurface: the N9300 Smart Switch fabric and the on-premises Kubernetes controller that manages its security policy + Localize faults across the layered architecture: switching andforwarding, security services and enforcement, and the control plane + Develop a deep understanding of each customer's architecture and configuration, and diagnose failures in unfamiliar production environments + Reproduce customer failures, partner with engineering to drive fixes, and own the fix back to the customer + Convert individual cases into systemic improvements: runbooks, diagnostics, knowledge-basecontent, and product feedback to engineering + Help build the team's proactive view of customer health, developing new monitoring, tooling, and reliability practices as the installed base grows ## Related Videos - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [Your Infrastructure Is Not a Playground: AI Agents for Infra Done Right](https://www.wearedevelopers.com/videos/2084-your-infrastructure-is-not-a-playground-ai-agents-for-infra-done-right) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)