> Markdown version of [/jobs/ext/1897463-principal-network-engineer-reliability](https://www.wearedevelopers.com/jobs/ext/1897463-principal-network-engineer-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Network Engineer - Reliability - **Company:** Wells Fargo - **Location:** Dallas, TX, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Border Gateway Protocol, Cloud Computing, Configuration Management, Computer Networks, Computer Engineering, Data Centers, Software Design Patterns, Domain Name System (DNS), Python (Programming Language), Network Security, Network Planning and Design, Routing, Network Service, Network Time Protocols, Open Shortest Path First (OSPF), Windows PowerShell, Reliability Engineering, Ansible, TCP/IP, Datadog, Scripting, Computer Networking Systems, Load Balancing, Computer Network Operations, Mttr, Reliability of Systems, Firewalls (Computer Science), Juniper, Information Technology, Performance Monitor, Terraform, Service Stack - **Published:** August 2, 2026 - **Apply:** https://dejobs.org/x/x/DF9C8DEF70D34F06A015E165B487CF48/job/ ## About the Role You should demonstrate principal-level technical depth, operational judgment, and cross-functional influence, including: * 7+ years of experience in network engineering, reliability engineering, or network operations supporting business-critical enterprise network services. * 5+ years of expert knowledge of routing, switching, TCP/IP, BGP, OSPF, high-availability design, and related network services including load balancing, DNS, NTP, firewalls, and network security. * 5+ years of Proven ability to lead complex incident response, guide troubleshooting, restore service, and communicate clearly under pressure. * 5 plus years of practical experience applying SRE or reliability engineering practices, including SLOs/SLIs, problem management, service health measurement, operational health metrics, and continuous improvement. * 5 plus years experience improving operational stability and reducing repeat incidents through root cause analysis, post-incident reviews, corrective action planning, systemic remediation, and measurable recurrence prevention. * 5 plus years experience owning corrective actions from post-incident review through implementation, validation, and closure. * 5 plus years experience improving change success rates through risk assessment, peer review, validation testing, rollback planning, and post-change verification. * 5 plus years experience conducting operational readiness reviews, production readiness assessments, failure-mode reviews, or service acceptance reviews. * 5 plus years experience managing vendor escalations, product defects, support cases, and platform lifecycle risks that impact service stability. * 5 plus years experience with capacity planning, performance analysis, traffic engineering, and resiliency planning for large-scale enterprise networks. * 5 plus years of hands-on automation or scripting experience with tools such as Python, PowerShell, Bash, Ansible, Terraform, or equivalent technologies. * 5 plus years of experience improving monitoring, telemetry, alerting, dashboards, metrics, service health reporting, or performance visibility for critical infrastructure services. Preferred Qualifications The following qualifications are beneficial and will help a candidate be successful in this role: * Experience influencing enterprise network architecture and aligning designs with infrastructure, cloud, security, compliance, and architecture standards. * Strong technical leadership, collaboration, mentoring, written communication, verbal communication, and stakeholder influence skills, including the ability to explain complex technical issues to senior stakeholders. * Engineering, or a related technical field, or equivalent practical experience. * Relevant networking, cloud, reliability, or IT service management certifications, such as CCNP/CCIE, Juniper, Arista, AWS/Azure networking, ITIL, or equivalent credentials. * Ability to prioritize reliability improvements based on business impact, operational risk, service criticality, and incident trends. * Experience contributing to technical strategy, engineering standards, reusable design patterns, or network reliability innovation. * Experience with network modernization, data center migration, SDN, cloud networking, or reliability transformation initiatives. * Experience in regulated or critical industries such as financial services, healthcare, telecommunications, or similar environments. * Master's degree in Telecommunications, Network Engineering, Computer Engineering, or a related field. * Experience reading and interpreting packet captures to diagnose network performance, connectivity, protocol, and application-impacting issues., Education: Bachelor's degree in Computer Science, Electrical/Network ## Description * Role Type: Individual contributor; this is not a people-manager role. * Work Model: Hybrid work model with three days per week in the office. * Eligible Locations: Dallas, TX metro; Charlotte, NC metro; or Chandler, AZ metro. * Escalation Expectations: Provides senior technical escalation for critical incidents and high-risk changes. This role is expected to support critical incident response as needed; any recurring on-call rotation will be clearly defined before offer acceptance. * Travel Expectations: 5% or less. Core Technology Stack This role is network-centric, with primary focus on routing and switching technologies, data center fabrics, WAN and campus networking, cloud and hybrid connectivity, load balancing, DNS, NTP, firewalls, automation, telemetry, monitoring, and observability tooling., You will provide senior technical leadership, hands-on engineering guidance, and cross-functional influence across the following areas: * Network Reliability Strategy: Define and drive reliability targets, SLOs/SLIs, risk measures, service health indicators, and improvement plans for critical network services. * Incident Leadership: Lead major incident response and service restoration, and drive root cause analysis, corrective action planning, and prevention of repeat incidents. * Architecture Influence: Review, challenge, and guide network designs to improve resiliency, scalability, security, capacity, failure isolation, and operational simplicity. * Automation and Observability: Expand telemetry, alerting, service health reporting, configuration management, automated validation, and self-service capabilities to reduce manual toil. * Operational Excellence: Strengthen standards, runbooks, documentation, change quality, failure readiness, compliance-aligned practices, and production readiness reviews. * Cross-Functional Leadership: Partner with network, cloud, security, infrastructure, application, architecture, and vendor teams while mentoring engineers and communicating complex technical issues clearly to senior stakeholders., Success in this role is measured by stronger reliability outcomes, improved operational readiness, and reduced network risk. * Improved availability, stability, MTTD, MTTR, change success rate, and incident trends for network services. * Fewer repeat incidents through stronger root cause analysis, corrective action tracking, and systemic prevention. * Broader adoption of reliability standards, runbooks, design patterns, operational playbooks, and production readiness practices. * Reduced manual operational toil through practical automation, improved validation, and clearer service health visibility. * Stronger confidence from engineering, application, business, and leadership stakeholders in network stability, resilience, and change readiness., Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Best Paying Remote Jobs](https://www.wearedevelopers.com/magazine/255-best-paying-remote-jobs) - [Top Characteristics of a Software Engineer](https://www.wearedevelopers.com/magazine/166-top-characteristics-of-a-software-engineer) - [How to Answer the Interview Question: “Why Do You Want to Be a Software Engineer?”](https://www.wearedevelopers.com/magazine/392-how-to-answer-the-interview-question-why-do-you-want-to-be-a-software-engineer)