> Markdown version of [/jobs/ext/3659510-network-engineer](https://www.wearedevelopers.com/jobs/ext/3659510-network-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # network engineer - **Company:** Crusoe Cloud - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $195,000.0 - $235,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Border Gateway Protocol, Common Lisp Object Systems, Data Centers, Monitoring of Systems, Multi-protocol Systems, Junos, Python (Programming Language), NetFlow, Open Shortest Path First (OSPF), Operational Databases, Overlay Transport Virtualization, Performance Tuning, Remote Direct Memory Access, Prometheus, Simple Network Management Protocols, TCP/IP, Traffic Analysis, Scripting, Cloud Platform System, High Performance Computing, Delivery Pipeline, Grafana, Reliability of Systems, Juniper, Information Technology, Ddos - **Published:** October 9, 2026 - **Apply:** https://www.thejobnetwork.com/job/d13fba62-add7-4a5a-89ba-d68c5f36dc24/staff-network-production-engineer-operations ## About the Role The ideal candidate is a seasoned network engineer with deep operational experience in large-scale environments who thrives in high-pressure situations and takes pride in keeping systems healthy. You'll contribute to defining SLIs and SLOs, improving observability tooling, building automation to reduce toil, and mentoring peers - all while serving as a key escalation point during high-severity network events., * 8+ years of production network engineering experience with a focus on operations, incident response, and reliability in large-scale or internet-scale environments. * Strong Python and scripting proficiency: comfortable writing diagnostic tooling, auto-remediation scripts, and automation pipelines from scratch, not just modifying existing scripts. * Hands-on experience with observability and monitoring tools including streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes, and the ability to script against their APIs. * Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including familiarity with PFC, ECN, and DCQCN tuning. * Expert hands-on knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments. * Proficiency with Arista (EOS) and Juniper (Junos) platforms in leaf-spine CLOS architectures across multi-vendor environments. * Comfort operating large device fleets across multi-region environments with on-call responsibility, including experience as an escalation point during critical events. * Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience. Bonus Points: * Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments. * Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility. * Experience defining or contributing to SLIs and SLOs in partnership with SRE or product teams. * Exposure to operating 10K+ device fleets across hyperscale or cloud environments. * Background contributing to post-incident learning programs or operational excellence initiatives org-wide. ## Description * Production Reliability: Help own uptime across Crusoe's global edge, backbone, data center, and GPU cluster network, directly supporting AI workloads at scale. * Operational Automation: Build and maintain Python-based tooling to reduce toil, automate common remediation workflows, and accelerate mean time to resolution across the operations org. * Automate Incident Response: Develop scripts and tooling that speed up detection, triage, and mitigation during high-severity network events, and contribute directly to the response itself, including stakeholder communication and postmortem documentation. * Automate Root Cause Analysis: Build tooling to surface systemic issues faster during RCAs, and drive remediation plans tracked through to closure. * Observability Automation: Write automation on top of Crusoe's monitoring stack (streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, ThousandEyes) to reduce manual triage and surface actionable signal faster. * Operational Standards: Author and maintain runbooks, escalation playbooks, and SOPs, with an eye toward what can be scripted or automated rather than performed manually. * SLI/SLO Contribution: Partner with Architecture and SRE teams to define and track network reliability metrics, and build the automated dashboards and alerting that back them. * Mentorship: Provide technical guidance to Senior engineers, particularly around automation practices and reducing operational toil, and contribute to a culture of operational excellence.