> Markdown version of [/jobs/ext/553310-senior-network-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/553310-senior-network-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Network & Site Reliability Engineer - **Company:** Alembic, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $210,000.0 - $240,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Bash Shell, Border Gateway Protocol, Cloud Computing, Configuration Management, Complex Networks, Computer Networks, Data Center Infrastructure Management (CIM), Software Debugging, Linux, Monitoring of Systems, InfiniBand, Networking Hardware, Internet Protocol Security (IP SEC), Storage Area Network (SAN), Virtual Private Networks (VPN), Multi-protocol Systems, Internet Small Computer System Interface (ISCSI), Python (Programming Language), Network Security, Machine Learning, Network Administration, Peering, Reliability Engineering, Ansible, Prometheus, Supercomputing, Wide Area Networks, Datadog, Load Balancing, Grafana, Apache Spark, Reliability of Systems, Firewalls (Computer Science), Kubernetes, Apache Kafka, Terraform, Open Network Automation Platform - **Published:** June 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d0228d464ec90d0e ## About the Role Do you have experience in WAN?, * 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration. * A strong background in network security, architecture, design, and operations. * Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols - BGP, QoS, MPLS, and IPsec VPNs. * Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP). * Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar). * WAN engineering - carrier circuit provisioning and external network peering. * Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure. * Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry). * Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication., * NVIDIA networking technologies - Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs (this is the fabric behind our SuperPOD). * Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI). * Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments. ## Description We're building infrastructure that has to perform under real-world scale, reliability, and security demands - and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role. You'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers - the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions. It's a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it. What You'll Do * Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads. * Own network device configuration management end to end, ensuring consistency and reliability across the fleet. * Improve system and network reliability and performance through automation, observability, and proactive capacity planning. * Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering. * Build and maintain comprehensive monitoring, alerting, and incident response - SLOs, runbooks, and on-call rotations - and drive post-incident analysis and continuous improvement. * Ensure security, compliance, and operational readiness across our network and cloud infrastructure. * Partner across engineering and data science to drive a culture of performance and reliability. ## Related Videos - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Transforming Education: A Journey from interactive Markdown to Remote-Labs](https://www.wearedevelopers.com/videos/941-transforming-education-a-journey-from-interactive-markdown-to-remote-labs) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)