> Markdown version of [/jobs/ext/1091123-remote](https://www.wearedevelopers.com/jobs/ext/1091123-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** Calix - **Location:** San Francisco Bay Area, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $136,000.0 - $266,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Application Layers, Border Gateway Protocol, Big Data, BigQuery, Cloud Computing, Computer Programming, Network Congestion, Data Infrastructure, Data Integration, Data Transport Utility, Software Debugging, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Identity and Access Management, OSI Models, Python (Programming Language), Transport Layer, PostgreSQL, Linux Kernel, Machine Learning, Network Architecture, Routing, Packet Analyzer, Open Shortest Path First (OSPF), Prometheus, Session Management, Data Streaming, Transmission Control Protocol (TCP), Wireshark, Network Switches, Computer Networking Systems, Google Cloud, System Availability, Grafana, Containerization, Kubernetes, Low Latency, Hashicorp, Apache Kafka, Terraform - **Published:** June 18, 2026 - **Apply:** https://www.builtincolorado.com/job/staff-site-reliability-operations-engineer/9785040 ## About the Role * Location/Work Style: Proven track record of high autonomy and successful delivery in a 100% remote engineering environment. * Experience: 8+ years in SRE, Production Engineering, or Distributed Systems infrastructure roles. * Networking Expertise (L1-L7): Deep technical knowledge and debugging mastery across all OSI layers, including: * L1-L3: Physical/fiber infrastructure awareness, switching, and advanced routing protocols (BGP, OSPF). * L4: Transport layer tuning (TCP congestion control algorithms, UDP, QUIC). * L5-L7: Session management, TLS termination, DNS architecture, and advanced application protocols (HTTP/3, gRPC). * Orchestration & Containerization: Expert-level mastery of Google Kubernetes Engine (GKE) internals, custom controllers, multi-cluster networking, and GitOps workflows. * Data Infrastructure: Proven track record managing high-throughput Apache Kafka pipelines and large-scale data environments across PostgreSQL, AlloyDB, and BigQuery. * Grafana Ecosystem: Deep, hands-on experience deploying and managing Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale. * AIOps Implementation: Track record applying AI/ML techniques for time-series anomaly detection, log clustering, and correlation (e.g., Grafana Adaptive Metrics, BigPanda). * Infrastructure as Code: Advanced, production-scale expertise utilizing HashiCorp Terraform exclusively to provision and manage multi-region GCP cloud architectures. * Programming: High proficiency in Go and Python for building custom infrastructure tooling, Kubernetes operators, and data integration scripts. Preferred Attributes * Remote Communicator: Exceptional written and verbal communication skills, with an emphasis on creating clear documentation for asynchronous alignment. * GCP Expert: Deep knowledge of Google Cloud architectural best practices, Cloud SDN, Cloud Armor, Interconnect, Identity and Access Management (IAM), and cost optimization. * Systems Thinker: Deep understanding of Linux internals, eBPF-based monitoring, kernel-level networking, and packet analysis tools (Wireshark, tcpdump). ## Description We are seeking a Staff Site Reliability Engineer (SRE) to lead our global platform reliability and drive our next-generation observability strategy on Google Cloud Platform (GCP). In this role, you will leverage Grafana Labs' complete telemetry stack and AIOps methodologies to build intelligent, self-healing infrastructure. You will bring deep expertise in scaling enterprise-grade Google Kubernetes Engine (GKE) topologies, managing high-throughput Kafka event streams, and maintaining high-performance PostgreSQL, AlloyDB, and BigQuery ecosystems at massive scale. Crucially, you will provide deep technical leadership across the entire networking stack, diagnosing complex issues from physical-layer transport up to application-layer protocols., * Full-Stack Network Architecture: Architect, optimize, and troubleshoot complex networking infrastructure spanning Layer 1 through Layer 7, ensuring low-latency data transport, secure edge routing, and seamless service mesh integration. * Grafana Stack Architecture: Design, scale, and optimize our unified observability platform using the Grafana Labs suite (Grafana, Mimir, Loki, Tempo, and Beyla). * AIOps & Intelligent Alerting: Deploy machine learning models and automated anomaly detection to cut through telemetry noise, reduce alert fatigue, and predict network or data pipeline bottlenecks. * GKE Platform Engineering: Drive the architecture, scaling, security, and networking of production Google Kubernetes Engine (GKE) clusters. * Data & Event Streaming Reliability: Tune, and maintain high-throughput Apache Kafka clusters to guarantee low-latency event delivery and high availability. * Large-Scale Database Management: Ensure the performance, scalability, and disaster recovery readiness of our transactional and analytical data tiers across PostgreSQL, AlloyDB, and BigQuery. * Automated Incident Response: Integrate AIOps insights with Grafana workflows to automate triage, accelerate root-cause analysis, and trigger auto-remediation scripts. * Technical Leadership: Champion the long-term technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards. * Mentorship: Coach senior and junior engineers on advanced debugging techniques, distributed systems thinking, and intelligent operations across a distributed workforce. ## Related Videos - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Paying Remote Jobs](https://www.wearedevelopers.com/magazine/255-best-paying-remote-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Mastering Remote Work: Tips for Developers](https://www.wearedevelopers.com/magazine/558-mastering-remote-work-tips-for-developers) - [Remote Work: Best Practices for Developers](https://www.wearedevelopers.com/magazine/315-remote-work-best-practices-for-developers)