> Markdown version of [/jobs/ext/2058230-staff-software-engineer-observability](https://www.wearedevelopers.com/jobs/ext/2058230-staff-software-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - Observability - **Company:** Snowflake Inc. - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Salary:** $236,000.0 - $339,200.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Cloud Computing, Profiling, Databases, Data Structures, Distributed Systems, Linux Kernel, Network Protocols, Performance Tuning, Prometheus, Pulumi, Google Cloud, Snowflake, Grafana, Concurrency, Multi-Cloud, Backend, Infrastructure Automation Frameworks, Apache Flink, Apache Kafka, Vertica, Terraform, Stream Processing, Dynatrace, Network Optimization, Golang - **Published:** August 14, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/86906484/1 ## About the Role You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry., * 10+ years of professional experience building infrastructure and backend distributed systems at scale using languages such as Go, C++, Java, or Rust. * Proven track record of architecting, deploying, and maintaining hyper-scale distributed platforms in public cloud environments (AWS, Azure, or GCP). * Deep theoretical and practical CS fundamentals (data structures, algorithms, concurrency patterns, storage engines, distributed consensus, networking protocols). * Strong experience with infrastructure-as-code (IaC) tools such as Terraform, or Pulumi. * Hands-on expertise with time-series databases, high-cardinality metric stores, distributed tracing frameworks (OpenTelemetry, Jaeger), or log streaming systems (Kafka, Flink, ClickHouse, Prometheus/Thanos). * Superior communication, collaboration, and diplomatic skills-ability to align cross-functional teams around architectural standards and lead technical decisions with empathy and clarity. Preferred Qualifications * Massive Scale Experience: Prior experience in high-performance computing (HPC) or handling global installations processing petabytes of telemetry per day. * Customer-Facing Observability: Experience building external or customer-exposed observability tools, APIs, and analytics dashboards with strict latency and access-control constraints. * Network & Systems Deep Dive: Deep operational understanding of high-performance Linux kernel tuning, ebpf, eBPF-based profiling, or low-level network optimization. * Multi-Cloud Expertise: Prior experience running cloud-agnostic platform infrastructure across AWS, Azure, and GCP simultaneously. ## Description At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don't just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset - who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform, Snowflake's Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems., * Architect & Scale Distributed Infrastructure: Lead the technical vision, design, and execution of Snowflake's External Observability platform capable of handling trillions of events per day across multi-cloud deployments (AWS, Azure, GCP). * Define Observability Standards: Standardize telemetry generation (metrics, logs, traces, and events) across all Snowflake application engineering teams, ensuring consistent schema, zero-data-loss ingestion, and optimized storage access. * Build High-Throughput Engines: Write scalable, reliable, and testable backend services to process time-series data, high-cardinality metrics, and distributed traces at massive scale. * Automate Infrastructure Lifecycle: Practice infrastructure-as-code (IaC) using tools like Terraform to deliver self-healing, automated telemetry pipelines and dynamic monitoring topology. * Drive Platform Adoption & Diplomacy: Collaborate closely with Application Engineering, Security, and Customer Support teams to make systems measurable, translate telemetry into actionable customer-facing insights, and resolve cross-organizational technical dependencies. * Technical Leadership & Mentorship: Drive engineering excellence through rigorous design reviews, technical roadmapping, performance tuning, and mentoring engineers across the broader organization. * Incident Escalation & Root Cause Analysis: Serve as an expert troubleshooter for critical, complex system failure modes across large distributed clusters, conducting deep-dive post-mortems and building automation to eliminate repeat incidents. ## Related Videos - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025)