> Markdown version of [/jobs/ext/3144046-senior-staff-software-engineer-data-platform-kubernetes-distributed-systems-federal](https://www.wearedevelopers.com/jobs/ext/3144046-senior-staff-software-engineer-data-platform-kubernetes-distributed-systems-federal). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal - **Company:** ServiceNow - **Location:** Kirkland, WA, United States - **Experience:** Expert - **Salary:** $149,800.0 - $262,200.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Code Review, Continuous Integration, Data as a Services, Data Infrastructure, Relational Databases, Distributed Systems, Failover, Identity and Access Management, Key Management, PostgreSQL, Queueing Systems, RabbitMQ, Redis, Service Design, Software Engineering, Data Streaming, Google Cloud, Amazon ElastiCache, Istio, System Availability, Kubernetes, Infrastructure Automation Frameworks, Apache Kafka, Terraform, Servicenow - **Published:** September 29, 2026 - **Apply:** https://dejobs.org/x/x/B70F809A0362487E929CA71BC2C1D8FD/job/ ## About the Role * Experience leveraging or critically thinking about how to integrate AI into engineering and platform work - whether using AI-powered tooling, automating operational workflows, building agentic systems for fleet visibility and operations, or reasoning about AI's impact on how infrastructure is built and run. * 12+ years of software development experience with a Bachelor's degree; OR 8+ years with a Master's degree; OR 5+ years with a PhD; OR equivalent work experience. * 8+ years building production software, with solid experience operating distributed systems and running stateful workloads on Kubernetes at scale. * A track record of leading the design and delivery of at least one shared, multi-tenant infrastructure service used broadly by other teams - a relational database service (Postgres or similar), a message queue or streaming platform (Kafka, NATS, RabbitMQ, or similar), or a key-value/cache service (Redis/Valkey, etcd, or similar) - including its HA, failover, and operating model. * Deep understanding of high availability and failure handling: replication topologies, leader election, quorum and consensus, split-brain avoidance, backup/restore and point-in-time recovery, and designing to explicit RPO/RTO and durability targets. * Strong system design skills - you can reason rigorously about consistency models, partitioning and rebalancing, capacity planning, tenant isolation, and noisy-neighbor mitigation, and communicate the trade-offs clearly to engineers and stakeholders. * Hands-on experience with at least one major hyperscaler (AWS, Azure, GCP), including its core compute, networking, storage, and IAM primitives. * Strong working knowledge of containers and Kubernetes (including stateful primitives: StatefulSets, CSI/persistent storage, PDBs, topology spread), CI/CD and GitOps-based delivery, and infrastructure-as-code. * Strong programming skills in Go (or strong systems-language skills with a willingness to work primarily in Go)., * Experience building or extending Kubernetes operators that manage stateful systems (e.g., CloudNativePG, Zalando/Crunchy Postgres operators, Strimzi, Redis/Valkey operators), and opinions on when to adopt versus build. * Deep expertise in one of the target systems: Postgres internals (WAL, streaming/logical replication, vacuum, connection pooling with PgBouncer/PgCat, major-version upgrades); Kafka/NATS (partitioning, ISR/replication, exactly-once semantics, consumer scaling); or Redis/Valkey (cluster mode, persistence, eviction, hot-key handling). * Experience running data services across multiple regions or clusters - cross-region replication, failover orchestration, and DR testing. * Experience with zero-downtime upgrades, schema/data migrations, and fleet-wide rollouts of stateful services. * Experience with observability and SLOs for stateful systems - replication lag, saturation, tail latency, error budgets - and with capacity and cost management for shared infrastructure. * Experience designing multi-tenancy: quotas, isolation, chargeback/showback, and self-service provisioning APIs or CRDs. * Experience with container networking (CNI) and/or service mesh, and with workload identity, mTLS, and secrets management as applied to data services. * Experience with managed Kubernetes (EKS/AKS/GKE), managed data services (RDS/Aurora, Cloud SQL, MSK, ElastiCache), and infrastructure-as-code tools such as Terraform or Crossplane. ## Description We are seeking a Senior Staff Software Engineer (IC5 Level) to join our Data Platform Engineering organization. You will be a hands-on technical lead who will lead t he design and delivery of shared, multi-tenant platform services - Postgres, queueing/streaming, and key-value stores - on Kubernetes, and serves as the technical anchor for their availability, resilience, and operating model. What you get to do in this role: * You will lead the design and delivery of shared platform services - managed Postgres, queueing/streaming, and key-value/cache - that run on Kubernetes across a large global fleet and that product teams across the company depend on, owning them from design through production operation. * You will act as the technical lead on major initiatives within your team, breaking down ambiguous problems - "offer HA Postgres as a service to every cluster," "make our queueing tier survive a zone loss" - into clear, executable designs with explicit availability, durability, and cost targets. * You will define the HA, failover, disaster-recovery, and multi-region architecture for the services you own, and the Kubernetes-native automation (operators, controllers, CRDs, self-service APIs) that provisions, scales, upgrades, and fails them over without human intervention. * You will partner with principal and distinguished engineers to align your work with the broader platform architecture and standards, and with product teams to set consumption contracts, tenancy models, and SLOs for shared services. * You will identify technical risks early and drive them to resolution, with a strong focus on reliability, scalability, and operability - leading failure-mode analysis, game days, and post-incident reviews for stateful systems. * You will spend significant time hands-on - designing, coding, and reviewing the core systems your team builds, such as operators, controllers, infrastructure automation, and platform services. * You will mentor mid-level and junior engineers and raise the engineering bar through code reviews, design feedback, and pairing - particularly around distributed-systems and data-service design., We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here (https://careers.servicenow.com/life-at-servicenow#workpersonas) . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service., For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. ## Related Videos - [Adjusting Pod Eviction Timings in Kubernetes](https://www.wearedevelopers.com/videos/426-adjusting-pod-eviction-timings-in-kubernetes) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)