> Markdown version of [/jobs/ext/1014583-principal-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1014583-principal-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - **Company:** Uniphore Inc. - **Location:** Palo Alto, CA, United States - **Experience:** Expert - **Salary:** $232,900.0 - $335,811.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, DevOps, Open Source Technology, Reliability Engineering, Software Engineering, Multi-Cloud, Kubernetes, Production Code, Api Design, Terraform, Webhooks - **Published:** June 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b3ed9c23d6183070 ## About the Role Do you have experience in Tooling?, * 10+ years in DevOps/SRE/Platform Engineering, with demonstrated Staff- or Principal-scope impact and a track record of transforming operational models. * Production Go: you write Go regularly, understand its concurrency model, and are comfortable owning Go services in production. * Kubernetes depth: operational expertise plus the ability to extend it - you understand the controller-runtime model and could write or maintain a Kubernetes Operator. * Cloud & infrastructure: expert-level AWS/GCP/Azure, Terraform, and multi-cloud architecture, with strong cost-optimization instincts. * Production excellence: deep incident management, RCA process, and on-call system design experience. * Software engineering fundamentals: API design, testing, observability instrumentation, and service lifecycle ownership - you treat internal tooling with the same rigor as customer-facing software. * Standards & documentation: strong technical writing; you create operational procedures teams can self-execute. * Architecture & planning : RFC/PRD review experience; you catch operational problems at design time. * Collaboration & coaching: you build team capability through tooling and knowledge transfer rather than doing the work for them. Nice to Haves: * Building Kubernetes Operators, controllers, or admission webhooks (controller-runtime, kubebuilder ). * Contributions to open-source infrastructure tooling. * AWS Solutions Architect Professional or equivalent GCP/Azure certifications. * Kubernetes certifications (CKA, CKAD, CKS). * Platform engineering, developer experience, or internal developer portals (Backstage, etc.). * GitOps patterns ( ArgoCD , Flux) and policy-as-code tooling (OPA, Kyverno ). ## Description We're looking for a Principal Site Reliability Engineer to join our Platform Engineering team - someone equally at home writing production Go as designing and operating cloud infrastructure. The highest-leverage work here isn't a runbook; it's the service that enforces the runbook automatically. You'll write Go that runs in production and multiplies your impact across hundreds of services. You'll build the standards, frameworks, automations, agentic workflows, and self-service capabilities that make engineering teams autonomous while maintaining enterprise-grade reliability and security. You won't just define standards - you'll implement them in code: a Kubernetes Operator that enforces service readiness criteria, a service that surfaces SLO health across the fleet, an internal platform service that automates task execution. You'll collaborate with feature teams as an expert advisor and standard-setter, helping them build operational maturity while you maintain oversight of our single/ multi-tenant, multi-cloud infrastructure. You'll be a bridge between software development and systems operations, focused on large-scale, resilient, automated infrastructure rather than daily firefighting. This is a senior individual-contributor role. You will not have direct reports. Your leadership is technical - exercised through architecture, production code, design reviews, and mentorship. This role participates in our on-call rotation, which covers all production systems. As a Principal , you'll own the hardest escalations and use what you learn on-call to drive the architectural fixes that eliminate whole classes of incidents. Responsibilities: Invention: * Define and execute long-term architectural strategy for our multi-cloud platforms. * Lead hands-on implementation of critical infrastructure projects, focusing on reliability, automation, and performance at scale. * Own multi-year technical roadmaps that establish the vision for infrastructure scalability, reliability, security, and engineering velocity. Own: * Provide technical leadership through design reviews and code contributions; set technical direction, eliminate architectural barriers to execution, and drive toward simplicity. * Maintain end-to-end technical stewardship of your systems, keeping execution aligned with architectural vision and best practices. * Act as a key technical advisor to Engineering Leadership and Product Management, influencing the strategic direction of Uniphore . * Lead design reviews across Infrastructure with a focus on automation, scalability, and reliability, and align architectural roadmaps across teams. * Partner with Security to build secure-by-default systems and remediate weaknesses. * Own the reliability of the systems under your technical stewardship. * Create the technical clarity - vision, standards, and tooling - that lets feature teams build, own and operate their services . * Participate in fleet-wide on-call, owning critical escalations across all production systems and converting recurring failure modes into permanent architectural fixes. Teach: * Establish and evangelize design principles for reliable, secure, scalable systems. * Grow other engineers through technical mentorship, architectural guidance, and design review. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Beyond Webhooks: The Future of Scalable API Event Delivery](https://www.wearedevelopers.com/videos/100312-beyond-webhooks-the-future-of-scalable-api-event-delivery) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Building a Cloud Platform Where Everything is Just Another Kubernetes Resource](https://www.wearedevelopers.com/videos/100137-building-a-cloud-platform-where-everything-is-just-another-kubernetes-resource) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)