> Markdown version of [/jobs/ext/3095677-principal-site-reliability-engineer-kubernetes-required-hybrid](https://www.wearedevelopers.com/jobs/ext/3095677-principal-site-reliability-engineer-kubernetes-required-hybrid). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer (Kubernetes Required) - Hybrid - **Company:** Factset Research Systems Inc. - **Location:** Norwalk, CT, United States - **Experience:** Expert - **Salary:** $190,000.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Packaging, Microsoft Azure, Bash Shell, Cloud Computing, Computer Programming, Continuous Integration, DevOps, Github, Monitoring of Systems, Python (Programming Language), Open Source Technology, Performance Tuning, Reliability Engineering, Ansible, Prometheus, Pulumi, Scripting, Google Cloud, Grafana, Kubernetes, Information Technology, Puppet, Terraform, Golang - **Published:** September 26, 2026 - **Apply:** https://www.careerjet.com/job/us068283b68a5e48ce32aad2b8d5c317fc/eaa ## About the Role * 8+ years' experience ensuring the reliability, scalability, and performance of our systems and services Required Technical Skills Kubernetes (Required) * Hands-on experience deploying, managing, and troubleshooting workloads in Kubernetes * Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and Ingress * Experience with Kubernetes cluster management and administration * Familiarity with Helm for application packaging and deployment * Understanding of Kubernetes networking, storage, and security best practices Additional Technical Skills * Cloud Platforms: (e.g. AWS, GCP, Azure) * CI/CD Tooling: (e.g. GitHub Actions, ArgoCD, Harness) * Monitoring & Observability: (e.g. Prometheus, Grafana, Coralogix, OpenTelemetry) * Infrastructure as Code: (e.g. Terraform, Pulumi) * Config Management: (e.g. Ansible, Puppet, Chef) * Programming/Scripting: (e.g. Python, Go, Bash) Soft Skills & General Requirements * Strong problem-solving and analytical skills with a methodical approach to troubleshooting * Excellent communication skills with the ability to collaborate across technical and non-technical teams * A proactive mindset with a focus on automation and continuous improvement * Ability to work effectively under pressure, particularly during incident response * Commitment to a blameless culture and continuous learning Nice to Have * Experience contributing to open-source projects * Familiarity with SRE principles as defined by the Google SRE handbook * Previous experience in a DevOps or Platform Engineering role Education: * Bachelor's degree in computer science or relevant degree. ## Description We are looking for a skilled and motivated Principal Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices. Key Responsibilities * Monitor, maintain, and improve the reliability and availability of production systems * Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrence * Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) * Collaborate with development teams to build reliability into services from the ground up * Design and implement automation to reduce toil and improve operational efficiency * Participate in an on-call rotation to support critical systems * Contribute to capacity planning and performance optimization efforts * Document systems, processes, and runbooks to support the wider team ## Related Videos - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)