Principal Site Reliability Engineer (Kubernetes Required) - Hybrid
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
We are looking for a skilled and motivated Principal Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices. Key Responsibilities
- Monitor, maintain, and improve the reliability and availability of production systems
- Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrence
- Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Collaborate with development teams to build reliability into services from the ground up
- Design and implement automation to reduce toil and improve operational efficiency
- Participate in an on-call rotation to support critical systems
- Contribute to capacity planning and performance optimization efforts
- Document systems, processes, and runbooks to support the wider team
Requirements
- 8+ years’ experience ensuring the reliability, scalability, and performance of our systems and services
Required Technical Skills Kubernetes (Required)
- Hands-on experience deploying, managing, and troubleshooting workloads in Kubernetes
- Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and Ingress
- Experience with Kubernetes cluster management and administration
- Familiarity with Helm for application packaging and deployment
- Understanding of Kubernetes networking, storage, and security best practices
Additional Technical Skills
- Cloud Platforms: (e.g. AWS, GCP, Azure)
- CI/CD Tooling: (e.g. GitHub Actions, ArgoCD, Harness)
- Monitoring & Observability: (e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)
- Infrastructure as Code: (e.g. Terraform, Pulumi)
- Config Management: (e.g. Ansible, Puppet, Chef)
- Programming/Scripting: (e.g. Python, Go, Bash)
Soft Skills & General Requirements
- Strong problem-solving and analytical skills with a methodical approach to troubleshooting
- Excellent communication skills with the ability to collaborate across technical and non-technical teams
- A proactive mindset with a focus on automation and continuous improvement
- Ability to work effectively under pressure, particularly during incident response
- Commitment to a blameless culture and continuous learning
Nice to Have
- Experience contributing to open-source projects
- Familiarity with SRE principles as defined by the Google SRE handbook
- Previous experience in a DevOps or Platform Engineering role
Education:
- Bachelor’s degree in computer science or relevant degree.
Benefits & conditions
At FactSet, we celebrate difference of thought, experience, and perspective. Qualified applicants will be considered for employment without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, disability, protected veteran status or other characteristics protected by law. FactSet participates in E-Verify
About the company
FactSet creates flexible, open data and software solutions for over 200,000 investment professionals worldwide, providing instant access to financial data and analytics that investors use to make crucial decisions. At FactSet, our values are the foundation of everything we do. They express how we act and operate, serve as a compass in our decision-making, and play a big role in how we treat each other, our clients, and our communities. We believe that the best ideas can come from anyone, anywhere, at any time, and that curiosity is the key to anticipating our clients’ needs and exceeding their expectations., FactSet (NYSE:FDS | NASDAQ:FDS) helps the financial community to see more, think bigger, and work better. Our digital platform and enterprise solutions deliver financial data, analytics, and open technology to more than 8,200 global clients, including over 200,000 individual users. Clients across the buy-side and sell-side, as well as wealth managers, private equity firms, and corporations, achieve more every day with our comprehensive and connected content, flexible next-generation workflow solutions, and client-centric specialized support. As a member of the S&P 500, we are committed to sustainable growth and have been recognized among the Best Places to Work in 2023 by Glassdoor as a Glassdoor Employees’ Choice Award winner. Learn more at and follow us on and .
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Fully Remote Software Engineer Jobs
Highest Paying Tech Companies for Developers