Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+25 more
Job description
We are looking for a Senior Site Reliability Engineer (SRE) to help scale and modernize platform operations across Windows, Linux, and cloud-native environments. In this role, you will drive the transition from application-specific support to platform-wide reliability engineering, focusing on automation, scalability, and resilience.
You will leverage your expertise in Google Cloud Platform (Google Cloud Platform), container orchestration, and infrastructure automation to build systems that are reliable, secure, and performant across a diverse enterprise landscape. What You’ll Do Reliability & Cloud Infrastructure
- Design, build, and maintain highly available, scalable systems across Windows, Linux, and Google Cloud Platform environments
- Operate and support containerized applications using Kubernetes (GKE) and Docker
- Provision and manage infrastructure using Terraform, Ansible, and Google Cloud Platform-native tools
Automation & Observability
- Develop tools and automation to reduce manual effort and improve system reliability
- Define and implement SLIs/SLOs to drive service performance and reliability
- Build monitoring and alerting solutions using Prometheus, Grafana, and Google Cloud Platform Operations Suite
Incident Response & Resilience
- Lead incident management, root cause analysis, and postmortems
- Design and implement self-healing systems and automated remediation workflows
- Improve system resilience through proactive reliability engineering practices
Security & Compliance
- Partner with security teams to enforce infrastructure hardening and vulnerability management
- Integrate security controls into CI/CD pipelines and container platforms
- Implement IAM, encryption, and policy enforcement across cloud environments
Collaboration & Enablement
- Work cross-functionally with developers, infrastructure teams, and stakeholders
- Create documentation, runbooks, and operational best practices
- Enable teams to adopt reliable, scalable platform solutions
Requirements
- 3+ years of experience in Windows or Linux production support/administration
- 5+ years of software engineering experience or equivalent combination of work, training, or education
- Experience with cloud platforms (Google Cloud Platform preferred) and distributed systems
Preferred Qualifications
- Strong scripting skills (e.g., Python, PowerShell, Shell)
- Hands-on experience with Google Cloud Platform services (GKE, IAM, Cloud Functions, Cloud Monitoring)
- Expertise in Docker and Kubernetes
- Experience with Infrastructure as Code (Terraform, Ansible)
- Knowledge of Active Directory, DNS, DHCP, and Windows security
- Experience with CI/CD tools (GitLab CI, Jenkins)
- Familiarity with ITIL practices and change management processes
- Exposure to ServiceNow, load balancing, certificate management, endpoint security tools
- Security certifications (e.g., CISSP, Security+, Google Cloud Platform Professional Cloud Security Engineer)
- Experience working in financial services or regulated industries
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Where to Find Entry-Level Software Engineering Jobs
Software Engineer Salary London