Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+32 more
Job description
We are looking for a Senior Site Reliability Engineer (SRE) to help drive the reliability, scalability, and security of enterprise platforms across Windows, Linux, and cloud-native environments. In this role, you will support the transformation from traditional application support models to modern platform engineering practices. You will leverage expertise in Google Cloud Platform (Google Cloud Platform), automation, containerization, and infrastructure engineering to build resilient systems that enable business-critical applications at scale.
As a member of the Site Reliability Engineering team, you will collaborate with software engineers, infrastructure teams, and security partners to improve platform availability, operational efficiency, and cloud adoption., * Design, implement, and maintain highly available, scalable, and secure production systems across Windows, Linux, and Google Cloud Platform environments.
- Build and support containerized platforms using Kubernetes (GKE) and Docker.
- Develop and manage infrastructure through Infrastructure-as-Code (IaC) tools including Terraform and Ansible.
- Improve platform performance, reliability, and capacity through proactive engineering and optimization.
Automation and Observability
- Create automation solutions to reduce operational overhead and improve incident response efficiency.
- Develop monitoring, alerting, and observability capabilities using SLIs, SLOs, Prometheus, Grafana, and Google Cloud Operations Suite.
- Implement telemetry and performance metrics across hybrid and cloud environments.
Incident Management and Resilience
- Lead incident response efforts, perform root cause analyses, and facilitate post-incident reviews.
- Design and implement self-healing systems and automated remediation workflows.
- Drive continuous improvement initiatives to enhance system reliability and operational excellence.
Security and Compliance
- Partner with Information Security teams to implement security best practices, vulnerability management, and compliance requirements.
- Integrate security controls into cloud platforms, infrastructure, and CI/CD pipelines.
- Support identity management, encryption, access controls, and policy enforcement across enterprise environments.
Cross-Functional Collaboration
- Work closely with software developers, application owners, and infrastructure engineers to build reliable cloud-native solutions.
- Develop and maintain technical documentation, operational procedures, and runbooks.
- Serve as a trusted technical advisor on platform reliability and operational best practices.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
- 5+ years of experience in Software Engineering, Site Reliability Engineering, Systems Engineering, or related technical roles.
- 3+ years of hands-on experience supporting production Windows and/or Linux environments.
- Experience administering and troubleshooting large-scale production systems.
- Experience with infrastructure automation and scripting using PowerShell, Python, Shell, or similar languages.
Preferred Qualifications
- Experience with Google Cloud Platform (Google Cloud Platform), including GKE, IAM, Cloud Functions, Cloud Monitoring, and related services.
- Experience with container orchestration technologies, including Kubernetes and Docker.
- Experience with Infrastructure-as-Code tools such as Terraform and Ansible.
- Strong understanding of Linux system administration and hybrid cloud architectures.
- Knowledge of Active Directory, DNS, DHCP, and Windows security concepts.
- Experience implementing CI/CD pipelines using tools such as GitLab CI, Jenkins, or similar platforms.
- Familiarity with ITIL practices, change management processes, and incident management frameworks.
- Experience with ServiceNow, load balancers, certificate management, and endpoint security solutions.
- Industry certifications such as CISSP, CompTIA Security+, or Google Professional Cloud Security Engineer.
- Experience working within financial services or other highly regulated industries.
Additional Information
- Participation in on-call rotations, including weekends and holidays as business needs require.
- This position requires strong problem-solving skills, a customer-focused mindset, and the ability to thrive in a collaborative, fast-paced environment.
What You’ll Bring
You are a proactive engineer with a passion for reliability, automation, and cloud technologies. You enjoy solving complex operational challenges, improving system resilience through engineering solutions, and enabling teams to deliver reliable services at scale.
Skills: Google Cloud Platform (Google Cloud Platform), Kubernetes, Docker, Terraform, Ansible, Windows Administration, Linux Administration, Site Reliability Engineering (SRE), Automation, CI/CD, Observability, Prometheus, Grafana, Cloud Operations, Python, PowerShell, Security, Infrastructure Engineering.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Data Engineer Salary UK
Is Software Engineering Over-Saturated?
Where to Find Entry-Level Software Engineering Jobs