Senior Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
As a Site Reliability Engineer, you will work with Agile engineering teams to provide production insight into running and operating software at-scale in a globally distributed and highly available cloud based system.
You will guide the team to consider resiliency, scalability and operability implications in the choices they make during the development cycle to help foster an ownership in production mentality (“You build it, you own it”).
You will also own and develop our platform by implementing and championing GitOps principles. You will play a hands-on role helping the team meet technical, operational, schedule, and business requirements.
The ideal candidate will be a systems problem solver with a passion for crafting products that deliver incredible customer experiences, have deep experience with infrastructure, operational automation, data driven metrics collection, modern platform management, and a true desire to automate it rather than do it repeatedly., * Design, build, and maintain highly scalable and reliable infrastructure using Infrastructure-as-Code (IaC).
- Manage and optimize Kubernetes clusters, leveraging Helm and ArgoCD for efficient application deployment and lifecycle management.
- Administer and troubleshoot Linux-based systems, ensuring their performance, security, and availability.
- Work extensively with both GCP and AWS services, architecting and managing cloud-native solutions.
- Implement robust observability practices, including monitoring, logging and alerting to proactively identify and resolve issues, using Splunk and Grafana.
- Develop and maintain automation tools using languages such as Python and Go to streamline operations and improve efficiency.
- Provide production support, responding to incidents and outages, and participating in a 24/7 on-call rotation.
- Troubleshoot complex networking issues and optimize network performance.
- Engage in communications across all areas of the organization.
Requirements
- Strong experience with IaC.
- Required: Terraform
- Nice to have: Terragrunt and Ansible
- Deep understanding of Kubernetes and Helm.
- Extensive experience with Linux administration and troubleshooting.
- Hands-on experience with GCP and AWS services and cloud-native architectures.
- Experience in implementing observability solutions.
- Solid understanding of networking concepts and protocols.
- Ability to work both independently and collaboratively.
- Ability to lead and work on projects.
- Ability to multitask and adapt quickly to changing priorities.
- Excellent communication and problem-solving skills. We have an international team, so good communication skills are essential.
- Willingness to participate in a 24/7 on-call rotation.
Benefits & conditions
- Salary Range: £ 60,000 - 70,000 base salary per year
- Bonus Plan
Benefits and Perks:
- Regional specific competitive benefits
- Build your own Benefits (BYOB) perk
- Local events, team building, and development opportunities
About the company
Symphony is an AI-powered communication and technology company fueled by interconnected platforms: messaging, voice, directory and analytics.
Our end-to-end encrypted technologies enable over 1,400 institutions to accelerate AI impact, prioritize data security, navigate complex regulatory compliance and optimize business interactions.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
Find a Developer Job: 12 Best Job Sites For Developers
Dev Digest 120 - Apple and peers