Principal SRE (AWS, Azure, Terraform, Kubernetes)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
- Work with Architecture, Product Owners, and development teams to elicit and define the practical needs and requirements for products being migrated to cloud.
- Design and implement scalable, highly available, and resilient infrastructure solutions to agreed standards.
- Innovate and enhance tooling for monitoring, alerting, incident management, and automation.
- Develop and maintain comprehensive documentation for infrastructure and processes, ensuring compliance with industry standards.
- Actively participate in incident management, driving rapid recovery, root cause analysis, and continuous improvement.
- Identify and implement strategies to future-proof infrastructure for growth and increased demand.
- Use AI tools effectively as a force multiplier and encourage others to do the same where appropriate.
- Assist in the career development of colleagues, acting as a role model and encouraging best practice.
Technologies:
- AI
- AWS
- Ansible
- Azure
- Cloud
- DevOps
- Docker
- Incident Management
- Support
- Kanban
- Kubernetes
- OOP
- Puppet
- Terraform
- REST
More:
We are Fourth, and in July 2019 we joined forces with HotSchedules to become a global leader in end-to-end restaurant and hospitality management technology solutions. Our combined SaaS suite covers scheduling, time and attendance, applicant tracking, training, inventory management and procurement, HR and benefits, and payroll services, supporting customers in 120,000 locations worldwide. We have a dedicated, unified team across offices in the US, UK, Bulgaria, China, Australia, and the UAE. We are seeking an experienced and pragmatic Principal SRE to join our worldwide team as we accelerate adoption of automated, highly reliable, zero-downtime infrastructure pipelines and move further into public cloud, particularly Azure. We offer a smart, fun, and talented team, flexible working, hybrid working environments, 25 basic holidays with the option to grow to 30 with service plus your birthday off and bank holidays, wellness activities, laptop and equipment, healthcare expense claim tools, certification support, annual meetups, an enhanced parenting scheme, cycle to work scheme, season ticket loan, pension and life insurance options, and on-demand pay tools. We are an equal opportunity employer.
Requirements
- Demonstrable experience in Site Reliability Engineering, Platform Engineering, DevOps, or a similar role, with experience mentoring others.
- Proven track record of designing, implementing, and managing large-scale, highly available, and scalable infrastructure.
- Relevant and recent experience with our main tech stack: Terraform, Configuration Management (Chef ideally, but we will consider Ansible or Puppet), Kubernetes (cloud based), and Docker (Kubernetes or AWS ECS Fargate).
- Extensive cloud experience, ideally with Azure and AWS.
- Programming and scripting skills, with a solid understanding of OOP principles.
- Strong analytical and problem-solving skills.
- Outstanding communication skills, with the ability to convey complex technical concepts to non-technical stakeholders.
- Ability to work in a fast-paced, dynamic environment and manage multiple priorities.
- Experience working in Agile or Kanban.
- A natural and positive team player.
- Bachelors degree in computer science, Engineering, or a related field.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where To Find Software Engineering Jobs
The 12 Best Jobs for Software Engineers
Is Software Engineering Over-Saturated?
Software Engineer Salary London