Senior Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Job description
Experteer Overview As Senior SRE II at FYUL Platform Infrastructure, you will own and drive reliability and scalability across multi-account AWS, Kubernetes (EKS), and core platform services.You will design large-scale automation, set engineering standards, and mentor junior SREs, while balancing hands-on platform work with technical leadership.You’ll lead cost optimization, security, and observability improvements to enable product teams to ship reliably at global scale.This role offers meaningful impact through DevOps enablement and cross-team collaboration.Compensaciones / Beneficios - Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts using infrastructure as code - Operate and design Amazon EKS clusters with networking, storage, and scaling strategies for containerized workloads - Own core platform services like cloud networking, Kubernetes, databases, and messaging systems - Drive large-scale automation with Terraform/Terragrunt and GitOps (ArgoCD); establish standards across teams - Lead adoption of automation to reduce manual work and promote repeatability - Lead on-call and incident response; author runbooks, ADRs, and postmortems - Improve reliability and observability with Grafana/Prometheus/Loki/Tempo/Mimir for scalable systems - Mentor mid-level SREs and support onboarding; communicate complex concepts to engineers and non-technical stakeholders - Collaborate with product squads to represent Platform Infrastructure in cross-team initiatives - Drive security (IAM; encryption; securelogging) and FinOps for cost efficiency; audit spend and optimize resources - Contribute to platform/DevEx roadmap and self-service initiatives for other teamsResponsabilidades - Solid Linux systems administration and Python scripting - Strong AWS knowledge (EKS, IAM, VPC, RDS, S3, SQS); multi-account environments a plus - Hands-on Kubernetes (EKS) operations, Helm, CNI (Cilium), IPAM, container security (ECR) - Terraform (modules, state) and Terragrunt; GitOps with ArgoCD - Postgres/MySQL/MongoDB in production; Aurora experience a plus - CI/CD with Jenkins and/or GitHub Actions; blue/green and canary deployment experience - Grafana/Prometheus/Loki/Tempo/Mimir for observability; dashboarding, alerting,tracing - Incident management experience; on-call rotations, runbooks, postmortems - Knowledge of 12-Factor App principles and FinOps awareness - Nice to have: GCP, Kafka/AWS MSK, regulated environments, self-service platform experience - Languages: PHP, Node.Js; Infra: AWS, Kubernetes, Terraform, Helm, AtlantisRequisitos principales - remote work options - private health insurance - extra days off for wellbeing - annual learning opportunities - mentorship and internal meetups - office lunch in Riga
Requirements
Experteer Overview As Senior SRE II at FYUL Platform Infrastructure, you will own and drive reliability and scalability across multi-account AWS, Kubernetes (EKS), and core platform services. You will design large-scale automation, set engineering standards, and mentor junior SREs, while balancing hands-on platform work with technical leadership. You’ll lead cost optimization, security, and observability improvements to enable product teams to ship reliably at global scale. This role offers meaningful impact through DevOps enablement and cross-team collaboration.Compensaciones / Beneficios - Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts using infrastructure as code - Operate and design Amazon EKS clusters with networking, storage, and scaling strategies for containerized workloads - Own core platform services like cloud networking, Kubernetes, databases, and messaging systems - Drive large-scale automation with Terraform/Terragrunt and GitOps (ArgoCD); establish standards across teams - Lead adoption of automation to reduce manual work and promote repeatability - Lead on-call and incident response; author runbooks, ADRs, and postmortems - Improve reliability and observability with Grafana/Prometheus/Loki/Tempo/Mimir for scalable systems - Mentor mid-level SREs and support onboarding; communicate complex concepts to engineers and non-technical stakeholders - Collaborate with product squads to represent Platform Infrastructure in cross-team initiatives - Drive security (IAM; encryption; securelogging) and FinOps for cost efficiency; audit spend and optimize resources - Contribute to platform/DevEx roadmap and self-service initiatives for other teamsResponsabilidades - Solid Linux systems administration and Python scripting - Strong AWS knowledge (EKS, IAM, VPC, RDS, S3, SQS); multi-account environments a plus - Hands-on Kubernetes (EKS) operations, Helm, CNI (Cilium), IPAM, container security (ECR) - Terraform (modules, state) and Terragrunt; GitOps with ArgoCD - Postgres/MySQL/MongoDB in production; Aurora experience a plus - CI/CD with Jenkins and/or GitHub Actions; blue/green and canary deployment experience - Grafana/Prometheus/Loki/Tempo/Mimir for observability; dashboarding, alerting,tracing - Incident management experience; on-call rotations, runbooks, postmortems - Knowledge of 12-Factor App principles and FinOps awareness - Nice to have: GCP, Kafka/AWS MSK, regulated environments, self-service platform experience - Languages: PHP, Node.Js; Infra: AWS, Kubernetes, Terraform, Helm, AtlantisRequisitos principales - remote work options - private health insurance - extra days off for wellbeing - annual learning opportunities - mentorship and internal meetups - office lunch in Riga
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Find a Developer Job: 12 Best Job Sites For Developers
Is Software Engineering Over-Saturated?
Where To Find Software Engineering Jobs
The 12 Best Jobs for Software Engineers