Site Reliability Engineer Id53670

Agileengine
Madrid, Spain
21 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
2 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Artificial Intelligence Amazon Web Services Confluence JIRA Bash Shell Software as a Service DevOps Github Python (Programming Language) Reliability Engineering Prometheus Software Engineering
+10 more
Circleci Grafana Git Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Terraform Pagerduty Jenkins

Job description

AgileEngine is an Inc. ** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries.We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.Por favor, lea detenidamente la información de esta oferta de empleo para entender exactamente qué se espera de los posibles candidatos.WHY JOIN USIf you’re looking for a place to grow, make an impact, and work with people who care, we’d love to meet you!ABOUT THE ROLEWe are looking for a Middle SRE Operations Engineer to maintain reliability across a cloud-based SaaS platform.You’ll handle live incidents, improve observability, and reduce toil through automation using Kubernetes, Terraform, Grafana, and AWS.Hands?on, execution?focused, with real ownership across CI/CD pipelines, GitOps workflows, and on?call rotations.WHAT YOU WILL DOMonitor and support production and staging environments to ensure availability, performance, and stability.Respond to incidents, perform triage and root cause analysis, and contribute to remediation efforts.Participate in on-call rotations with defined SLAs.Handle operational requests from internal teams.Maintain and improve monitoring, alerting, dashboards, logs, and metrics.Support CI/CD pipelines, production releases, and GitOps workflows.Contribute to automation initiatives to reduce operational overhead.Maintain and improve Kubernetes?based infrastructure and containerized workloads.Support Infrastructure as Code practices and environment improvements.MUST HAVES2+ years of experiencein Site Reliability Engineering, DevOps, or Production Operations.Experience withAWSsupporting production environments.Experience supportingproduction SaaS applications.Strong understanding ofCI/CD systems(GitHub Actions, Jenkins, CircleCI).Experience withGitOps and Git fundamentals.Experience usingGitHub, Jira, and Confluence.Experience withKubernetes(EKS, kOps or similar).Experience withDocker and containerization.Experience withobservability tools(Grafana, Prometheus, Loki, PagerDuty).Proficiency inscripting(Bash, Python, or Go).Experience withInfrastructure as Code(Terraform, Helm).Ability to work within structured operational processes and SLAs.Strong written and verbal English communication skills.Self?driven with a growth mindset.NICE TO HAVESAWS certifications such as Solutions Architect, DevOps Engineer, or SysOps Administrator.Experience with multi?tenant SaaS environments.Experience working in globally distributed teams.Familiarity with ChatOps practices.Experience improving monitoring quality and reducing alert fatigue.PERKS AND BENEFITSProfessional growth: Mentorship, TechTalks, and personalized growth roadmaps.Competitive compensation: USD?based pay with education, fitness, and team activity budgets.Exciting projects: Modern solutions with Fortune 500 and top product companies.xcskxljFlextime: Flexible schedule with remote and office options.#J-***-Ljbffr

Requirements

2+ years of experiencein Site Reliability Engineering, DevOps, or Production Operations. Experience withAWSsupporting production environments. Experience supportingproduction SaaS applications. Strong understanding ofCI/CD systems(GitHub Actions, Jenkins, CircleCI). Experience withGitOps and Git fundamentals. Experience usingGitHub, Jira, and Confluence. Experience withKubernetes(EKS, kOps or similar). Experience withDocker and containerization. Experience withobservability tools(Grafana, Prometheus, Loki, PagerDuty). Proficiency inscripting(Bash, Python, or Go). Experience withInfrastructure as Code(Terraform, Helm). Ability to work within structured operational processes and SLAs. Strong written and verbal English communication skills. Self?driven with a growth mindset. NICE TO HAVES AWS certifications such as Solutions Architect, DevOps Engineer, or SysOps Administrator. Experience with multi?tenant SaaS environments. Experience working in globally distributed teams. Familiarity with ChatOps practices. Experience improving monitoring quality and reducing alert fatigue.

Benefits & conditions

Competitive compensation: USD?based pay with education, fitness, and team activity budgets. Exciting projects: Modern solutions with Fortune 500 and top product companies. xcskxlj Flextime: Flexible schedule with remote and office options. #J-*****-Ljbffr

About the company

Madrid, España

AgileEngine is an Inc. ** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all