Site Reliability Engineer - Fixed Term Contract

Mmt
Laza, Spain
2 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Laza, Spain

Tech stack

Amazon Web Services (AWS)
Application Performance Management
Azure
Cloud Computing
Cloud Engineering
Program Optimization
Computer Programming
Fault Tolerance
Github
Monitoring of Systems
Python
Nagios
Software Architecture
Reliability Engineering
Ansible
Software Engineering
Systems Architecture
Datadog
Cloud Platform System
System Availability
Multi-Cloud
Reliability of Systems
Infrastructure as Code (IaC)
Cloudformation
Kubernetes
Infrastructure Automation Frameworks
Data Analytics
Cloudwatch
Terraform
Azure
Docker

Job description

The RoleWe are seeking an experienced Site Reliability Engineer to play a pivotal role in bridging the gap between software engineering and operations. This role emphasizes designing robust solutions, mentoring teams, and driving performance improvements for both internal and client systems through expertise in automation, scalability, and system reliability.As a Site Reliability Engineer, you will be responsible for owning the uptime and performance of critical infrastructure and applications while working closely with clients to align reliability goals with their business objectives.Key ResponsibilitiesSystem Reliability & PerformanceOwn the uptime and performance of critical infrastructure and applicationsDesign scalable, fault-tolerant architectures that meet business needs while maximizing operational efficiencyDefine and govern Non-Functional Requirements (NFRs) such as availability, performance, and maintainability for internal and client systemsProactively identify opportunities for system optimization, scalability, and cost reductionAutomation & InfrastructureDesign and implement automation for monitoring, incident response, and repetitive operational tasksImplement Infrastructure as Code (IaC) practices with tools like Terraform, ARM templates and CloudFormation for consistent cloud environment provisioningDesign and deploy containerized solutions using Docker and Kubernetes on cloud platformsSet up and manage CI/CD pipelines and frameworks tailored for cloud-native applications using GitHub Actions and Azure DevOpsCloud Infrastructure & OperationsParticipate in architectural decisions and implement cloud infrastructure solutions using Azure and AWS services, ensuring high availability and scalabilityManage and optimize cloud resources to improve performance, cost efficiency, and securityApply cloud-native best practices to secure and govern cloud environments, ensuring compliance with industry standardsIntegrate advanced monitoring and alerting tools (e.g., Datadog, CloudWatch, Application Insights) to maintain system observability in multi-cloud environmentsIncident Management & AnalysisLead incident response, conduct root cause analyses, and produce blameless postmortems to prevent future occurrencesCollaborate with development teams to integrate observability and performance metrics into the development lifecycleBuild and maintain executive-level and developer-centric dashboards to visualize key metricsTechnical ExpertiseProven experience in running and maintaining production systems with expertise in triaging and solving incidentsProficiency in automation and configuration management tools (e.g., Terraform, Ansible)Expertise in cloud platforms, particularly Azure and AWS, and their associated toolsStrong programming skills, with a primary focus on Python, for developing automation scripts, creating custom tooling, and optimizing operational workflowsExperience with modern observability platforms such as DatadogSkills & ExperienceA solid foundation in system architecture, with a focus on scalability and reliabilityExceptional problem-solving skills and a data-driven mindsetDesirable RequirementsExperience with container orchestration tools such as KubernetesFamiliarity with CI/CD pipelines and tools like GitHub Actions and Azure DevOpsKnowledge of security best practices in cloud and hybrid environments

Requirements

Proven experience in running and maintaining production systems with expertise in triaging and solving incidents Proficiency in automation and configuration management tools (e.g., Terraform, Ansible) Expertise in cloud platforms, particularly Azure and AWS, and their associated tools Strong programming skills, with a primary focus on Python, for developing automation scripts, creating custom tooling, and optimizing operational workflows Experience with modern observability platforms such as Datadog Skills & Experience A solid foundation in system architecture, with a focus on scalability and reliability Exceptional problem-solving skills and a data-driven mindset Desirable Requirements Experience with container orchestration tools such as Kubernetes Familiarity with CI/CD pipelines and tools like GitHub Actions and Azure DevOps Knowledge of security best practices in cloud and hybrid environments

Apply for this position