Senior Infrastructure & Storage Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
Senior Infrastructure & Storage EngineerSummaryLocation:Barcelona (Hybrid)Rate:NegotiableDuration:6 MonthsAbout the ClientOur client is a global leader in aviation technology, delivering innovative IT and communications solutions that connect airlines, airports, aircraft and governments worldwide. Operating in over 200 countries and territories, they play a critical role in enabling the safe, secure and efficient movement of millions of passengers every day.About the RoleThe Senior Infrastructure & Storage Engineer owns the implementation, integration, testing, and engineering lifecycle of on-premises compute, Linux, distributed storage, and related data-centre infrastructure.The role translates approved architecture into reliable, resilient, supportable platforms and provides deep engineering support for complex infrastructure incidents.This is a senior, hands-on role with particular emphasis on bare-metal infrastructure, Ceph, MinIO, Linux, Kubernetes storage integration, capacity, backup and recovery, and operational knowledge transfer. The engineer works closely with Platform Architecture, Operations, Kubernetes & Platform Engineering, Network Engineering, Security, vendors, and delivery teamsKey Responsibilities:On-Premises Infrastructure Engineering* Build, configure, troubleshoot, and lifecycle-manage physical and virtual servers across heterogeneous data-centre environments.* Engineer and support CentOS, RHEL, or equivalent Linux platforms, including operatingsystem, service, package, certificate, disk, memory, process, and performance issues.* Execute data-centre refreshes, hardware deployments, firmware and operating-system changes, and replacement activities in accordance with approved designs.* Validate compute, network, storage, power, capacity, resilience, monitoring, and support dependencies before production acceptance.* Support selected Windows infrastructure where required.Distributed Storage and Data Protection* Administer and troubleshoot Ceph, including cluster health, OSDs, monitors, placement groups, pools, capacity, performance, replication, and failure domains.* Engineer and support distributed MinIO deployments, object-storage availability, capacity, healing, replication, and recovery.* Diagnose storage latency, degraded redundancy, disk and node failures, data-path issues, and capacity risks using an evidence-based approach.* Plan, document, and test backup, restoration, disaster recovery, and data-protection procedures; distinguish platform redundancy from backup.* Coordinate specialist and vendor escalation for high-risk or complex storage problemsKubernetes Infrastructure Integration* Support bare-metal Kubernetes nodes and their operating-system, hardware, network, containerruntime, and storage dependencies.* Engineer and troubleshoot Kubernetes persistent storage, CSI drivers, StorageClasses, persistent volumes, mounts, and Ceph-backed workloads.* Collaborate with the Kubernetes & Platform Engineer during cluster upgrades, node maintenance, capacity changes, recovery testing, and stateful workload incidents.* Understand core Kubernetes concepts sufficiently to diagnose whether failures originate in the workload, node, network, CSI, storage, or underlying infrastructure.Automation, Monitoring, and Capacity* Automate repeatable infrastructure provisioning and configuration using Terraform, Ansible, Bash, Python, or equivalent tools.* Use Git-based workflows so infrastructure changes are reviewable, auditable, and reproducible.* Implement Prometheus metrics, dashboards, and actionable alerts for hosts, hardware, Ceph, MinIO, storage paths, capacity, and backup health.* Track utilization, saturation, growth, hardware health, and redundancy; implement shortand mediumterm capacity actions.* Integrate infrastructure logs and telemetry with platforms such as New Relic and Elasticsearch where appropriate.Testing, Handover, and Engineering Support* Plan and execute functional, integration, resilience, failover, capacity, backup, restoration, upgrade, and rollback testing.* Define acceptance criteria, evidence, abort conditions, maintenance procedures, and safe recovery paths.* Provide Level 3 support, lead root-cause analysis, and implement permanent corrective actions for complex infrastructure incidents.* Produce practical SOPs, runbooks, diagrams, asset and dependency records, maintenance procedures, and escalation paths.* Conduct hands-on knowledge transfer and ensure routine operations can be performed safely without dependency on one engineer.What we are looking forRequired SkillsStrong recent, hands-on experience engineering and troubleshooting production on-premises infrastructure.* Advanced Linux systems administration and troubleshooting capability.* Strong experience with storage systems and data-protection concepts, including replication, quorum, failure domains, capacity, performance, backup, and recovery.* Practical production experience with Ceph or a closely comparable distributed storage platform.* Experience working with physical servers, disks, controllers/HBAs, firmware, operating systems, virtualization, and data-centre dependencies.* Experience automating infrastructure through Ansible, Terraform, Bash, Python, or equivalent technologies.* Experience implementing monitoring, alerting, capacity management, and operational procedures.* Strong incident troubleshooting, root-cause analysis, risk management, and vendor escalation skills.* Ability to produce clear documentation and transfer specialist knowledge effectively.Desirable Experience* Direct administration of Ceph in production, including recovery from degraded states and performance investigations.* Distributed MinIO on dedicated servers or bare-metal infrastructure.* Kubernetes, Rancher, RKE2/RKE, CSI, and persistent-volume integration.* Prometheus, New Relic, Elasticsearch, or comparable observability platforms.* MariaDB, PostgreSQL, or other infrastructure database administration.* CentOS/RHEL, virtualization, selected Windows infrastructure, and hybrid Azure integration.* Business-continuity and disaster-recovery exercises across multiple data centres.Candidates do not need to be application-platform or cloud specialists. Depth in Linux, on-premises infrastructure, distributed storage, safe recovery, and operational knowledge transfer is more important than matching every desirable product.
Requirements
Strong recent, hands-on experience engineering and troubleshooting production on-premises infrastructure.
- Advanced Linux systems administration and troubleshooting capability.
- Strong experience with storage systems and data-protection concepts, including replication, quorum, failure domains, capacity, performance, backup, and recovery.
- Practical production experience with Ceph or a closely comparable distributed storage platform.
- Experience working with physical servers, disks, controllers/HBAs, firmware, operating systems, virtualization, and data-centre dependencies.
- Experience automating infrastructure through Ansible, Terraform, Bash, Python, or equivalent technologies.
- Experience implementing monitoring, alerting, capacity management, and operational procedures.
- Strong incident troubleshooting, root-cause analysis, risk management, and vendor escalation skills.
-
Ability to produce clear documentation and transfer specialist knowledge effectively. Desirable Experience
- Direct administration of Ceph in production, including recovery from degraded states and performance investigations.
- Distributed MinIO on dedicated servers or bare-metal infrastructure.
- Kubernetes, Rancher, RKE2/RKE, CSI, and persistent-volume integration.
- Prometheus, New Relic, Elasticsearch, or comparable observability platforms.
- MariaDB, PostgreSQL, or other infrastructure database administration.
- CentOS/RHEL, virtualization, selected Windows infrastructure, and hybrid Azure integration.
- Business-continuity and disaster-recovery exercises across multiple data centres. Candidates do not need to be application-platform or cloud specialists. Depth in Linux, on-premises infrastructure, distributed storage, safe recovery, and operational knowledge transfer is more important than matching every desirable product.
About the company
Our client is a global leader in aviation technology, delivering innovative IT and communications solutions that connect airlines, airports, aircraft and governments worldwide. Operating in over 200 countries and territories, they play a critical role in enabling the safe, secure and efficient movement of millions of passengers every day.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Is Software Engineering Over-Saturated?
Data Engineer Salary UK
Why Upskilling And Reskilling is Important For Developers