Senior Site Reliability Engineer (Hardware Automation)

Jobgether
Madrid, Spain
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Bash Shell Continuous Integration Data Centers Distributed Systems Fault Tolerance Python (Programming Language) Linux System Administration Automation of Marketing Performance Tuning Reliability Engineering AI Infrastructure
+3 more
System Availability Backend Hardware Infrastructure

Job description

Senior Site Reliability Engineer (Hardware Automation) based in Spain.This is an exciting opportunity to build and operate the automation platforms that enable reliable, large-scale hardware infrastructure operations.In this role, you will develop tools and systems that reduce manual processes, improve operational visibility, and increase the reliability of critical infrastructure environments.Working as part of a product-focused engineering team, you will own solutions from initial design through deployment and long-term maintenance.You will collaborate closely with hardware engineers, data center operations teams, and infrastructure specialists to solve complex technical challenges.This position offers the chance to work with modern technologies, improve large-scale systems, and directly contribute to the evolution of AI infrastructure.AccountabilitiesBuild and maintain automation platforms and internal tooling that support large-scale hardware infrastructure operations.Develop solutions that eliminate manual processes, reduce operational risks, and improve visibility and control across infrastructure systems.Ensure high availability, fault tolerance, and reliable operation of critical services supporting hardware environments.Design, implement, and improve CI/CD processes to enable efficient and dependable software delivery.Troubleshoot complex infrastructure issues involving hardware, software, and networking components.Collaborate with hardware infrastructure teams, data center operations, and engineering stakeholders to gather requirements and deliver scalable solutions.Own the complete lifecycle of engineering projects, including design, development, deployment, monitoring, and ongoing reliability improvements.Continuously optimize system performance and operational workflows through automation and engineering best practices.RequirementsStrong experience working with Linux-based systems and production infrastructure environments.Proficiency in Python and Bash scripting for automation, tooling, and system management.Demonstrated ability to diagnose and resolve complex technical issues across hardware, software, and networking layers.Strong analytical and problem-solving skills with a focus on improving reliability, scalability, and performance.Experience designing, developing, or operating reliable infrastructure services and automated systems.Good communication skills and professional working proficiency in English.Ability to collaborate effectively within international, cross-functional engineering teams.Experience with backend development or building high-load distributed systems is considered an advantage.BenefitsCompetitive compensation package based on experience and expertise.Flexible working environment with a strong focus on autonomy and ownership.Opportunities for professional growth, continuous learning, and career development.Opportunity to contribute to impactful AI infrastructure projects.Collaborative and innovative international team culture.Exposure to challenging technical problems involving large-scale systems and automation.Environment built around trust, meaningful impact, and the opportunity to shape future technologies.#J-*****-Ljbffr

Requirements

monitoring, and ongoing reliability improvements.Continuously optimize system performance and operational workflows through automation and engineering best practices.RequirementsStrong experience working with Linux-based systems and production infrastructure environments.Proficiency in Python and Bash scripting for automation, tooling, and system management.Demonstrated ability to diagnose and resolve complex technical issues across hardware, software, and networking layers.Strong analytical and problem-solving skills with a focus on improving reliability, scalability, and performance.Experience designing, developing, or operating reliable infrastructure services and automated systems.Good communication skills and professional working proficiency in English.Ability to collaborate effectively within international, cross-functional engineering teams.Experience with backend development or building high-load distributed systems is considered an advantage.BenefitsCompetitive compensation package based on experience and expertise.Flexible working environment with a strong focus on autonomy and ownership.Opportunities for professional growth, continuous learning, and career development.Opportunity to contribute to impactful AI infrastructure projects.Collaborative and innovative international team culture.Exposure to challenging technical problems involving large-scale systems and automation.Environment built around trust, meaningful impact, and the opportunity to shape future technologies.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · WWC 2022

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

3:43 min

Designing a reliable technical hiring and interview process

Andreas Klinger Andreas Klinger +1 · Coffee With Developers

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · WWC 2024

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all