Platform Operations Engineer (DevOps - Ops Focus)

Data-driven Ionos Group
Barcelona, Spain
7 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Agile Methodology Airflow Amazon Web Services Command-Line Interface Cloud Computing Computer Networks Data Infrastructure Software Debugging Linux Domain Name System (DNS) Apache Hive Log Analysis
+15 more
Networking Basics Open Source Technology OpenShift Reliability Engineering Prometheus TCP/IP Data Logging Load Balancing DevOps Tools - Open-source Grafana Mttr Containerization Kubernetes Terraform Docker

Job description

Currently, we are building our new cloud Big Data platform completely based on the IONOS cloud infrastructure. This means you will work with up-to-date, non-commercial technologies and have the unique opportunity to operate on our own internal cloud ecosystem, completely independent of the usual hyperscalers (AWS/GCP)., * Platform Operations & Stability: Take full responsibility for running, maintaining, and maintaining the operational health of our Kubernetes-based cloud Big Data platform.

  • Troubleshooting & L2/L3 Incident Response: Act as the primary point of contact for platform users (data engineers, analysts, internal teams). You will dive deep into application logs, investigate container failures, analyze pod behavior, and diagnose performance bottlenecks or process crashes.
  • User Application Operations: Support and maintain internal user applications running on top of Docker and Kubernetes. Ensure deployed services are healthy, scalable, and resilient.
  • Root Cause Analysis (RCA): Apply an investigation-first mindset to identify why platform or application failures occur, preventing recurring incidents rather than just applying temporary fixes.
  • Agile Collaboration & Support: Work in 2-week sprints, addressing operational tickets, incidents, and platform health tasks while keeping team members and internal stakeholders aligned.
  • Continuous Operational Improvement: Continuously refine operational runbooks, monitoring setups, and incident response definitions to improve mean time to resolution (MTTR).

Requirements

We are looking for a troubleshooting-minded Senior Engineer-someone who thrives when diagnosing complex system problems, loves digging under the hood of containerized environments, and takes pride in keeping platforms rock-solid.

  • Experience: Minimum 5+ years of proven experience in Systems Operations, L2/L3 Application Support, Site Reliability Engineering (SRE), or Infrastructure Operations.
  • Kubernetes & Docker (Core): Deep hands-on experience operating, debugging, and troubleshooting containerized workloads in Kubernetes and Docker environments (e.g., pod lifecycle, ingress/service issues, resource limits, volume mounts).
  • Troubleshooting Mindset: Exceptional diagnostic skills. You enjoy reading logs, analyzing metrics, tracing application failure modes, and asking the right questions to solve user-reported problems.
  • Linux Fundamentals: Confident command-line usage, high-level handling of Linux systems, process management, log analysis, and system recovery.
  • Operational DevOps Tooling: Practical experience managing deployments and platform states using tools like ArgoCD, Helm, or Terraform from an operations perspective.
  • Observability & Open Source Stack: Hands-on experience using monitoring and logging stacks (e.g., Prometheus, Grafana) to diagnose issues. Familiarity with applications like Airflow, Trino, Hive, or OpenShift is a strong plus.
  • Networking Basics: Practical understanding of network concepts (DNS, TCP/IP, load balancers, network policies) to troubleshoot connection or traffic routing issues.
  • Mindset: Strong user-empathy, autonomy to take ownership of production issues, and a supportive team player attitude focused on knowledge sharing.
  • Communication: Good English skills (written and spoken) to collaborate effectively with platform users and international teams.

Benefits & conditions

  • Flexible compensation package
  • Discounts on Arsys products
  • Discounts on technology brands
  • Challenging technical environment with real high-performance system problems
  • Access to advanced AI tools with no usage restrictions
  • Close collaboration with senior profiles and a strong culture of continuous improvement
  • Personalized training and development plans
  • Biannual team events

About the company

Our department’s vision is to build a Data-Driven IONOS Group. Our mission is to connect AI and BI, evolve our platforms, and empower our customers. As the central provider of the BI platform for Corporate, CTO topics, and state-of-the-art AI solutions, we support data-driven decision-making all across the IONOS Group.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobleads.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all