Principal DevOps Engineer for Industrial AI Cloud

Deutsche Telekom Group
Madrid, Spain
2 days ago
Apply on www.recruit.net
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English, German

Tech stack

Artificial Intelligence Bash Shell Software as a Service Cloud Computing Computer Clusters System Configuration Continuous Integration Data Visualization Linux DevOps Github Infrastructure as a Service (IaaS)
+21 more
Python (Programming Language) Platform as a Service (PAAS) Performance Tuning Ansible Prometheus Software Configuration Management Scripting Delivery Pipeline Saltstack Large Language Models Grafana Gitlab Git AI Platforms Gitlab-ci Kubernetes Machine Learning Operations Hardware Infrastructure Virtual Agents Terraform Data Pipelines

Job description

Gain full access to exclusive job listings from leading companies worldwide.

  • Verified, High-Quality Jobs Only No ads, scams, or junk-just genuine opportunities.

  • Focus on Real Opportunities Explore thousands of open positions tailored to your lifestyle, including flexible remote jobs.

  • Exclusive Resume Review Receive expert feedback with personalized suggestions to enhance your resume., As DevOps Engineer Principal you will guide enterprise customers through onboarding, training, and early adoption of the AI platform. Your responsibility includes understanding customer requirements, supporting solution design, executing Proofs of Concept (PoCs), and ensuring smooth integration of customer workloads (LLMs, GPU compute, AI pipelines). You act as a trusted technical advisor, helping customers efficiently use their GPU clusters and AI toolchains. We are looking for a DevOps Engineer Principal position who will play a key role in helping our enterprise customers successfully adopt and scale our AI platform. In this position, you will guide clients through onboarding, training, and early-stage implementation of cutting-edge AI solutions. You’ll work closely with them to understand their technical and business needs, support the design of tailored architectures, and lead Proofs of Concept that validate real-world value. You will ensure smooth integration of complex workloads - from LLM deployment and GPU compute optimization to building end-to-end AI/ML pipelines. As a trusted technical advisor, you will empower customers to use their GPU clusters and AI toolchains efficiently, troubleshoot challenges, and adopt best practices that accelerate their AI journey.

What will you do?

  • Consult customers on all technical aspects related to GPU infrastructure and platform usage.
  • Lead onboarding and training, mentoring customer specialists on optimal usage of their GPU clusters and AI environments.
  • Design and implement PoCs, including environment setup, data processing pipelines, and deployment workflows.
  • Conduct requirement engineering, translating business needs into technical specifications.
  • Assist customers with performance optimization, troubleshooting, fine-tuning, and validation of delivered solutions.
  • Act as the key technical point of contact, coordinating cross-functional teams across infrastructure, networking, automation, security, and AI services.
  • Propose and develop automation concepts to improve services, processes, and operating models.
  • Ensure best practices in reliability, scalability, and security are applied across the customer lifecycle.
  • Support monitoring, observability, and capacity planning for AI workloads and GPU utilization., If you are looking for a new challenge, do not hesitate to send us your CV! Please send CV in English. Join our team! T-Systems Iberia will only process the CVs of candidates who meet the requirements specified for each offer. Solicitar ahora Guardar trabajo Automatically Get Matched to Devops Engineer Jobs Let our AI agent search and match you to the best jobs from across the web Try it now

Requirements

  • Have 3+ years of experience in the design and delivery of systems based on IaaS, PaaS and SaaS.
  • Have experience with GPU based infrastructure.
  • Possess solid knowledge of Kubernetes container-based technologies.
  • Have expert knowledge of scripting languages (Python/Bash).
  • Possess expert knowledge of Automation tools and automation deployments (Ansible/Salt-stack/Terraform/Helm).
  • Have expert knowledge of CI/CD in a Kubernetes environment and repository management.
  • Are expert in automation with Git (GitHub, GitLab) and CI/CD tools like GitHub Actions or GitLab CI/CD.
  • Have strong experience with Linux OS.
  • Have experience with monitoring and visualization tools (Grafana, Prometheus,…)
  • Speak English at B2 level (German is an advantage).

Other skills:

  • Good communication skills, analytical thinking, team cooperation, presentation skills, negotiation skills.
  • Project Management- Basic.
  • Leadership skills- Basic.
  • Quality management- Intermediate.

Benefits & conditions

Work environment & flexibility

  • International, dynamic and collaborative environment.
  • T-Social: social initiatives (sports, community, health, …).
  • Hybrid work model (remote/on-site).
  • Flexible working hours.

Growth & development

  • Customized training: access to Coursera to learn whatever you want, whenever you want.
  • Weekly language classes (English, Spanish & German).
  • International Mentoring Sessions & Experience Days.

Compensation & benefits

  • Flexible compensation plan (health insurance, meal vouchers, childcare, transport).
  • Telemedicine.
  • Life and accident insurance.
  • Social fund.

Wellbeing & time off

  • 26+ working days of vacation per year.
  • Free access to specialist services (medical, legal, wellness).
  • 100% salary coverage during medical leave.

About the company

T-Systems is part of the Deutsche Telekom Group, with around 30.000 employees worldwide. We create technology with purpose to generate a positive impact on society. We are looking for curious talent, eager to learn, take on challenges, and contribute ideas that transform our customers’ experience.

We trust people: we offer autonomy, continuous support, and a collaborative environment where you can grow without limits. We are one global team, guided by respect, integrity, and a passion for doing better every day., NVIDIA and Deutsche Telekom are jointly developing industrial AI cloud for Europe. This AI factory in Germany will host 10,000 GPUs across NVIDIA DGX B200 systems and RTX Pro Servers. Deutsche Telekom provides secure, sovereign and fast infrastructure, including data centers, operations, security, and AI solutions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.recruit.net
Prepare application

Good distractions

Loading talks and stories from around this role…