Senior Software Engineer (all genders)

Zalando SE
Berlin, Germany
28 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Batch Processing Continuous Integration Data Governance Programming Tools Distributed Systems Identity and Access Management Key Management Azure Machine Learning Data Streaming Systems Integration
+13 more
Cloud Platform System Feature Engineering Apache Spark Caching Containerization AI Platforms Kubernetes Infrastructure Automation Frameworks Apache Flink Deployment Automation Apache Kafka Machine Learning Operations Docker

Job description

Our ML Platform team builds the core ML platform capabilities powering Zalando’s AI-native experiences. We provide low-latency features, embeddings, real-time inference infrastructure, and scalable ML platform capabilities that enable applied science and product teams to deliver search, recommendations, personalization, forecasting, and emerging GenAI use cases.

Today, we operate Zalando’s central Feature Store and are evolving the next generation of Kubernetes-native AI runtime infrastructure, enabling scalable online serving, distributed GPU workloads, and self-service ML platform operations across the company.

As a Senior Software Engineer (ML Platform), you will play a key role in designing, building, and scaling these core ML infrastructure services. You’ll work hands-on with distributed systems, streaming pipelines, Kubernetes-native serving infrastructure, and platform automation, while also mentoring peers and contributing to engineering best practices across the team., Own the design and implementation of scalable real-time feature platforms, online serving infrastructure, and distributed ML runtime systems. Bring strong technical judgment to ensure our platform foundations are reliable, reusable, and operationally mature.

  • Deliver and maintain SLOs for feature freshness, data quality, online/offline consistency, and runtime reliability; implement monitoring, observability, and safe deployment practices.
  • Drive automation and self-service (IaC, GitOps, CI/CD), reusable deployment templates, and operational tooling that reduce friction and accelerate time-to-first-success for applied scientists and engineers. Contribute to reusable platform integrations and deployment automation that improve how ML systems interact with developer tooling and internal AI platform capabilities.

  • Implement identity and access management, secrets management, network isolation, and data governance built in from the start to ensure compliance and trustworthiness by default.
  • Act as a key technical contributor for complex ML infrastructure challenges, mentor junior colleagues, and raise the engineering bar through reviews, pairing, and knowledge sharing.
  • Take ownership of technical design decisions within the team and bring informed input to long-term platform and runtime infrastructure strategy decisions with product and senior engineering leadership.
  • Play an active role in hiring, onboarding, and mentoring engineers, helping to build a strong technical culture around ML infrastructure and platform engineering.

Requirements

  • You have 5+ years of experience building and operating ML Infrastructure or large-scale distributed systems on a cloud platform (AWS/EKS or equivalent), with strong skills in containerization (Docker), Kubernetes, and streaming/batch processing (e.g., Kafka/Kinesis, Spark/Flink).
  • You have hands-on experience with data/feature engineering pipelines, schema evolution, and ensuring online/offline consistency, with familiarity with feature stores (e.g., Feast, SageMaker).
  • You are experienced in designing and operating low-latency, high-scale distributed systems that meet strict throughput targets, including caching, request shaping, and traffic management.
  • You have experience operationalizing Kubernetes-native ML workloads (e.g., model serving, deployment, runtime) using technologies like NVIDIA Triton, MLflow, Kubeflow, or ZenML.
  • You have a background in building or integrating developer tooling, platform automation, or workflow systems, including emerging AI-assisted development or agentic workflows.
  • You have a track record of building reliable systems with SLOs, monitoring, and deployment safeguards, and are comfortable handling incident response and capacity planning.
  • You are proficient in security and governance (e.g., IAM, secrets management, network boundaries) and have experience embedding compliance into engineering workflows.
  • You have strong collaboration and communication skills, enabling you to work effectively with engineers, applied scientists, and product partners to translate requirements into reliable platform capabilities.

Benefits & conditions

Zalando provides a range of benefits, here’s an overview of what you can expect. Ask your Talent Acquisition Partner to learn more about what we offer.

  • 27 days of holiday a year to start for full-time employees (+1 day for every calendar year up to 30 days)
  • 2 paid volunteering days a year
  • Employee shares program
  • 40% off fashion and beauty products sold and shipped by Zalando, 30% off Lounge by Zalando, discounts from external partners
  • Relocation assistance available (subject to prior agreement)
  • Family services, including counseling and support
  • Health and wellbeing options (including Wellhub, formerly Gympass)
  • Mental health support and coaching available
  • Drive your development through our training platform and biannual peer-to-peer review

About the company

At Zalando, our vision is to be the leading pan-European ecosystem for fashion and lifestyle e-commerce - one that thrives on diversity and is truly inclusive by design. We believe that diverse teams fuel innovation and creativity, and we actively seek out talent from all backgrounds.

We actively seek to reduce bias in our hiring and employment processes, focusing on your qualifications, skills, and contributions. To support this, we kindly ask that you refrain from including personal details such as your photo, age, or marital status in your CV, ensuring a fair and equitable evaluation based solely on your abilities and potential.

We are committed to providing an exceptional and accessible candidate experience for everyone. If you require any accommodations to support you throughout the hiring process, please let us know - we are here to assist you.

Discover more about our commitment to creating a diverse and inclusive workplace: https://jobs.zalando.com/en/our-culture/diversity-and-inclusion

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · WWC Europe 2026

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all