Staff Machine Learning Ops Engineer

Preply Inc.
Barcelona, Spain
4 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Big Data Cloud Computing Continuous Integration Distributed Computing Environment Machine Learning Cloud Services Azure Machine Learning Management of Software Versions Autoscaling Large Language Models Machine Learning Operations GPT

Job description

OverviewIn this role you will architect and evolve Preply’s ML platform to move research into production at scale.You’ll design cloud-native systems for distributed training and inference, with strong emphasis on observability, testing, and cost-efficiency.You will partner with ML leads and engineers to set standards, enable experimentation, and deliver modular, measurable ML services.This position offers impact across product teams and a chance to shape the future of a global learning platform.Compensaciones / Beneficioslesson allowance (monthly)Learning & Development budgettime off for self-developmentequityhealth insurancemental health support platformsResponsabilidadesDesign and implement the ML platform architecture (experiment tracking, artifact management, scalable deployment)Build cloud-native distributed training and inference solutions with GPU support and autoscalingOwn CI/CD for ML, including testing, validation, and performance checksEmbed observability across ML workflows (metrics, alerts, drift detection, lineage)Mentor engineers, influence standards, and reduce risk in complex decisionsAlign platform direction with experimentation velocity, cost-efficiency, and user impactDevelop modular, testable ML services with monitoring from day oneContribute to LLM platform capabilities (RAG pipelines, latency-optimized inference, prompt experimentation frameworks)Requisitos principales8+ years of engineering experience in large-scale Data/ML platformsDeep knowledge of cloud services and end-to-end ML workflows (versioning, monitoring, performance benchmarking)Experience collaborating with scientists and building enabling toolsExcellent communication and cross-functional influence; mentoring ML engineersFamiliarity with LLM frameworks (LangChain, LlamaIndex), vector stores, and retrieval infrastructurecommunication and influencementoring and coachingcross-functional collaborationcloud-native ML platform designdistributed training and inference on cloudCI/CD for ML

Requirements

Requisitos principales8+ years of engineering experience in large-scale Data/ML platforms Deep knowledge of cloud services and end-to-end ML workflows (versioning, monitoring, performance benchmarking) Experience collaborating with scientists and building enabling tools Excellent communication and cross-functional influence; mentoring ML engineers Familiarity with LLM frameworks (LangChain, LlamaIndex), vector stores, and retrieval infrastructure communication and influence mentoring and coaching cross-functional collaboration cloud-native ML platform design distributed training and inference on cloud CI/CD for ML

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Bridging the gap between model management and devops

Joy Joy · World Congress 2024

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:33 min

Managing node capacity with default cluster autoscaling

Mario-Leander Reimer · World Congress 2023

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all