Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+21 more
Job description
We are looking for a MLOps / AIOps / LLMOps / AgentOps Engineer to join a multidisciplinary Data & AI team.The main mission of this role is to design, operate, and continuously evolve our AIOps platform, ensuring that our AI products run in a reliable, scalable, and cost?efficient way.This position is strongly focused on platform, infrastructure, automation, observability, and operations rather than on building ML models or AI products themselves.You will work with modern cloud technologies (mainly AWS, with some Azure exposure) and collaborate closely with Data Scientists, Data Engineers, and Product teams to bring AI solutions into production and keep them running smoothly.We are open to candidates with strong expertise in at least one core area (e.g. cloud, DevOps, platform engineering, or ML operations) and solid foundational knowledge in the others, with motivation to grow across the full AI operations stack.Key ResponsibilitiesDesign, maintain, and evolve the AIOps platform supporting:Traditional machine learning models in productionLLM?based solutions such as RAG pipelines and AI AgentsSpeech Analytics use cases (ASR, conversation analysis, NLP)Build and operate ML and LLM pipelines with a strong focus on:Reliability, automation, and observabilityModel and LLM quality, performance, and drift monitoringCloud cost control and optimizationImplement LLMOps / AgentOps practices, including:LLM evaluation and observabilityPrompt management, traceability, and specialized loggingAgent integration, orchestration, and lifecycle managementEnsure continuous operation of AI products, including:Alerts, dashboards, SLOs / SLIsScalability strategies and basic auto?remediation mechanismsManage deployments in cloud environments (AWS / Azure) and container platforms (Docker / Kubernetes)Collaborate closely with Data Scientists and Data Engineers to productionize robust, scalable AI solutionsContribute to internal standards, automation, and best practices across the AI and data ecosystemRequired Skills (Must Have)Hands?on experience in MLOps, AIOps, or operating ML systems in productionSolid understanding of LLMOps and AgentOps concepts (RAGs, agents, evaluation, monitoring)Experience working with AWS and/or Azure in production environmentsPractical knowledge of containers and Kubernetes (Docker, basic Helm usage, etc.)Experience with CI/CD pipelines (GitHub Actions, GitLab CI, Azure DevOps, Jenkins, or similar)Familiarity with observability and monitoring concepts (CloudWatch, OpenTelemetry, Prometheus, etc.)Experience managing infrastructure as code (Terraform, Bicep, CDK, or similar)Python experience and familiarity with the ML ecosystem (e.g. scikit?learn, PyTorch), even if not a Data ScientistGood understanding of the ML / LLM lifecycle, from development to production and monitoringFluent English to work in an international environmentNice To Have (Not Required, But Valuable)Experience with ML/AI platforms such as SageMaker, Azure ML, MLflow, KubeflowExposure to Speech Analytics technologies (ASR, diarization, conversational NLP)Experience with cloud cost optimization / FinOps, especially for AI workloadsExperience building or operating AI agents, copilots, or conversational systemsFamiliarity with LLM frameworks (LangChain, LlamaIndex, Semantic Kernel, etc.)Experience with workflow and orchestration tools (Airflow, Argo, Step Functions, Durable Functions)Professional Skills & MindsetStrong focus on reliability, automation, and scalabilityAbility to collaborate effectively in multidisciplinary teamsClear communication and documentation?oriented mindsetPlatform mindset: building reusable, maintainable, and robust solutionsProactive, analytical, and continuous?improvement drivenStrong sense of ownership and end?to?end responsibilityMotivation to learn and grow across the AI operations stackTechnology EnvironmentCloud: AWS, AzureOrchestration & Containers: Kubernetes, DockerCI/CD: GitHub Actions, GitLab CI, Azure DevOpsObservability: Prometheus, Grafana, ELK/EFK, OpenTelemetryInfrastructure as Code: Terraform, Bicep, CloudFormationAI / ML Tools: MLflow, Azure ML, SageMaker, LangChain, LlamaIndex, Semantic KernelPrimary Language: Python#J-*****-Ljbffr
Requirements
Required Skills (Must Have) Hands?on experience in MLOps, AIOps, or operating ML systems in production Solid understanding of LLMOps and AgentOps concepts (RAGs, agents, evaluation, monitoring) Experience working with AWS and/or Azure in production environments Practical knowledge of containers and Kubernetes (Docker, basic Helm usage, etc.) Experience with CI/CD pipelines (GitHub Actions, GitLab CI, Azure DevOps, Jenkins, or similar) Familiarity with observability and monitoring concepts (CloudWatch, OpenTelemetry, Prometheus, etc.) Experience managing infrastructure as code (Terraform, Bicep, CDK, or similar) Python experience and familiarity with the ML ecosystem (e.g. scikit?learn, PyTorch), even if not a Data Scientist Good understanding of the ML / LLM lifecycle, from development to production and monitoring Fluent English to work in an international environment Nice To Have (Not Required, But Valuable) Experience with ML/AI platforms such as SageMaker, Azure ML, MLflow, Kubeflow Exposure to Speech Analytics technologies (ASR, diarization, conversational NLP) Experience with cloud cost optimization / FinOps, especially for AI workloads Experience building or operating AI agents, copilots, or conversational systems Familiarity with LLM frameworks (LangChain, LlamaIndex, Semantic Kernel, etc.) Experience with workflow and orchestration tools (Airflow, Argo, Step Functions, Durable Functions) Professional Skills & Mindset Strong focus on reliability, automation, and scalability Ability to collaborate effectively in multidisciplinary teams Clear communication and documentation?oriented mindset Platform mindset: building reusable, maintainable, and robust solutions Proactive, analytical, and continuous?improvement driven Strong sense of ownership and end?to?end responsibility Motivation to learn and grow across the AI operations stack Technology Environment Cloud: AWS, Azure Orchestration & Containers: Kubernetes, Docker CI/CD: GitHub Actions, GitLab CI, Azure DevOps Observability: Prometheus, Grafana, ELK/EFK, OpenTelemetry Infrastructure as Code: Terraform, Bicep, CloudFormation AI / ML Tools: MLflow, Azure ML, SageMaker, LangChain, LlamaIndex, Semantic Kernel Primary Language: Python #J-*****-Ljbffr
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.buscojobs.com.esGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps – What’s the deal behind it?
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
What Are Large Language Models?
MLOps And AI Driven Development