> Markdown version of [/jobs/ext/1878959-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1878959-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Bull - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Confluence, JIRA, Distributed Systems, Github, Monitoring of Systems, InfiniBand, Python (Programming Language), Linux System Administration, Machine Learning, Scrum Methodology, Tensorflow, Prometheus, Graphics Processing Unit (GPU), High Performance Computing, Pytorch, Model Validation, Reliability of Systems, Gitlab, Scikit Learn, Kubernetes, Information Technology, Data Analytics, Machine Learning Operations, Software Version Control - **Published:** July 31, 2026 - **Apply:** https://www.jobleads.com/es/job/ee8317a3b1bcd9af9147da8f672398824 ## About the Role * Master's or PhD in Computer Science, Artificial Intelligence, Data Science, Telecommunications or a related field. * Strong experience with Machine Learning and Deep Learning frameworks (e.g., TensorFlow, PyTorch, Scikit-learn). * Experience working with time-series data and anomaly detection models. * Proficiency in Python and data science ecosystems. * Experience with Prometheus or similar monitoring/telemetry systems. * Familiarity with containerization and orchestration technologies, especially Kubernetes. * Experience building production-grade ML pipelines. * Experience handling large-scale monitoring and operational datasets. * Understanding of distributed systems and infrastructure monitoring. * Knowledge of HPC environments, GPUs, and high-speed interconnects (e.g., Infiniband) is highly desirable. * Proficiency with Git-based version control systems (GitHub, GitLab). * Solid experience working in Linux environments. * Good understanding of Scrum methodology and experience with Jira and Confluence. ## Description * Design and develop ML/DL models for predicting hardware failures and detecting software or behavioral anomalies in HPC systems. * Apply advanced analytics techniques such as time-series forecasting, anomaly detection, classification, and predictive maintenance using large-scale monitoring data. * Build and maintain data pipelines and features from infrastructure telemetry and logs. * Perform rigorous model validation to ensure robustness, reliability, and production readiness. * Deploy and operationalize models within a Kubernetes-based environment, including scalable inference services and lifecycle management. * Contribute to AI-driven cybersecurity use cases, such as detecting abnormal behaviors, potential intrusions, or security-related anomalies in infrastructure and system activity. * Work within an Agile/Scrum environment, participating in sprint planning, stand-ups, and retrospectives. * Collaborate with system administrators, support teams, and data engineers to translate operational challenges into data-driven solutions that enhance system reliability and automation. ## Related Videos - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)