Data & Machine Learning Engineer

Medium
Málaga, Spain
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
2 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Microsoft Windows Application Programming Interfaces (APIs) Agile Methodology Artificial Intelligence Amazon Web Services Big Data Cloud Computing Information Engineering Data Stores Data Warehousing Relational Databases Distributed Computing Environment
+26 more
Apache Hadoop JSON Python (Programming Language) Machine Learning Payment Service Provider Azure Machine Learning Software Engineering PL-SQL SQL Databases Systems Integration Unstructured Data Software Repository Data Processing Datastax Data Ingestion Large Language Models Snowflake Prompt Engineering Apache Spark Deployment Automation HuggingFace Apache Kafka Machine Learning Operations GPT Software Version Control Data Pipelines

Job description

This is a full-time opportunity for a data/ML Engineer from LATAM. In?person verification will be conducted.IDT is an American telecommunications company founded in ** and headquartered in New Jersey. It is an industry leader in prepaid communication and payment services and one of the world’s largest international voice carriers. The company is listed on the NYSE, employs over 1,300 people across 20+ countries, and has revenues in excess of $1.5 billion.We are looking for a skilled Data/ML Engineer to join our BI team and take an active role in designing, building, and maintaining the end?to?end data pipeline, architecture, and design that powers our warehouse, LLM?driven applications, and AI?based BI.ResponsibilitiesDesign, develop, and maintain scalable data pipelines to support ingestion, transformation, and delivery into centralized feature stores, model?training workflows, and real?time inference services.Build and optimize workflows for extracting, storing, and retrieving semantic representations of unstructured data to enable advanced search and retrieval patterns.Architect and implement lightweight analytics and dashboarding solutions that deliver natural language query experience and AI?backed insights.Define and execute processes for managing prompt engineering techniques, orchestration flows, and model fine?tuning routines to power conversational interfaces.Oversee vector data stores and develop efficient indexing methodologies to support retrieval?augmented generation (RAG) workflows.Partner with data stakeholders to gather requirements for language?model initiatives and translate them into scalable solutions.Create and maintain comprehensive documentation for all data processes, workflows, and model deployment routines.Stay informed and learn emerging methodologies in data engineering, MLOps, and LLM operations.Requirements8+ years of experience as a Data Engineer with 2+ years focused on MLOps.Excellent English communication skills.Effective oral and written communication skills with the BI team and user community.Demonstrated experience in utilizing Python for data engineering tasks, including transformation, advanced data manipulation, and large?scale data processing.Deep understanding of vector databases and RAG architectures, and how they drive semantic retrieval workflows.Skilled at integrating open?source LLM frameworks into data engineering workflows for end?to?end model training, customization, and scalable inference.Experience with cloud platforms like AWS or Azure Machine Learning for managed LLM deployments.Hands?on experience with big data technologies including Apache Spark, Hadoop, and Kafka for distributed processing and real?time data ingestion.Experience designing complex data pipelines extracting data from RDBMS, JSON, API, and flat?file sources.Demonstrated skills in SQL and PL/SQL programming, with advanced mastery in Business Intelligence and data warehouse methodologies, and hands?on experience in one or more relational database systems and cloud?based database services such as Snowflake or Redshift.Understanding of software engineering principles and experience working on Unix/Linux/Windows operating systems, and experience with Agile methodologies.Proficiency in version control systems, with experience in managing code repositories, branching, merging, and collaborating within a distributed development environment.Interest in business operations and comprehensive understanding of how robust BI systems drive corporate profitability by enabling data?driven decision?making and strategic insights.PlusesExperience with vector databases such as DataStax AstraDB, and developing LLM?powered applications using popular open?source frameworks like LangChain and LlamaIndex - including prompt engineering, retrieval?augmented generation (RAG), and orchestration of intelligent workflows.Familiarity with evaluating and integrating open?source LLM frameworks - such as Hugging Face Transformers or LLaMA?4 - across end?to?end workflows, including fine?tuning and inference optimization.Knowledge of MLOps tooling and CI/CD pipelines to manage model versioning and automated deployments.Only accepting applicants from LATAM.#J-***-Ljbffr

Requirements

8+ years of experience as a Data Engineer with 2+ years focused on MLOps. Excellent English communication skills. Effective oral and written communication skills with the BI team and user community. Demonstrated experience in utilizing Python for data engineering tasks, including transformation, advanced data manipulation, and large?scale data processing. Deep understanding of vector databases and RAG architectures, and how they drive semantic retrieval workflows. Skilled at integrating open?source LLM frameworks into data engineering workflows for end?to?end model training, customization, and scalable inference. Experience with cloud platforms like AWS or Azure Machine Learning for managed LLM deployments. Hands?on experience with big data technologies including Apache Spark, Hadoop, and Kafka for distributed processing and real?time data ingestion. Experience designing complex data pipelines extracting data from RDBMS, JSON, API, and flat?file sources. Demonstrated skills in SQL and PL/SQL programming, with advanced mastery in Business Intelligence and data warehouse methodologies, and hands?on experience in one or more relational database systems and cloud?based database services such as Snowflake or Redshift. Understanding of software engineering principles and experience working on Unix/Linux/Windows operating systems, and experience with Agile methodologies. Proficiency in version control systems, with experience in managing code repositories, branching, merging, and collaborating within a distributed development environment. Interest in business operations and comprehensive understanding of how robust BI systems drive corporate profitability by enabling data?driven decision?making and strategic insights. Pluses Experience with vector databases such as DataStax AstraDB, and developing LLM?powered applications using popular open?source frameworks like LangChain and LlamaIndex - including prompt engineering, retrieval?augmented generation (RAG), and orchestration of intelligent workflows. Familiarity with evaluating and integrating open?source LLM frameworks - such as Hugging Face Transformers or LLaMA?4 - across end?to?end workflows, including fine?tuning and inference optimization. Knowledge of MLOps tooling and CI/CD pipelines to manage model versioning and automated deployments. Only accepting applicants from LATAM. #J-*****-Ljbffr

About the company

IDT is an American telecommunications company founded in ** and headquartered in New Jersey. It is an industry leader in prepaid communication and payment services and one of the world’s largest international voice carriers. The company is listed on the NYSE, employs over 1,300 people across 20+ countries, and has revenues in excess of $1.5 billion.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · WWC 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all