> Markdown version of [/jobs/ext/2091478-data-engineer](https://www.wearedevelopers.com/jobs/ext/2091478-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Thetaray - **Location:** Valladolid, Spain - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, Apache HTTP Server, Big Data, Cloudera Impala, Information Systems, Customer Data Management, Information Engineering, Data Files, Data Transformation, Data Structures, Linux, Elasticsearch, Apache Hadoop, Hadoop Distributed File System, Apache Hive, Machine Learning, Metadata, SQL Databases, Sqoop, Data Streaming, Feature Engineering, Apache Spark, Jupyter, Git, Pandas, Pyspark, Information Technology, Build Process, Software Version Control, Data Pipelines, Docker, Jenkins, Microservices - **Published:** August 16, 2026 - **Apply:** https://www.buscojobs.com.es/data-engineer-en-valladolid-ID-367552822 ## About the Role 2+ years of Hands-on experience working with Apache Spark - must Hands-on experience with SQL Hands-on experience with version-control tools such as GIT Hands-on experience with Apache Hadoop Ecosystem including Hive, Impala, Hue, HDFS, Sqoop etc.. Experience with Python (Pandas) Experience with PySpark/Scala/Java/R Hands-on experience with data transformation, validations, cleansing, and ML feature engineering BSc degree or higher in Computer Science, Statistics, Informatics, Information Systems, Engineering, or another quantitative field Experience working with and optimizing big data pipelines, architectures, and data sets - an advantage Strong analytic skills related to working with structured and semi-structured datasets Build processes supporting data transformation, data structures, metadata, dependency, and workload management Experience performing root cause analysis on internal and external data and processes to answer specific business questions and identify opportunities for improvement Business-oriented and able to work with external customers and cross-functional teams Fluent in English & Spanish both written and spoken Nice to have Experience with Linux Experience in building Machine Learning pipeline Experience with Elasticsearch Experience with Zeppelin/Jupyter Experience with workflow automation platforms such as Jenkins or Apache Airflow Experience with Microservices architecture components, including Docker and Kubernetes. ## Description At ThetaRay, our purpose is to make the world a safer place by protecting the integrity of the global financial system.We do this by putting AI at the core of both our technology and our way of working.Our AI-driven solutions help banks and fintech companies worldwide detect and stop serious financial crime, from human trafficking and terrorist financing to sophisticated money laundering, while advanced technology, automation, and AI-driven tools help our teams collaborate smarter, move faster, and continuously improve how we build, deliver, and innovate.About the role:We are looking for aData Engineerto turn expertise, initiative, and bold thinking into real impact on the next generation of AI-driven financial crime detection.If you combine strong data engineering capabilities with hands-on experience in building and optimizing data pipelines and transformations at scale, and if you are motivated by designing the data flows that power real-world money laundering detection for global financial institutions, ThetaRay could be your next challenge.Responsibilities:Implement and maintain data pipeline flows in production within the ThetaRay system based on the data scientist's designDesign and implement solution-based data flows for specific use cases, enabling the applicability of implementations within the ThetaRay productBuilding a Machine Learning data pipelineCreate data tools for analytics and data scientist team members that assist them in building and optimizing our product into an innovative industry leaderWork with product, R&D, data, and analytics experts to strive for greater functionality in our systemsTrain customer data scientists and engineers to maintain and amend data pipelines within the productTravel to customer locations both domestically and abroadBuild and manage technical relationships with customers and partnersRequirements:2+ years of Hands-on experience working with Apache Spark - mustHands-on experience with SQLHands-on experience with version-control tools such as GITHands-on experience with Apache Hadoop Ecosystem including Hive, Impala, Hue, HDFS, Sqoop etc..Experience with Python (Pandas)Experience with PySpark/Scala/Java/RHands-on experience with data transformation, validations, cleansing, and ML feature engineeringBSc degree or higher in Computer Science, Statistics, Informatics, Information Systems, Engineering, or another quantitative fieldExperience working with and optimizing big data pipelines, architectures, and data sets - an advantageStrong analytic skills related to working with structured and semi-structured datasetsBuild processes supporting data transformation, data structures, metadata, dependency, and workload managementExperience performing root cause analysis on internal and external data and processes to answer specific business questions and identify opportunities for improvementBusiness-oriented and able to work with external customers and cross-functional teamsFluent in English & Spanish both written and spokenNice to haveExperience with LinuxExperience in building Machine Learning pipelineExperience with ElasticsearchExperience with Zeppelin/JupyterExperience with workflow automation platforms such as Jenkins or Apache AirflowExperience with Microservices architecture components, including Docker and Kubernetes. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)