Data Engineer
Role details
Job location
Tech stack
Requirements
Location: ThetaRay Madrid, Community of Madrid, Spain About ThetaRay ThetaRay is a trailblazer in AI-powered Anti-Money Laundering (AML) solutions, offering cutting-edge technology to fintechs, banks, and regulatory bodies worldwide. Our mission is to enhance trust in financial transactions, ensuring compliant and innovative business growth. Why Join ThetaRay? At ThetaRay, you'll be part of a dynamic global team committed to redefining the financial services sector through technological innovation. You will contribute to creating safer financial environments and have the opportunity to work with some of the brightest minds in AI, ML, and fintech. We offer a collaborative, inclusive, and forward-thinking work environment where your ideas and contributions are valued and encouraged. Position: Data Engineer As a Data Engineer you will design, implement, and optimize data pipeline flows within the ThetaRay system, supporting data scientists with data flow implementation based on their feature design and constructing complex rules to detect money-laundering activity. You will build pipeline solutions from the ground up, support multiple production implementations and train customer data scientists and engineers. Responsibilities - Implement and maintain production data pipelines in the ThetaRay system based on data-scientist designs. - Design and implement solution-based data flows for specific use cases, enabling product applicability. - Build a machine-learning data pipeline. - Create data tools for analytics and data-science team members. - Collaborate with product, R&D, data, and analytics experts to enhance system functionality. - Train customer data scientists and engineers to maintain and amend pipelines. - Travel to customer locations domestically and abroad. - Build and manage technical relationships with customers and partners. Requirements - 2+ years of hands-on experience with Apache Spark. - Hands-on experience with SQL. - Version control experience (Git). - Experience with the Apache Hadoop ecosystem (Hive, Impala, Hue, HDFS, Sqoop, etc.). - Python (Pandas) experience. - Experience with PySpark/Scala/Java/R. - Hands-on experience with data transformation, validation, cleansing, and ML feature engineering. - BSc or higher in Computer Science, Statistics, Informatics, Information Systems, Engineering, or another quantitative field. - Experience optimizing big-data pipelines, architectures, and data sets (advantage). - Strong analytical skills with structured and semi-structured datasets. - Experience building processes for data transformation, structures, metadata, dependency, and workload management. - Root-cause analysis experience on internal and external data and processes. - Business-oriented, able to work with external customers and cross-functional teams. - Fluent in English and Spanish (written & spoken). Nice to Have - Linux proficiency. - Experience building machine-learning pipelines. - Experience with Elasticsearch. - Experience with Zeppelin or Jupyter. - Experience with workflow automation platforms (Jenkins, Apache Airflow). - Experience with microservices architecture, Docker, and Kubernetes. Employment Details Seniority level: Not applicable Employment type: Full-time Job function: Information Technology Industries: Software Development