> Markdown version of [/jobs/ext/3612127-data-engineer](https://www.wearedevelopers.com/jobs/ext/3612127-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Ocean Infinity - **Location:** London, UK - **Experience:** Experienced - **Salary:** £65,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Business Logic, Automation of Tests, Microsoft Azure, Big Data, Cloud Computing, Software Quality, Code Review, Continuous Integration, Data as a Services, Data Validation, Data Cleansing, Data Integration, Data Transformation, Data Systems, Software Debugging, DevOps, Distributed Systems, JSON, Python (Programming Language), PostgreSQL, MongoDB, NoSQL, Operational Databases, Query Optimization, Redis, Standard Sql, Software Engineering, Data Streaming, Data Processing, Cloud Platform System, Feature Engineering, Indexer, Backend, Containerization, Data Lakes, Semi-structured Data, Git Flow, Information Technology, Machine Learning Operations, Data Delivery, Software Version Control, Data Pipelines, Docker - **Published:** October 8, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5917309955 ## About the Role * A degree in Computer Science, Mathematics, Engineering, or a related field, or equivalent practical experience. * Minimum 2 years of experience as a Data Engineer, Backend Engineer, or Software Engineer working with data-intensive systems and data pipelines. * Solid software engineering skills in Python, with experience writing reliable, maintainable data processing code and automation workflows. * Good SQL skills, with an understanding of query optimization and working with analytical datasets. * Experience building and maintaining data pipelines and transformation workflows, ideally in cloud-based environments. * Hands-on experience with at least one cloud platform (e.g., AWS, GCP, or Azure), particularly for building and running data pipelines, storage systems, and compute workloads. * Familiarity with data lake and lakehouse environments, including object storage systems (e.g., S3 or equivalent) and structured transformation layers (e.g., Bronze/Silver/Gold patterns). * Understanding of unstructured and semi-structured data storage patterns (e.g., JSON, logs, event data, files in object storage) and how to process them efficiently. * Experience with relational and/or NoSQL databases (e.g., Postgres, MongoDB, Redis), including schema design and indexing. * Working knowledge of containerization technologies such as Docker. * Good understanding of data modelling principles and how to structure datasets for analytics and downstream consumption. * Exposure to distributed systems concepts and large-scale data processing patterns, with emphasis on practical implementation and debugging. * Strong problem-solving skills with a focus on reliability, performance, and automation in production systems. * Ability to collaborate effectively with analysts, product teams, and engineers, translating requirements into robust, testable implementations. * Interest in improving and refactoring existing pipelines and workflows to increase reliability, maintainability, and automation. * Comfortable working in modern engineering practices including Git-based workflows, code reviews, CI/CD pipelines, and automated testing for data systems. * Exposure to workflow orchestration tools such as Airflow, Prefect, or similar, including scheduling, dependency management, retries, and observability. Nice to have! * Experience working with telemetry, sensor, location, and operational technology (OT) data workloads. * Experience processing and managing large-scale video, image, and unstructured media datasets, including streaming ingestion and analytics workflows. * Understanding of feature engineering patterns and data preparation workflows supporting machine learning systems. * Experience working with maritime, fleet, vessel operations, logistics, or industrial operational domains, including telematics, tracking, asset monitoring, or operational analytics use cases. ## Description * Develop and implement scalable data pipelines that automate manual data processes and improve reliability and efficiency across the organization. * Build and maintain end-to-end data transformation workflows (Bronze * Silver * Gold), focusing on implementing business logic, data validation, and reusable transformation components. * Develop robust Python-based data processing solutions, including scripts, services, and automation tools that support ingestion, transformation, and data delivery. * Implement data integration pipelines across a variety of sources, including operational databases, external APIs, and event/streaming data, using both batch and near-real-time approaches. * Develop and maintain data quality checks, validation rules, and testing frameworks to ensure accuracy, consistency, and reliability of data products. * Build backend data services and lightweight APIs that expose curated datasets to analytics tools, reporting systems, and downstream applications. * Collaborate closely with data analysts, product teams, and AI engineers to translate data requirements into efficient, production-ready implementations. * Optimize and refactor existing pipelines and processes to improve performance, reliability, and maintainability, with a strong focus on code quality and automation. * Follow and contribute to engineering best practices, including version control standards, modular pipeline design, CI/CD for data workflows, and documentation of implemented solutions. * Participate in code reviews and pairing, learning from more experienced engineers and sharing implementation knowledge and reusable patterns with the team.