> Markdown version of [/jobs/ext/3293412-data-engineer](https://www.wearedevelopers.com/jobs/ext/3293412-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Ripjar - **Location:** Cheltenham, UK (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Airflow, Confluence, JIRA, Code Review, Continuous Integration, Information Engineering, Data Structures, Software Debugging, Linux, Distributed Systems, Github, Apache Hadoop, Hadoop Distributed File System, Apache HBase, Integrated Development Environments, Python (Programming Language), MongoDB, Node.Js, Octopus Deploy, Cloud Services, Ansible, Management of Software Versions, Workflow Management Systems, Apache Spark, Indexer, Pyspark, Integration Tests, Kubernetes, Information Technology, Apache Nifi, Data Management, Rundeck, Terraform, Software Version Control, Data Pipelines, Service Stack, Programming Languages - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/data-engineer-ripjar-7762969 ## About the Role We're looking for someone with 2+ years of industry experience building and operating production software who enjoys working across data pipelines, distributed systems, and operational reliability., Essential: * 2+ years building and operating production software systems * Fluency in at least one programming language (Python/Node.js a plus) * Experience debugging moderately complex systems and improving reliability/performance * Strong fundamentals: data structures, testing, version control, Linux basics Nice to have: * Spark/PySpark experience * Hadoop ecosystem exposure (HDFS/HBase) * Workflow orchestration (Airflow/Dagster/NiFi) * Search/indexing (OpenSearch, MongoDB) * Kubernetes and infrastructure-as-code * Degree in Computer Science or numerical degree ## Description We see a Data Engineer as a software engineer who specialises in distributed data systems. You'll join the Data Engineering team, whose prime responsibility is the development and operation of the Data Collection Hub, a platform that ingests data from many sources, processes/enriches it, and distributes it to multiple downstream systems., * Engineer distributed ingestion services that reliably pull data from diverse sources, handle messy real-world edge cases, and deliver clean, well-structured outputs to multiple downstream products. * Build high-throughput processing components (batch and/or near-real-time) with a focus on performance, scalability, and predictable cost, using strong profiling and measurement practices. * Design and evolve data contracts (schemas, validation rules, versioning, backward compatibility) so downstream teams can build with confidence. * Own production quality: write maintainable code, strong unit/integration tests, and add the observability you need (metrics/logs/tracing) to diagnose issues quickly. * Improve platform reliability by hardening pipelines against partial failures, retries, rate limits, data drift, and infrastructure issues-then codify those learnings into better tooling and guardrails. * Contribute to CI/CD and developer experience: faster builds, better test signal, safer releases, and automated operational checks. * Participate in design reviews, code reviews, incident retrospectives, and iterative delivery-making pragmatic trade-offs and documenting them clearly. Technology Stack * Languages: Predominantly Python and Node.js * Distributed/data platforms: HDFS, HBase, Spark, plus increasing use of Kubernetes and cloud services * Storage/search: MongoDB, OpenSearch * Orchestration: Airflow, Dagster, NiFi * Tooling: GitHub, GitHub Actions, Rundeck, Jira, Confluence * Deployment/config: Ansible (physical), Terraform / Argo CD / Helm (Kubernetes) * Development environment: MacBook (typical) ## Related Videos - [GitOps for the people](https://www.wearedevelopers.com/videos/461-gitops-for-the-people) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)