> Markdown version of [/jobs/ext/613164-data-engineer](https://www.wearedevelopers.com/jobs/ext/613164-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** RIPJAR LIMITED - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Airflow, Confluence, JIRA, Code Review, Continuous Integration, Information Engineering, Data Structures, Software Debugging, Linux, Distributed Data Store, Distributed Systems, Github, Apache Hadoop, Hadoop Distributed File System, Apache HBase, Integrated Development Environments, Python (Programming Language), MongoDB, Node.Js, Octopus Deploy, Cloud Services, Ansible, Management of Software Versions, Workflow Management Systems, Apache Spark, Indexer, Pyspark, Integration Tests, Kubernetes, Information Technology, Apache Nifi, Data Management, Rundeck, Terraform, Software Version Control, Data Pipelines, Service Stack, Programming Languages - **Published:** June 16, 2026 - **Apply:** https://arc.dev/remote-jobs/j/ripjar-data-engineer-oxa71e3v4w ## About the Role We're looking for someone with 2+ years of industry experience building and operating production software who enjoys working across data pipelines, distributed systems, and operational reliability., Essential: * 2+ years building and operating production software systems * Fluency in at least one programming language (Python/Node.js a plus) * Experience debugging moderately complex systems and improving reliability/performance * Strong fundamentals: data structures, testing, version control, Linux basics Nice to have: * Spark/PySpark experience * Hadoop ecosystem exposure (HDFS/HBase) * Workflow orchestration (Airflow/Dagster/NiFi) * Search/indexing (OpenSearch, MongoDB) * Kubernetes and infrastructure-as-code * Degree in Computer Science or numerical degree ## Description We see a Data Engineer as a software engineer who specialises in distributed data systems. You'll join the Data Engineering team, whose prime responsibility is the development and operation of the Data Collection Hub, a platform that ingests data from many sources, processes/enriches it, and distributes it to multiple downstream systems., * Engineer distributed ingestion services that reliably pull data from diverse sources, handle messy real-world edge cases, and deliver clean, well-structured outputs to multiple downstream products. * Build high-throughput processing components (batch and/or near-real-time) with a focus on performance, scalability, and predictable cost, using strong profiling and measurement practices. * Design and evolve data contracts (schemas, validation rules, versioning, backward compatibility) so downstream teams can build with confidence. * Own production quality: write maintainable code, strong unit/integration tests, and add the observability you need (metrics/logs/tracing) to diagnose issues quickly. * Improve platform reliability by hardening pipelines against partial failures, retries, rate limits, data drift, and infrastructure issues-then codify those learnings into better tooling and guardrails. * Contribute to CI/CD and developer experience: faster builds, better test signal, safer releases, and automated operational checks. * Participate in design reviews, code reviews, incident retrospectives, and iterative delivery-making pragmatic trade-offs and documenting them clearly. Technology Stack * Languages: Predominantly Python and Node.js * Distributed/data platforms: HDFS, HBase, Spark, plus increasing use of Kubernetes and cloud services * Storage/search: MongoDB, OpenSearch * Orchestration: Airflow, Dagster, NiFi * Tooling: GitHub, GitHub Actions, Rundeck, Jira, Confluence * Deployment/config: Ansible (physical), Terraform / Argo CD / Helm (Kubernetes) * Development environment: MacBook (typical) ## Related Videos - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)