> Markdown version of [/jobs/ext/1447893-data-engineer](https://www.wearedevelopers.com/jobs/ext/1447893-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Reality Defender - **Location:** New York, NY, United States (Remote available) - **Salary:** $140,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Big Data, Computer Programming, Databases, Data Validation, Data Infrastructure, Extract Transform Load (ETL), Data Transformation, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Performance Tuning, SQL Databases, Transcoding, Apache Spark, Containerization, Kubernetes, Luigi, Machine Learning Operations, Feature Extraction, Stream Processing, Data Pipelines, Golang - **Published:** July 26, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5ff1fc115883b1b6 ## About the Role * Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments. * Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets. * Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines. * Strong programming skills in Python and SQL; experience with Golang is a plus. * Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets. * Experience working with audio or video data at scale is a strong plus (e.g., transcoding, feature extraction, or preprocessing pipelines). * Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks). * Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines. * Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling). ## Description We're looking for a Data Engineer who can build and scale the infrastructure powering our data platform, with a strong foundation in distributed systems and cloud-native tooling. You'll design and operate the pipelines that move, process, and prepare multi-terabyte and streaming datasets - including audio and video - for Reality Defender's detection models, and you'll work closely with ML engineers and researchers to keep that data flowing reliably at scale. Responsibilities include: * Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows. * Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments. * Build and maintain distributed data processing jobs using frameworks such as Spark and Ray. * Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines. * Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies. * Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)