> Markdown version of [/jobs/ext/2565613-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/2565613-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Nira, Inc. - **Location:** Washington, DC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, Audit Trail, Automation of Tests, Big Data, Configuration Management, Databases, Continuous Integration, Data Architecture, Data Deduplication, Information Engineering, Data Integration, Data Integrity, Digital Forensics, Distributed Computing Environment, Python (Programming Language), Metadata, Named Entity Recognition, Parsing, Performance Tuning, Scala (Programming Language), Software Engineering, SQL Databases, Unstructured Data, Cloud Platform System, Data Ingestion, Large Language Models, Apache Spark, Data Lineage, Deployment Automation, Integration Frameworks, Apache Kafka, Data Pipelines, Databricks - **Published:** August 5, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-data-engineer-washington-dc-usa-58838525 ## About the Role retry, reconciliation, and reprocessing * Preserve source identifiers, lineage, provenance, and chain-of-custody events * Support OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment * Develop schemas, canonical models, and open, portable export formats * Support search, analytics, and AI/ML-enabled processing with auditability * Automate deployment and testing via CI/CD and configuration management * Tune performance and troubleshoot across DEV/TEST/UAT/Prod/DR environments * Maintain technical documentation and collaborate with architects, cloud integrators, cybersecurity, and product specialists Tasks * Bachelor's degree in engineering, mathematics, or science * 15+ years in data engineering, data integration, data architecture, software engineering, or large-scale data processing * Experience designing and maintaining enterprise data pipelines * Experience with structured and unstructured data from multiple sources * Experience with a and modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging * Proficiency in Python, Java, Scala, SQL, or equivalent data-engineering languages * Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments * Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting * U.S. Citizenship and ability to obtain DOJ residency and security clearances * Preferred: experience with digital evidence, forensic data, eDiscovery, and federal cloud environments * Preferred: experience with OCR, NLP, entity extraction, transcription, translation, AI/ML enrichment, or LLM/RAG pipelines * Preferred: experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow Key requirements * ## Description Experteer Overview In this role you will design and operate large-scale data ingestion and processing pipelines for the DOJ Digital Evidence Review Platform. You will ensure data integrity, provenance, auditability, and secure chain-of-custody across diverse evidence datasets. You will collaborate with architects, engineers, and cybersecurity teams to enable search, analytics, and AI-enabled processing while maintaining strong governance. This is a mission-critical, hands-on position that shapes how digital evidence is ingested, secured, and analyzed at scale. Compensation / Benefits * Design, develop, test, operate, and optimize high-volume data ingestion, transformation, and processing pipelines * Ingest, normalize, enrich, validate, store, search, export, archive, and restore structured and unstructured data * Process sources including forensic packages, documents, multimedia, call records, geolocation, metadata, and parser outputs * Implement data quality controls, error handling, retry, reconciliation, and reprocessing * Preserve source identifiers, lineage, provenance, and chain-of-custody events * Support OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment * Develop schemas, canonical models, and open, portable export formats * Support search, analytics, and AI/ML-enabled processing with auditability * Automate deployment and testing via CI/CD and configuration management * Tune performance and troubleshoot across DEV/TEST/UAT/Prod/DR environments * Maintain technical documentation and collaborate with architects, cloud integrators, cybersecurity, and product specialists Tasks * Bachelor's degree in engineering, mathematics, or science * 15+ years in data engineering, data integration, data architecture, software engineering, or large-scale data processing * Experience designing and maintaining enterprise data pipelines * Experience with structured and unstructured data from multiple sources * Experience with data modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging * Proficiency in Python, Java, Scala, SQL, or equivalent data-engineering languages * Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments * Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting * U.S. Citizenship and ability to obtain DOJ residency and security clearances * Preferred: experience with digital evidence, forensic data, eDiscovery, and federal cloud environments * Preferred: experience with OCR, NLP, entity extraction, transcription, translation, AI/ML enrichment, or LLM/RAG pipelines * Preferred: experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow Key requirements * ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [A Data Mesh needs Open Metadata](https://www.wearedevelopers.com/videos/505-a-data-mesh-needs-open-metadata) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Crafting Custom Frameworks with Rust: A Deep Dive into Procedural Macros](https://www.wearedevelopers.com/videos/849-crafting-custom-frameworks-with-rust-a-deep-dive-into-procedural-macros) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)