Senior Data Engineer

Nira, Inc.
Washington, DC, United States
about 1 month ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Audit Trail Automation of Tests Big Data Configuration Management Databases Continuous Integration Data Architecture Data Deduplication Information Engineering
+23 more
Data Integration Data Integrity Digital Forensics Distributed Computing Environment Python (Programming Language) Metadata Named Entity Recognition Parsing Performance Tuning Scala (Programming Language) Software Engineering SQL Databases Unstructured Data Cloud Platform System Data Ingestion Large Language Models Apache Spark Data Lineage Deployment Automation Integration Frameworks Apache Kafka Data Pipelines Databricks

Job description

Experteer Overview In this role you will design and operate large-scale data ingestion and processing pipelines for the DOJ Digital Evidence Review Platform. You will ensure data integrity, provenance, auditability, and secure chain-of-custody across diverse evidence datasets. You will collaborate with architects, engineers, and cybersecurity teams to enable search, analytics, and AI-enabled processing while maintaining strong governance. This is a mission-critical, hands-on position that shapes how digital evidence is ingested, secured, and analyzed at scale. Compensation / Benefits * Design, develop, test, operate, and optimize high-volume data ingestion, transformation, and processing pipelines * Ingest, normalize, enrich, validate, store, search, export, archive, and restore structured and unstructured data * Process sources including forensic packages, documents, multimedia, call records, geolocation, metadata, and parser outputs * Implement data quality controls, error handling, retry, reconciliation, and reprocessing * Preserve source identifiers, lineage, provenance, and chain-of-custody events * Support OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment * Develop schemas, canonical models, and open, portable export formats * Support search, analytics, and AI/ML-enabled processing with auditability * Automate deployment and testing via CI/CD and configuration management * Tune performance and troubleshoot across DEV/TEST/UAT/Prod/DR environments * Maintain technical documentation and collaborate with architects, cloud integrators, cybersecurity, and product specialists Tasks * Bachelor’s degree in engineering, mathematics, or science * 15+ years in data engineering, data integration, data architecture, software engineering, or large-scale data processing * Experience designing and maintaining enterprise data pipelines * Experience with structured and unstructured data from multiple sources * Experience with data modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging * Proficiency in Python, Java, Scala, SQL, or equivalent data-engineering languages * Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments * Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting * U.S. Citizenship and ability to obtain DOJ residency and security clearances * Preferred: experience with digital evidence, forensic data, eDiscovery, and federal cloud environments * Preferred: experience with OCR, NLP, entity extraction, transcription, translation, AI/ML enrichment, or LLM/RAG pipelines * Preferred: experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow Key requirements *

Requirements

retry, reconciliation, and reprocessing * Preserve source identifiers, lineage, provenance, and chain-of-custody events * Support OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment * Develop schemas, canonical models, and open, portable export formats * Support search, analytics, and AI/ML-enabled processing with auditability * Automate deployment and testing via CI/CD and configuration management * Tune performance and troubleshoot across DEV/TEST/UAT/Prod/DR environments * Maintain technical documentation and collaborate with architects, cloud integrators, cybersecurity, and product specialists Tasks * Bachelor’s degree in engineering, mathematics, or science * 15+ years in data engineering, data integration, data architecture, software engineering, or large-scale data processing * Experience designing and maintaining enterprise data pipelines * Experience with structured and unstructured data from multiple sources * Experience with a and modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging * Proficiency in Python, Java, Scala, SQL, or equivalent data-engineering languages * Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments * Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting * U.S. Citizenship and ability to obtain DOJ residency and security clearances * Preferred: experience with digital evidence, forensic data, eDiscovery, and federal cloud environments * Preferred: experience with OCR, NLP, entity extraction, transcription, translation, AI/ML enrichment, or LLM/RAG pipelines * Preferred: experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:56 min

Open-sourcing a complex parsing library for game data

Johan Hutting Johan Hutting · World Congress 2024

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

1:47 min

Comparing Egeria to alternative open metadata solutions

Ferd Scheepers · World Congress 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:36 min

Managing complex operation sequence weights using recursive parsing

Florian Rappl · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all