Senior Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+23 more
Job description
Experteer Overview In this role you will design and operate large-scale data ingestion and processing pipelines for the DOJ Digital Evidence Review Platform. You will ensure data integrity, provenance, auditability, and secure chain-of-custody across diverse evidence datasets. You will collaborate with architects, engineers, and cybersecurity teams to enable search, analytics, and AI-enabled processing while maintaining strong governance. This is a mission-critical, hands-on position that shapes how digital evidence is ingested, secured, and analyzed at scale. Compensation / Benefits * Design, develop, test, operate, and optimize high-volume data ingestion, transformation, and processing pipelines * Ingest, normalize, enrich, validate, store, search, export, archive, and restore structured and unstructured data * Process sources including forensic packages, documents, multimedia, call records, geolocation, metadata, and parser outputs * Implement data quality controls, error handling, retry, reconciliation, and reprocessing * Preserve source identifiers, lineage, provenance, and chain-of-custody events * Support OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment * Develop schemas, canonical models, and open, portable export formats * Support search, analytics, and AI/ML-enabled processing with auditability * Automate deployment and testing via CI/CD and configuration management * Tune performance and troubleshoot across DEV/TEST/UAT/Prod/DR environments * Maintain technical documentation and collaborate with architects, cloud integrators, cybersecurity, and product specialists Tasks * Bachelor’s degree in engineering, mathematics, or science * 15+ years in data engineering, data integration, data architecture, software engineering, or large-scale data processing * Experience designing and maintaining enterprise data pipelines * Experience with structured and unstructured data from multiple sources * Experience with data modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging * Proficiency in Python, Java, Scala, SQL, or equivalent data-engineering languages * Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments * Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting * U.S. Citizenship and ability to obtain DOJ residency and security clearances * Preferred: experience with digital evidence, forensic data, eDiscovery, and federal cloud environments * Preferred: experience with OCR, NLP, entity extraction, transcription, translation, AI/ML enrichment, or LLM/RAG pipelines * Preferred: experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow Key requirements *
Requirements
retry, reconciliation, and reprocessing * Preserve source identifiers, lineage, provenance, and chain-of-custody events * Support OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment * Develop schemas, canonical models, and open, portable export formats * Support search, analytics, and AI/ML-enabled processing with auditability * Automate deployment and testing via CI/CD and configuration management * Tune performance and troubleshoot across DEV/TEST/UAT/Prod/DR environments * Maintain technical documentation and collaborate with architects, cloud integrators, cybersecurity, and product specialists Tasks * Bachelor’s degree in engineering, mathematics, or science * 15+ years in data engineering, data integration, data architecture, software engineering, or large-scale data processing * Experience designing and maintaining enterprise data pipelines * Experience with structured and unstructured data from multiple sources * Experience with a and modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging * Proficiency in Python, Java, Scala, SQL, or equivalent data-engineering languages * Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments * Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting * U.S. Citizenship and ability to obtain DOJ residency and security clearances * Preferred: experience with digital evidence, forensic data, eDiscovery, and federal cloud environments * Preferred: experience with OCR, NLP, entity extraction, transcription, translation, AI/ML enrichment, or LLM/RAG pipelines * Preferred: experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Data Engineer Salary UK