> Markdown version of [/jobs/ext/1144044-data-engineer](https://www.wearedevelopers.com/jobs/ext/1144044-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** BLN24 - **Location:** McLean, VA, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Airflow, Big Data, Data Validation, Information Engineering, Extract Transform Load (ETL), Data Transformation, Data Stores, Distributed Data Store, Document-Oriented Databases, R (Programming Language), Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Operational Data Store, Operational Databases, Standard Sql, SAS (Software), Workflow Management Systems, Enterprise Data Management, Data Logging, Data Ingestion, Apache Spark, Pandas, Data Lakes, Pyspark, Scikit Learn, Information Technology, Statistics Packages, Data Management, Data Pipelines, Databricks - **Published:** July 2, 2026 - **Apply:** http://bln24.applytojob.com/apply/jobs/details/DR8iZx7QeE ## About the Role * Bachelor's degree in Data Science, Statistics, Computer Science, Engineering, or related field (or equivalent experience) * 3-5 years of experience spanning both data engineering and data science/statistical analysis * Strong proficiency in Python, including experience with data engineering libraries (e.g., pandas, PySpark) and statistical/ML libraries (e.g., scikit-learn, statsmodels) * Hands-on experience building and maintaining ETL/ELT pipelines, including ingestion, transformation, and validation logic * Solid grounding in classical statistical methods (hypothesis testing, regression, distributional analysis) and practical machine learning techniques * Experience working with SQL and relational/distributed data systems * Ability to work within a federal data environment, including familiarity with data sensitivity tiers and access/disclosure constraints * Strong communication skills, with the ability to explain technical/statistical concepts to non-technical stakeholders Preferred Qualifications: * Prior experience supporting federal statistical agencies or other federal data programs * Familiarity with Databricks or modern lakehouse architectures (Spark, Delta Lake, etc.) * Experience with workflow orchestration tools (e.g., Airflow, Databricks Workflows) * Experience designing anomaly-detection or outlier-detection approaches beyond standard threshold-based methods * Exposure to disclosure avoidance concepts or working with regulated/protected government data * Experience working across multiple coding environments (Python, R, SAS) within the same analytics platform * Background in requirements gathering or systems design for enterprise data platforms Work Environment: * Contract position supporting a federal agency data modernization engagement * Collaborative, cross-functional environment working alongside data engineers, data scientists, architects, and program SMEs * Requires U.S. citizenship and ability to obtain a public trust or other clearance/suitability determination typical of federal contractor engagements ## Description BLN24 is seeking a mid-level Data Engineer to support a large-scale data and analytics platform modernization effort for a federal statistical agency client. This is a hybrid role: data engineering (building and maintaining the pipelines that bring data into the platform) and applied data science (using classical statistics and machine learning to analyze that data once it's available). The ideal candidate is equally comfortable writing production-grade ingestion and transformation code as they are designing and validating a statistical or ML model. This role works closely with SMEs across multiple program areas to understand source data, build reliable ETL/ingestion pipelines, and apply analytical methods - anomaly detection, statistical modeling, and machine learning - to support operational decision-making., Data Engineering * Design, build, and maintain ETL/ELT pipelines to ingest data from multiple source systems into the platform's central data store * Develop and maintain data ingestion workflows for both batch and near-real-time sources * Implement data validation, cleaning, and transformation logic to ensure data quality and consistency across pipelines * Work within a modern lakehouse/cloud data architecture, optimizing pipeline performance and reliability * Build and maintain data models and schemas that support downstream analytics and reporting needs * Monitor pipeline health, troubleshoot failures, and implement logging/alerting for data quality issues * Document data lineage, transformation logic, and pipeline architecture for governance and reproducibility Data Science / Statistics & ML * Apply classical statistical methods (hypothesis testing, regression, time-series analysis, distributional comparisons) to identify trends, anomalies, and outliers in operational data * Design and implement benchmarking approaches that compare production data against historical, modeled, or external reference values * Develop and evaluate machine learning models where appropriate, balancing predictive performance with interpretability for non-technical stakeholders * Investigate flagged anomalies by digging into underlying data to identify root causes and contributing factors * Work with SMEs to translate operational questions into analytical approaches, and clearly communicate statistical/ML findings and their limitations * Account for data sensitivity classifications and governance requirements when designing analyses and models * Collaborate with visualization-focused team members to ensure outputs of statistical/ML work are presented clearly to stakeholders ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)