> Markdown version of [/jobs/ext/613496-data-engineer](https://www.wearedevelopers.com/jobs/ext/613496-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Staffed4U LLC - **Location:** Chantilly, VA, United States - **Experience:** Experienced - **Contract:** Temporary contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Amazon S3, Big Data, Cloud Database, Code Review, Databases, Continuous Integration, Data Architecture, Information Engineering, Web Scraping, Extract Transform Load (ETL), Distributed Computing Environment, Document-Oriented Databases, Elasticsearch, Graph Database, Python (Programming Language), Linux System Administration, Machine Learning, Microsoft Office, Neo4j, Performance Tuning, Systems Development Life Cycle, SQL Databases, Data Streaming, Systems Integration, Unstructured Data, Management of Software Versions, Data Storage Technologies, Feature Engineering, Data Ingestion, Git, Pyspark, Kubernetes, Data Lineage, Data Management, Machine Learning Operations, Restful APIs, Stream Processing, Software Version Control, Data Pipelines, Docker - **Published:** June 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=eafd6da3be73f364 ## About the Role Do you have experience in Version control systems?, * 3-5+ years of professional experience in Data Engineering or a related technical field. * Experience designing and implementing ETL/ELT pipelines. * Experience processing and managing large-scale structured and unstructured datasets. * Experience working in cloud-based data environments. Technical Skills * Strong proficiency with Python and SQL. * Experience with PySpark or other distributed processing frameworks (highly desired). * Experience with ElasticSearch/OpenSearch technologies. * Experience working within AWS cloud environments. * Experience supporting Linux-based systems. * Proficiency with Git for version control and collaborative development. * Understanding of machine learning workflows and MLOps concepts. * Experience integrating and consuming REST APIs. * Familiarity with Docker, Kubernetes, and CI/CD pipelines. Clearance Requirements * Active TS/SCI with Full Scope Polygraph (FSP) is required. * U.S. Citizenship required. Professional Skills * Strong collaboration and communication skills. * Ability to communicate complex technical concepts to non-technical audiences. * Detail-oriented with a strong commitment to data quality and integrity. * Ability to manage multiple priorities in a fast-paced mission environment. * Strong analytical and problem-solving capabilities. Desired Qualifications * Hands-on experience with graph databases. * Experience modeling, querying, and optimizing Neo4j databases. * Experience supporting advanced analytics, knowledge graphs, or entity resolution systems. * Experience working within Intelligence Community environments. ## Description We are seeking a talented and mission-focused Data Engineer to join our growing team supporting cutting-edge intelligence community initiatives in Chantilly, VA. This role offers the opportunity to work with large-scale datasets and contribute to the development of a custom enterprise platform supporting critical mission objectives., The selected candidate will play a key role in designing, building, and optimizing scalable data pipelines and architectures that support analytics, machine learning, and enterprise data integration efforts. This position is funded for an initial 9-12 month period aligned with defined mission deliverables and system development timelines, with all development performed onsite at the customer location., * Design, develop, and maintain ETL/ELT pipelines for both batch and real-time data processing using Python and SQL. * Integrate data from a variety of structured and unstructured sources, including databases, APIs, streaming platforms, PDFs, and Microsoft Office files. * Build scalable and maintainable data architectures to support analytics and machine learning workloads. * Optimize data processing workflows and queries for performance, scalability, and cost efficiency within AWS environments. * Support future pipeline scalability through exposure to PySpark and other distributed data processing frameworks. * Develop and maintain web scraping and data ingestion workflows to collect and process open-source data. * Transform collected information into structured datasets and visualizations for stakeholder analysis and decision-making., * Collect, clean, validate, and manage large volumes of structured and unstructured data. * Implement data quality controls, validation procedures, and version management practices. * Design and optimize data storage solutions utilizing AWS S3 for raw, intermediate, and production datasets. * Implement data governance best practices including documentation, cataloging, lineage tracking, and security controls. * Ensure compliance with customer and security requirements for data management and handling., * Partner closely with Data Scientists, Analysts, and Engineering teams to understand business and mission requirements. * Prepare clean, structured, and feature-ready datasets for analytics and machine learning applications. * Support feature engineering, aggregation, and large-scale data transformations. * Assist with deploying machine learning models into production environments while supporting monitoring, versioning, and performance optimization. * Integrate and consume REST APIs to support data acquisition and application workflows. * Utilize Docker, Kubernetes, Git, and CI/CD pipelines to support deployment and operational workflows. Documentation & Communication * Document data pipelines, architectures, schemas, and transformation processes. * Communicate technical concepts effectively to both technical and non-technical stakeholders. * Participate in code reviews and promote engineering best practices across the team. * Contribute to continuous improvement efforts related to data engineering, automation, and platform development. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)