> Markdown version of [/jobs/ext/1283155-data-engineer](https://www.wearedevelopers.com/jobs/ext/1283155-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Edgesource Corporation - **Location:** McLean, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Geographic Information Systems, Airflow, Amazon Web Services, Amazon S3, Apache HTTP Server, Bash Shell, Big Data, Computer Programming, System Configuration, Data Centers, Information Engineering, Data Governance, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Data Security, Data Visualization, Software Debugging, Software Design Patterns, Distributed Computing Environment, Amazon DynamoDB, Python (Programming Language), PostgreSQL, Metadata Repositories, MySQL, NoSQL, NumPy, Operational Databases, Performance Tuning, PostGIS, Query Optimization, DataOps, Azure Machine Learning, Software Deployment, Software Engineering, SQL Databases, Data Streaming, Systems Integration, Strategies of Testing, Speech Recognition, Data Processing, Cloud Platform System, Large Language Models, Apache Spark, Topic Modeling, Git, Cloudformation, Pandas, Containerization, Pyspark, Data Lineage, Data Lakehouse, Terraform, Software Version Control, Data Pipelines, Docker - **Published:** July 15, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9030384/data-engineer ## About the Role * Minimum of 5 years' experience * Demonstrated experience building production data pipelines and ETL/EL workflows at scale * Proficiency with Apache Spark and PySpark for distributed data processing * Advanced Python programming skills including data manipulation libraries (Pandas, NumPy) and data engineering best practices * Understanding of data security, privacy, governance, and compliance principles * Experience with workflow orchestration tools (such as Step Functions, Airflow) * Familiarity with containerization (such as Docker or Podman) and deploying data applications in cloud environments * Experience with AWS services (S3, Lambda, Step Functions) * Experience with PostgreSQL and MySQL in production environments, including performance tuning and schema design * Demonstrated experience with SQL and query optimization for complex analytical workloads * Experience with version control (Git) and Cl/CD practices for data pipelines * Demonstrated ability to work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight * Strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks * Experience with data Lakehouse architecture using Apache Iceberg * Hands-on experience configuring, deploying, and integrating data platform components: * Apache Ranger (access control and data governance) * Trino (distributed SQL query engine) * Data catalogs (Unity Catalog OSS, Apache Polaris, etc.) * Apache Superset {data visualization and dashboarding) * Proficiency with Bash scripting for automation and data processing tasks * Experience with Infrastructure as Code (Terraform or CloudFormation) for data infrastructure * Familiarity or experience with tracking data lineage and associated tooling such as Open lineage * Familiarity or experience with Java * Familiarity with data quality frameworks, testing methodologies, and validation strategies * Background with large-scale data migrations or platform modernization efforts * Experience integrating Al/ML services and models (translation, OCR, speech-to-text, NLP, language detection, topic modeling), LLMs, and RAG {retrieval-augmented generation) pipelines * Familiarity with geospatial data processing {H3, PostGIS, or similar) * Contributions to data engineering documentation, best practices, and design patterns * Experience with NoSQL databases (DynamoDB, etc.) As an ISO 9001:2015 certified and CMMI Level 3 appraised small business, Edgesource specializes in providing a variety of technical solutions to include software development, database services, enterprise networking, data center virtualization, and management support. We are always seeking top-talent to join our team in helping to address the most critical technical challenges facing our nation. ## Description We are seeking a Data Engineer to work with a small team to build complex data flows for a custom application. Successful candidate will have advanced Python programming skills, familiarity with Java, an understanding of data security, privacy, governance and compliance principles and a demonstrated history of building production data pipelines and ETL workflows at scale. * Building end-to-end data pipelines leveraging Python * Using orchestration tools to deploy data pipelines, including configuring and updating Spark Jobs * Containerizing and deploying applications in cloud environments like AWS. * Working with MySQL and PostgreSQL including performance tuning, schema design, and query optimization for complex, analytical workloads. * Leveraging industry standard tools for code control (Git, IaaC control, etc.) * Working with data catalogs, tracking data lineage and handling a variety of data formats, including Geospatial. * Using Bash scripting for automation and data processing tasks * Integrating Al/ML services and models * Work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight * Leverage strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks * Leverage a background in large-scale data migration or platform modernization efforts * Contribute to data engineering documentation, best practices, and design patterns. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)