Data & Software Engineer

NewGen Technologies
Chantilly, VA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Geographic Information Systems Airflow Amazon Web Services Amazon S3 Apache HTTP Server Bash Shell Big Data Computer Programming System Configuration Information Engineering Data Governance
+42 more
Data Infrastructure Data Integration Extract Transform Load (ETL) Data Security Data Visualization Software Debugging Software Design Patterns Distributed Computing Environment Amazon DynamoDB Python (Programming Language) PostgreSQL Metadata Repositories MySQL NoSQL NumPy Operational Databases Performance Tuning PostGIS Query Optimization Azure Machine Learning Software Deployment SQL Databases Data Streaming Systems Integration Strategies of Testing Speech Recognition Data Processing Cloud Platform System Large Language Models Apache Spark Topic Modeling Git Cloudformation Pandas Containerization Pyspark Data Lineage Data Lakehouse Terraform Software Version Control Data Pipelines Docker

Job description

  • Work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight
  • Leverage strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks
  • Leverage a background in large-scale data migration or platform modernization efforts
  • Contribute to data engineering documentation, best practices, and design patterns

Requirements

The Data & Software Engineer works with a small team to build complex data flows for a custom application. Successful candidate will have advanced Python programming skills, familiarity with Java, an understanding of data security, privacy, governance and compliance principles and a demonstrated history of building production data pipelines and ETL workflows at scale.

The candidate must have experience: Building end-to-end data pipelines leveraging Python. Using orchestration tools to deploy data pipelines, including configuring and updating Spark Jobs. Containerizing and deploying applications in cloud environments like AWS. Working with MySQL and PostgreSQL including performance tuning, schema design, and query optimization for complex, analytical workloads. Leveraging industry standard tools for code control (Git, IaaC control, etc.). Working with data catalogs, tracking data lineage and handling a variety of data formats, including Geospatial. Using Bash scripting for automation and data processing tasks. Integrating Al/ML services and models., * TS/SCI FSP Clearance

  • Minimum of 5 years’ experience:
  • Demonstrated experience building production data pipelines and ETL/EL workflows at scale
  • Proficiency with Apache Spark and PySpark for distributed data processing
  • Advanced Python programming skills including data manipulation libraries (Pandas, NumPy) and data engineering best practices
  • Understanding of data security, privacy, governance, and compliance principles
  • Experience with workflow orchestration tools (such as Step Functions, Airflow)
  • Familiarity with containerization (such as Docker or Podman) and deploying data applications in cloud environments
  • Experience with AWS services (S3, Lambda, Step Functions)
  • Experience with PostgreSQL and MySQL in production environments, including performance tuning and schema design
  • Demonstrated experience with SQL and query optimization for complex analytical workloads
  • Experience with version control (Git) and Cl/CD practices for data pipelines
  • Demonstrated ability to work with stakeholders to understand data requirements, assess feasibility, and design appropriate solutions with minimal oversight
  • Strong problem-solving and debugging skills for data quality issues, pipeline failures, and performance bottlenecks

Highly Desired Skills

  • Experience with data Lakehouse architecture using Apache Iceberg
  • Hands-on experience configuring, deploying, and integrating data platform components:
  • Apache Ranger (access control and data governance)
  • Trino (distributed SQL query engine)
  • Data catalogs (Unity Catalog OSS, Apache Polaris, etc.)
  • Apache Superset {data visualization and dashboarding)
  • Proficiency with Bash scripting for automation and data processing tasks
  • Experience with Infrastructure as Code (Terraform or CloudFormation) for data infrastructure
  • Familiarity or experience with tracking data lineage and associated tooling such as Open lineage
  • Familiarity or experience with Java
  • Familiarity with data quality frameworks, testing methodologies, and validation strategies
  • Background with large-scale data migrations or platform modernization efforts
  • Experience integrating Al/ML services and models (translation, OCR, speech-to-text, NLP, language detection, topic modeling), LLMs, and RAG {retrieval-augmented generation) pipelines
  • Familiarity with geospatial data processing {H3, PostGIS, or similar)
  • Contributions to data engineering documentation, best practices, and design patterns
  • Experience with NoSQL databases (DynamoDB, etc.)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all