Data Engineer- Veterans Affairs

ThunderYard Solutions LLC
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$100,000.0 - $120,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow Data Analysis Microsoft Azure Cloud Storage Continuous Integration Data Architecture Data Validation Data Deduplication Data Governance Extract Transform Load (ETL) Data Transformation
+29 more
Data Systems Data Warehousing Relational Databases Python (Programming Language) Query Optimization Power BI Standard Sql Software Deployment SQL Databases Data Streaming Tableau (Software) Workflow Management Systems Privacy Controls Azure Service Bus Cloud Platform System Azure Data Factory Apache Spark Git Data Lakes Pyspark Real Time Data Apache Kafka Data Management Tools for Reporting Azure Synapse Analytics Looker Analytics Software Version Control Data Pipelines Databricks

Job description

ThunderYard Solutions is seeking a Data Engineer to support the U.S. Department of Veterans Affairs in designing, developing, and maintaining scalable data solutions that support mission-critical healthcare and business operations. The ideal candidate will have expertise in data architecture, ETL/ELT processes, cloud-based data platforms, and analytics technologies, with a strong commitment to delivering secure, high-quality data solutions in a federal environment., * Design, develop, and maintain ETL/ELT pipelines to ingest, transform, and load data from multiple sources such as APIs, relational databases, cloud storage, and streaming platforms

  • Build scalable batch and near real time data pipelines using Databricks and Apache Spark (PySpark / SQL)

  • Implement data transformation logic following best practices for performance, reliability, and reusability

  • Support schema evolution, data validation, deduplication, and error handling in ETL workflows

Databricks Platform Development

  • Develop and optimize pipelines using Delta Lake and medallion (Bronze / Silver / Gold) architecture patterns

  • Use Databricks Workflows / Jobs or similar orchestration tools to schedule and monitor pipelines

  • Optimize Spark jobs for performance and cost (partitioning, caching, file sizing, query tuning)

  • Collaborate on data governance initiatives using Unity Catalog, access controls, and lineage where applicable

Collaboration & Operations

  • Work closely with data architects, analytics teams, and downstream consumers to define data requirements

  • Troubleshoot pipeline failures and data quality issues and implement long term fixes

  • Produce documentation for pipelines, datasets, and operational runbooks

  • Participate in CI/CD practices using Git based version control for notebooks and code deployments

Requirements

This role will collaborate with cross-functional teams, including data analysts, software developers, and government stakeholders, to optimize data pipelines, improve data accessibility, and ensure compliance with federal security and privacy standards. The successful candidate will demonstrate strong problem-solving skills, technical leadership, and the ability to work effectively in an agile environment supporting veteran-focused initiatives., * 3+ years of experience as a Data Engineer or in a similar data focused role

  • Hands on experience with Databricks

  • Strong experience building ETL/ELT pipelines

  • Proficiency in Python and SQL

  • Experience with Apache Spark / PySpark

  • Experience with Azure, Azure Synapse Analytics, and Azure Data Factory

  • Solid understanding of data modeling, data warehousing, and analytics use cases, Preferred / Nice to Have
  • Experience with Delta Live Tables (DLT) or Databricks Auto Loader

  • Experience with orchestration tools such as Airflow

  • Familiarity with streaming data technologies (Kafka, Event Hubs, Kinesis)

  • Experience supporting analytics tools (Power BI, Tableau, Looker) connected to Databricks

  • Databricks certification (Associate or Professional), * Azure: 3 years (Required)
  • Databricks: 3 years (Required)
  • Azure Data Factory: 3 years (Required)

Benefits & conditions

$100,000 - $120,000 a year - Full-time, Pulled from the full job description

  • Professional development assistance
  • Tuition reimbursement
  • 401(k)
  • Health insurance
  • Retirement plan
  • 401(k) matching
  • Paid time off, The salary is budgeted at $100,000-$120,000 annually, plus benefits. ThunderYard offers benefits including medical, dental and vision insurance, 401k matching, PTO, certification reimbursement and more.

Vetting:

Candidates selected will be subject to a background investigation for clearance eligibility by our government client.

ThunderYard Solutions is proud to be an Equal Opportunity Employer. We don’t just accept difference - we celebrate it, we support it, and we thrive on it for the benefit of our employees, our community, and our customers. All applicants will be considered for employment without discrimination of race, color, religion, or belief, national, social, or ethnic origin, sex, age, physical, mental, or sensory disability, HIV status, sexual orientation, gender identity and/or expression, marital, civil union, or domestic partnership status, protected veteran status, family medical history or genetic information.

Pay: $100,000.00 - $120,000.00 per year, * 401(k)

  • 401(k) matching
  • Dental insurance
  • Flexible spending account
  • Health insurance
  • Health savings account
  • Life insurance
  • Paid time off
  • Professional development assistance
  • Retirement plan
  • Tuition reimbursement
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all