Data Engineer

Bartech Staffing
Charlotte, NC, United States
25 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Amazon S3 Business Logic Continuous Integration Data Deduplication Data Governance Data Integrity Extract Transform Load (ETL) Data Warehousing Github Identity and Access Management
+18 more
Python (Programming Language) Operational Databases Regression Testing Standard Sql Amazon Simple Notification Service (SNS) User Environment Management Delivery Pipeline State Machines Pytest Data Lakes Pyspark Apache Kafka Cloudwatch Amazon Simple Queue Service (SQS) Terraform Data Pipelines Confluent Amazon Redshift

Job description

  • Collaborate with Lead Developers (Data Engineer, Software Engineer, Data Scientist, Technical Test Lead) to understand requirements and use cases, outline technical scope, and deliver technical solutions
  • Collaborate with Data and Solution architects on key technical decisions
  • Develop data pipelines with focus on long-term reliability and maintaining high data quality
  • Design data lake and warehousing solutions with the end-user in mind, ensuring ease of use without compromising on performance
  • Manage and resolve issues in production data warehouse environments on AWS

Core Experience and Abilities:

  • Perform hands-on development and peer review for certain components and tech stack
  • Set up development instances and migration paths with required security, access, and roles
  • Develop components and related processes (e.g., data pipelines, ETL processes, workflows)
  • Build new data pipelines, identify existing data gaps, and provide automated solutions to deliver analytical capabilities and enriched data to applications
  • Implement data pipelines with attentiveness to durability and data quality
  • Implement data warehousing products with focus on end-user experience (ease of use with appropriate performance)
  • Implement data quality frameworks and validation rules (e.g., schema validation, null checks, referential integrity, deduplication)
  • Design and implement automated data tests for ETL pipelines using Python, PySpark, and SQL
  • Write unit, integration, and regression tests for data pipelines (e.g., pytest-based testing for transformations and business rules)
  • Demonstrate familiarity with data observability and monitoring concepts, including freshness, volume, and anomaly detection
  • Understand data reconciliation and source-to-target validation techniques to ensure business logic accuracy Embed data quality checks into CI/CD pipelines to prevent defective data from reaching downstream consumers *

Requirements

  • 2+ years of AWS experience
  • AWS services: S3, EMR, Glue Jobs, Lambda, Athena, CloudTrail, SNS, SQS, CloudWatch, Step Functions, Redshift
  • Experience with Kafka, preferably Confluent Kafka
  • Experience with Lake Formation, Amazon Redshift and Amazon Athena
  • Strong SQL and data modeling skills, executing ETL processes tailored for data warehousing
  • Competence in developing and refining data pipelines within AWS
  • Extensive understanding of database management fundamentals
  • Tools and Languages: Python (good experience in PySpark), SQL
  • Infrastructure as Code technology: Terraform
  • DevOps pipeline (CI/CD): GitHub
  • Deep knowledge of IAM roles and policies
  • Experience with AWS workflow orchestration tools like Airflow or Step Functions

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:05 min

Tagging and organizing execution scenarios with pytest markers

Florian Bruhin · World Congress 2021

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all