Senior Data Engineer- US Citizen

VYTWO TECHNOLOGIES INC.
Prosper, TX, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Batch Processing Databases Data Validation Information Engineering Extract Transform Load (ETL) Database Queries Performance Tuning Apache Spark Git Data Lakes Pyspark Software Version Control
+2 more
Data Pipelines Databricks

Job description

Design, develop, and maintain scalable data pipelines using Apache Spark (PySpark and/or Scala) Build and optimize data workflows on Databricks, including Delta Lake, notebooks, and scheduled jobs Ingest, transform, and curate large-scale structured and semi-structured datasets Perform performance tuning and cost optimization of Spark workloads and Databricks clusters Implement data quality checks, monitoring, and error handling Collaborate with analytics and business stakeholders to deliver well-modeled, analytics-ready data Support batch processing and, where applicable, streaming data pipelines Follow best practices for testing, documentation, security, and version control

Requirements

5+ years of experience in Data Engineering or related roles Strong hands-on experience with Apache Spark (PySpark or Scala) Proven experience working in Databricks environments Strong SQL skills and experience with relational and analytical databases Experience building and maintaining ETL/ELT pipelines at scale Familiarity with modern data lake architectures (Delta Lake preferred) Experience with Git-based version control Must be a US Citizen Flexible work from home options available.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all