Mid Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Keep production Databricks pipelines running, including ingestion, transformation, and delivery to downstream consumersDiagnose and resolve pipeline failures and data quality issuesReverse-engineer and document existing transformation logic and business rulesMigrate legacy tables from Hive Metastore to Unity CatalogMaintain Iceberg-enabled table sharing between Databricks and SnowflakeBuild and test dbt models, including incremental materializations and data testsDevelop and maintain Airflow DAGs for orchestrationValidate migrated pipelines against Databricks outputsContribute to Snowflake modeling, performance, and cost decisionsWork directly with client stakeholders on technical topics alongside the team leadRequirements3+ years operating production data pipelinesPySpark and SQL - able to read, debug, and modify existing pipelinesDelta Lake: MERGE/upsert patterns, table properties, OPTIMIZE, partitioningDatabricks Workflows, cluster configuration, job troubleshootingUnity Catalog: catalogs, schemas, grants, lineage, and the metastore modelSnowflake warehouses, roles and grants, and general operating modelSnowflake query performance and awareness of compute cost behaviordbt models, sources, tests, and incremental materializationsdbt project structure and deployment workflowAirflow DAGs, operators, scheduling, and dependency managementAirflow retries, backfills, and idempotent task designStrong SQL, including window functions, complex joins, and reading transformation logicPython for scripting, automation, and API integrationIncremental loading patterns, idempotency, late-arriving data, and reprocessingAWS S3 and IAM basicsBasic working knowledge of Redshift and its role in wider architectureFluent EnglishSelf-directed and able to progress on an unfamiliar codebase without structured onboardingAble to explain production incidents to non-technical stakeholders and provide realistic ETAsCore Competencies Demonstrates expertise in managing production data pipelines using Databricks, including ingestion, transformation, and delivery processes.Proficient in SQL, PySpark, and dbt for building and testing data models, with a strong understanding of Snowflake and Airflow for orchestration and performance optimization.Highest-signal resume keywordsDatabricks Pipeline ManagementSQL ProficiencyPySpark DevelopmentAirflow DAG DevelopmentDbt Model BuildingHard SkillsSQLPySparkDbtAirflowDelta LakeUnity CatalogSnowflakeAWS S3Incremental Loading PatternsData Quality DiagnosisSoft SkillsSelf-DirectedEffective CommunicationStakeholder EngagementIndustry KeywordsData PipelineData TransformationData QualityData ModelingOrchestrationTools & TechnologiesDatabricksSnowflakeAirflowHive MetastoreRedshift#J-*****-Ljbffr
Requirements
3+ years operating production data pipelines PySpark and SQL - able to read, debug, and modify existing pipelines, Airflow DAGs, operators, scheduling, and dependency management Airflow retries, backfills, and idempotent task design Strong SQL, including window functions, complex joins, and reading transformation logic Python for scripting, automation, and API integration Incremental loading patterns, idempotency, late-arriving data, and reprocessing AWS S3 and IAM basics Basic working knowledge of Redshift and its role in wider architecture Fluent English Self-directed and able to progress on an unfamiliar codebase without structured onboarding Able to explain production incidents to non-technical stakeholders and provide realistic ETAs Core Competencies Demonstrates expertise in managing production data pipelines using Databricks, including ingestion, transformation, and delivery processes. Proficient in SQL, PySpark, and dbt for building and testing data models, with a strong understanding of Snowflake and Airflow for orchestration and performance optimization. Highest-signal resume keywords Databricks Pipeline Management SQL Proficiency PySpark Development Airflow DAG Development Dbt Model Building Hard Skills SQL PySpark Dbt Airflow Delta Lake Unity Catalog Snowflake AWS S3 Incremental Loading Patterns Data Quality Diagnosis Soft Skills Self-Directed Effective Communication Stakeholder Engagement Industry Keywords Data Pipeline Data Transformation Data Quality Data Modeling
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Making Data Warehouses Fast: A Developer’s Story
Top Big Data Technologies That You Need to Know
Highest Paying Tech Companies for Developers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again