Databricks Data Engineer

NewGen Technologies
United States
4 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Agile Methodology Amazon Web Services Cloud Computing Information Engineering Extract Transform Load (ETL) Data Warehousing Apache Hive SQL Databases Data Ingestion IT Architecture SC Clearance Data Lakes
+2 more
Pyspark Databricks

Job description

  • Build and maintain Databricks pipelines (60% of the time)
  • Gather and analyze requirements (20% of the time)
  • Monitor and troubleshoot data jobs (10% of the time)
  • Document processes and workflows (10% of the time)
  • Perform additional assigned tasks

Requirements

  • US Citizenship
  • Secret Clearance
  • Excellent interpersonal skills and ability in overseeing/prioritizing work requests
  • Excellent written and oral communication skills
  • Strong data engineering, ETL/ELT development, or data warehousing
  • Strong PySpark and/or Spark SQL skills; solid SQL fundamentals
  • Solid understanding of the Medallion Architecture (Bronze/Silver/Gold layering)
  • Proficiency with Delta Lake (ACID transactions, schema evolution, time travel, MERGE/Upsert patterns)
  • Experience designing both batch and incremental/streaming ingestion pipelines
  • Familiarity with the Agile process

Desired Skills

  • Databricks Certification/experience/familiarity preferred
  • Knowledge of and experience with the DoD
  • Familiar with Customer and/or DoD IT architectures
  • Federal experience
  • Experience in post go-live production support and break/fix
  • Focus on customer service and responsiveness, with sound customer handling skills
  • Experience working within AWS Gov Cloud
  • Experience working migrating code from non-accredited to accredited environments within AWS

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:24 min

The governance failures of centralized data lakes

Mario Meir-Huber · LIVE

5:14 min

Executing Databricks jobs with built-in Airflow operators

Alan Mazankiewicz · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:38 min

Overview of Databricks and interactive data processing

Alan Mazankiewicz · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all