Data Engineer Databricks / Spark / AI

Innova Software Services Inc.
San Francisco, CA, United States
4 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Agile Methodology Artificial Intelligence Business Analytics Applications Data Analysis JIRA Big Data Code Review Databases Information Engineering Data Files Extract Transform Load (ETL) Data Transformation
+20 more
Distributed Data Store Python (Programming Language) Performance Tuning Scrum Methodology Systems Development Life Cycle Requirements Management Software Engineering SQL Databases Enterprise Data Management Data Processing Scripting Freeform SQL Apache Spark Generative AI Pyspark Atlassian Tools Data Analytics Data Management Data Pipelines Databricks

Job description

  • Design, develop, and maintain scalable data engineering solutions using Databricks, Apache Spark, SQL, and Python.
  • Build and optimize high-volume ETL/ELT data pipelines and data-processing workflows.
  • Develop complex SQL queries and transformations for large-scale datasets.
  • Build distributed data-processing solutions using PySpark/Spark.
  • Design reliable and scalable data architectures within the Databricks ecosystem.
  • Work hands-on with Databricks Genie to enable AI-powered conversational data and analytics capabilities.
  • Integrate AI/GenAI capabilities with enterprise data platforms and analytics workflows.
  • Optimize Databricks workloads for performance, scalability, reliability, and cost efficiency.
  • Perform data transformation, cleansing, validation, and quality checks.
  • Troubleshoot performance issues across Spark jobs, SQL workloads, pipelines, and Databricks environments.
  • Work with structured and semi-structured datasets from multiple enterprise sources.
  • Collaborate with Data Engineering, Analytics, AI/ML, Product, and business teams.
  • Translate business and analytical requirements into scalable data solutions.
  • Participate in technical design, code reviews, testing, deployment, and production support.
  • Maintain engineering standards, documentation, and data-development best practices.
  • Participate in Agile/Scrum development activities using Jira and Confluence.

Requirements

We are seeking a highly skilled Senior Data Engineer with strong, recent hands-on experience in Databricks, Apache Spark, SQL, and Python.

This is primarily a Data Engineering role with additional exposure to AI and Generative AI capabilities within the Databricks ecosystem. The ideal candidate will have extensive experience designing and developing scalable data pipelines and data processing solutions using Databricks and Spark.

Candidates should also have hands-on experience with Databricks Genie and an understanding of how AI capabilities can be applied to enterprise data and analytics solutions.

Strong, recent Databricks experience is essential for this position., Agile Programming Methodologies, Apache Spark, Artificial Intelligence (AI), Atlassian JIRA, Code Reviews, Data Analysis, Data Management, Data Processing, Data Sets, Database Extract Transform and Load (ETL), Ecosystems, Identify Issues, Performance Tuning/Optimization, Production Support, Python Programming/Scripting Language, Requirements Management, SQL (Structured Query Language), Scalable System Development, Scrum Project Management and Software Development, Technical/Engineering Design, Testing, Validation Testing, Workflow Analysis

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all