Apache Spark Developer

Bright Vision Technologies
United States
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$125,000.0 - $185,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Agile Methodology Artificial Intelligence Airflow Amazon Web Services Amazon S3 Apache HTTP Server Microsoft Azure Big Data Catalyst (Software) Cloud Computing Cloud Engineering
+72 more
Cloud Storage Code Review Databases Continuous Integration Information Engineering Data Governance Extract Transform Load (ETL) Data Transformation Data Warehousing Software Debugging Dimensional Modeling Distributed Computing Environment Distributed Systems Fault Tolerance Fraud Prevention and Detection Github Apache Hadoop Hadoop Distributed File System Monitoring of Systems Apache Hive PostgreSQL Machine Learning Microsoft SQL Server Oracle (Applications) Performance Tuning Scrum Methodology Prometheus Cloudera Azure Data Lake Software Engineering Data Streaming Teradata SQL Enterprise Data Management Azure Service Bus Parquet Datadog Data Processing Cloud Platform System Feature Engineering Data Ingestion Apache Yarn Azure Data Factory Sql Optimization Snowflake Grafana Apache Spark Caching Electronic Medical Records Git Spark Mllib Containerization Data Lakes Pyspark Kubernetes Apache Flink AWS Glue Integration Frameworks Apache Kafka Apache Nifi Spark Streaming Machine Learning Operations Presto Tools for Reporting Cloud Migration Terraform Azure Synapse Analytics Data Pipelines Amazon Elastic Mapreduce (EMR) Docker Jenkins Databricks Control M

Job description

We are seeking an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting enterprise analytics, machine learning, real-time reporting, and cloud-based data platforms. This role focuses on building high-performance Spark applications capable of processing billions of records across structured and semi-structured data sources while delivering scalable, reliable, and cost-efficient data pipelines. You will work closely with data architects, data engineers, cloud platform teams, machine learning engineers, and business intelligence developers to build modern data processing solutions leveraging Apache Spark, cloud-native technologies, and distributed computing frameworks. The ideal candidate possesses deep expertise in Spark architecture, distributed systems, performance optimization, and cloud-based big data ecosystems., * Design, develop, and maintain high-performance distributed data processing applications using Apache Spark.

  • Build scalable batch and real-time ETL/ELT pipelines processing large volumes of enterprise data.
  • Develop Spark applications using PySpark, Scala, or Spark SQL for data transformation, aggregation, and analytics.
  • Optimize Spark jobs for memory utilization, partitioning strategies, shuffle performance, and execution efficiency.
  • Process structured, semi-structured, and streaming data from enterprise databases, APIs, Kafka, cloud storage, and data lakes.
  • Develop reusable Spark libraries, data processing frameworks, and metadata-driven ingestion pipelines.
  • Collaborate with cloud engineering teams to deploy Spark workloads on Databricks, EMR, Azure Synapse, or Kubernetes.
  • Implement data quality validation, reconciliation, monitoring, and automated error handling across distributed pipelines.
  • Integrate Spark applications with enterprise data warehouses, lakehouses, and reporting platforms.
  • Participate in architecture reviews, code reviews, technical design discussions, and Agile development activities.
  • Troubleshoot production issues involving distributed processing, cluster performance, resource utilization, and data quality.
  • Support cloud migration initiatives by modernizing legacy ETL workloads into Spark-based architectures., You will be joining a modern data engineering team responsible for building cloud-native big data platforms supporting enterprise analytics, AI, and business intelligence initiatives. Current projects include:
  • Enterprise data lakehouse implementation using Databricks and Delta Lake
  • Real-time streaming analytics processing billions of daily events
  • Large-scale customer analytics and behavioral data platforms
  • Financial risk modeling and fraud detection pipelines
  • Healthcare clinical and operational analytics solutions
  • Cloud migration of legacy Hadoop and ETL workloads
  • Machine learning feature engineering and model training pipelines
  • Enterprise reporting platforms supporting executive dashboards and self-service analytics
  • Distributed data processing infrastructure deployed on Azure and AWS

This is a hands-on engineering role where you will contribute to distributed system architecture, Spark application development, cloud migration, performance optimization, production support, and continuous improvement of enterprise-scale data processing platforms.

Requirements

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Six or more years of professional software or data engineering experience.

  • Four or more years of hands-on Apache Spark development experience in enterprise production environments.
  • Strong proficiency in PySpark, Scala, or Spark SQL for distributed data processing.
  • Deep understanding of Apache Spark architecture including RDDs, DataFrames, Datasets, Catalyst Optimizer, DAG execution, and Tungsten engine.
  • Strong experience with distributed computing concepts including partitioning, shuffling, caching, broadcast joins, and fault tolerance.
  • Advanced SQL skills with databases such as SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.
  • Experience working with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet.
  • Experience processing streaming data using Spark Structured Streaming, Apache Kafka, or Event Hubs.
  • Hands-on experience with cloud platforms including Azure Databricks, AWS EMR, AWS Glue, Azure Synapse Analytics, or Google Dataproc.
  • Experience integrating Spark applications with Delta Lake, Apache Iceberg, or Apache Hudi.
  • Strong understanding of data warehousing concepts, dimensional modeling, and data lake architecture.
  • Experience using Git, CI/CD pipelines, Azure DevOps, GitHub Actions, or Jenkins.
  • Strong debugging, troubleshooting, and Spark performance tuning skills.
  • Experience working in Agile Scrum development environments., * Experience building enterprise Lakehouse architectures using Databricks or Delta Lake.
  • Familiarity with Apache Airflow, Azure Data Factory, AWS Step Functions, or Control-M for workflow orchestration.
  • Experience with machine learning workflows using Spark MLlib, MLflow, or feature engineering pipelines.
  • Knowledge of Kubernetes, Docker, and containerized Spark deployments.
  • Experience implementing Data Quality frameworks using Great Expectations or Deequ.
  • Familiarity with Apache NiFi, Apache Flink, Trino, or Presto.
  • Experience working with cloud object storage including Amazon S3, Azure Data Lake Storage (ADLS Gen2), or Google Cloud Storage.
  • Knowledge of Infrastructure as Code using Terraform or ARM templates.
  • Experience with enterprise monitoring tools including Prometheus, Grafana, Datadog, or OpenTelemetry.
  • Cloud certifications in Azure, AWS, Databricks, or Apache Spark-related technologies are highly desirable.

Benefits & conditions

4.24.2 out of 5 stars Remote $125,000 - $185,000 a year - Full-time

About the company

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all