Data Engineer

Phaxis LLC
Edison, NJ, United States
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$140,000.0 - $170,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon S3 Data Analysis Microsoft Azure Big Data BigQuery Cloud Computing Databases Continuous Delivery Continuous Integration
+33 more
Information Engineering Data Files Data Integration Extract Transform Load (ETL) Dataspaces Data Systems Data Warehousing DevOps Distributed Computing Environment Electronic Data Interchange (EDI) Python (Programming Language) Performance Tuning Systems Development Life Cycle Software Engineering SQL Databases Data Streaming Enterprise Data Management Data Processing Scripting Enterprise Software Applications Data Storage Technologies Snowflake Apache Spark Reliability of Systems Data Lakes Pyspark Information Technology Deployment Automation Data Management Data Pipelines Serverless Computing Databricks Programming Languages

Job description

This role focuses on building efficient data pipelines, implementing best practices in data engineering, and ensuring system reliability and performance. The Data Engineer works closely with data scientists, analysts, and software engineers to deliver optimized data solutions for both batch and real-time processing., * Design, develop, and maintain end-to-end data pipelines and ETL/ELT workflows using Python, SQL, and modern orchestration tools (e.g., Airflow, dbt).

  • Architect and optimize data storage solutions, including data lakes and data warehouses, using Databricks, Delta Lake, and cloud-native services.
  • Build scalable data processing solutions leveraging Databricks notebooks, jobs, and clusters for both batch and streaming data workloads.
  • Develop and manage Databricks workflows using Spark (PySpark, SQL, or Scala) to transform, cleanse, and aggregate large datasets.
  • Design, develop, and implement complex data integrations across Databricks, cloud platforms (AWS, Azure, or GCP), enterprise applications, APIs, and modern data ecosystems to enable reliable, scalable data exchange and processing.
  • Implement data quality checks, schema validation, and monitoring to ensure data accuracy and reliability.
  • Optimize Databricks cluster configurations and job performance to minimize cost and maximize throughput.
  • Collaborate with DevOps teams to automate deployments, CI/CD pipelines, and infrastructure-as-code (IaC) for data systems.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related technical field.
  • Advanced proficiency in Python and SQL for data manipulation, transformation, and automation.
  • Deep experience with Databricks, including Spark optimization, Delta Lake management, job orchestration, and workspace administration.
  • Strong understanding of distributed data processing, partitioning strategies, and performance tuning in Databricks and Spark.
  • Demonstrated experience designing and implementing enterprise-scale data integrations across Databricks, cloud platforms (AWS, Azure, or GCP), enterprise applications, APIs, and modern data ecosystems.
  • Hands-on experience with cloud platforms (AWS, Azure, or GCP) and their data ecosystems (e.g., S3, ADLS, BigQuery, Snowflake).

Skills: Amazon Web Services (AWS), Application Programming Interface (API), Automation, Best Practices, Cloud Computing, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Data Analysis, Data Management, Data Partitioning, Data Processing, Data Quality, Data Science, Data Sets, Data Storage, Data Warehousing, Database Extract Transform and Load (ETL), DevOps, Ecosystems, Enterprise Applications, Enterprise Data Integration, GCP (Good Clinical Practices), Microsoft Windows Azure, Performance Tuning/Optimization, Python Programming/Scripting Language, SQL (Structured Query Language), Scala Programming Language, Scalable System Development, Software Engineering, Systems Reliability, Validation Testing

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

Videos

See all

Related articles

See all