Senior Data Engineer
CONFIDENTIAL COPY SERVICES
Newark, United States of America
4 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
Senior Compensation
$ 150KJob location
Newark, United States of America
Tech stack
Agile Methodologies
Amazon Web Services (AWS)
Data analysis
Big Data
Business Software
Software Quality
Information Engineering
ETL
Data Transformation
Data Warehousing
Memory Management
Python
Scrum
Power BI
DataOps
SQL Databases
Data Streaming
Systems Integration
Strategies of Testing
Data Processing
Sql Optimization
Snowflake
Spark
PySpark
Semi-structured Data
Information Technology
Optimization Algorithms
Data Analytics
Enterprise Integration
Amazon Web Services (AWS)
Kafka
Data Management
Tools for Reporting
Stream Processing
Data Pipelines
Confluent
Databricks
Job description
- Actively work with business and other BI stakeholders to understand their needs and solution enhancements to the data warehouse
- Be available for addressing issues or bugs in the present implementation
- Design, build, and maintain data infrastructure, ETL processes, and data pipelines.
- Design, build, and maintain data pipelines using Databricks
- Configure and implement data sourcing events on Kafka
- Build and optimize python/Pyspark framework
- Maintain Snowflake and AWS data infrastructure
- Work with Unstructured/Semi-Structured data and build reporting tables
- Assist the BI Analyst to manage client reporting requirements
- Build Data models and assist the BI Analyst to offer data-driven insights
- Design data models and automate manual processes.
- Create and execute a test strategy to ensure robustness of data pipelines
- Ensure effectiveness in infrastructure consumption by implementing solutions that are optimized and scalable
- Foster an environment that emphasizes trust, open communication, creative thinking, and cohesive team effort, across the business, IT and vendor teams
Requirements
- Proficient in Databricks Spark: Exceptional skills in Databricks Spark for sophisticated data processing. Proven experience in leveraging Spark for complex ETL tasks, surpassing traditional data processing methods.
- ETL Pipeline Mastery: Demonstrated excellence in designing and implementing ETL pipelines specifically within Databricks. Ability to utilize Spark's full capabilities to create efficient, scalable data pipelines.
- Data Transformation and Analysis: Expert in data transformation using Databricks Spark, skilled in performing advanced data analytics and processing large datasets with high efficiency.
- Optimization Techniques: Deep understanding of optimizing Databricks Spark applications for maximum performance, including fine-tuning Spark configurations, and memory management.
Integration with Confluent Kafka:
- Kafka Exposure: Solid background in working with Confluent Kafka, particularly in integrating it with Spark-based systems for real-time data streaming and processing.
- Efficient Data Pipelines: Proficiency in creating and managing data pipelines that seamlessly integrate Kafka with Databricks Spark, ensuring efficient data flow and processing.
- Supporting integrations with business applications, using Kafka
DataOps and Agile Methodologies:
- DataOps Principles: Strong grasp of DataOps methodologies, with a focus on improving the efficiency and quality of data analytics via automation, collaboration, and process optimization.
- Agile Development: Experienced in Agile software development practices, adept at implementing Agile methodologies like Scrum or Kanban in data-centric projects for improved collaboration and rapid delivery.
Proficiency in SQL and Platform Integrations:
- Advanced SQL Skills: Expertise in SQL, particularly for querying and managing tables/data warehouses within Snowflake. Ability to seamlessly integrate these with Databricks Spark.
- Platform Adaptability: Skilled in adapting to and integrating various data platforms and technologies, aligning them with strategic organizational goals.
Collaborative and Best Practice-Oriented:
- Adherence to High Standards: Committed to maintaining high standards in code quality, documentation, and adhering to DataOps and Agile best practices.
- Team Collaboration and Leadership: Ability to work collaboratively in a team, fostering a culture of continuous learning and improvement., * Bachelor's or master's degree in computer science, statistics, or analytics.
- Over 8 years of experience in the field of data engineering with capabilities on working through the entire development lifecycle
- Ability to work with senior business stakeholders to be able to ascertain data needs out of business opportunities or challenges
- Minimum of 5 years of cloud experience on any of the major cloud platforms preferably AWS
- Hands-on Knowledge of Data Modeling, Warehousing and Power BI (or equivalent) as a Reporting Tool
- Advanced proficiency in SQL with the ability to read and write queries is required
- Demonstrable experience with Databricks, Snowflake, Python, and PySpark
- Preferably to have experience with Kafka and Power BI