Data Engineer

Cornerstone OnDemand
Boston, MA, United States
13 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Sql Data Warehouse Java (Programming Language) Airflow Amazon Web Services Microsoft Azure Bash Shell Big Data C++ (Programming Language) Cloud Computing Data Control Data Infrastructure Extract Transform Load (ETL)
+28 more
Data Systems Data Warehousing Relational Databases Database Design Distributed Systems Apache Hadoop MapReduce Apache Hive Python (Programming Language) Machine Learning MongoDB Neo4j NoSQL Object-Oriented Software Development Raw Data Redis Standard Sql Scala (Programming Language) Workflow Management Systems Scripting Google Cloud Data Ingestion Snowflake Apache Spark Cassandra Presto Data Pipelines Databricks

Job description

We are seeking a talented Data Engineer with strong communication skills, passion for solving business problems with data and has domain knowledge in Finance, Human Resources, and Customer Success. You have empathy, curiosity and desire to improve and constantly learn. Should be hands-on with dbt, Snowflake, Airflow, Fivetran, and has a proven track record of driving the best practices and processes, building data models and ETL loads.

In this role you will…

· Design, build and maintain batch or real-time data pipelines in production.

· Maintain and optimize the data infrastructure required for accurate extraction, transformation, and loading of data from a wide variety of data sources.

· Develop ETL (extract, transform, load) processes to help extract and manipulate data from multiple sources.

· Automate data workflows such as data ingestion, aggregation, and ETL processing.

· Prepare raw data in Data Warehouses into a consumable dataset for both technical and non-technical stakeholders.

· Partner with data scientists and functional leaders in sales, marketing, and product to deploy machine learning models in production.

· Build, maintain, and deploy data products for analytics and data science teams on cloud platforms (e.g. AWS, Azure, GCP).

· Ensure data accuracy, integrity, privacy, security, and compliance through quality control procedures.

· Monitor data systems performance and implement optimization strategies.

· Leverage data controls to maintain data privacy, security, compliance, and quality for allocated areas of ownership.

Requirements

· 3+ years of SQL skills and experience with relational databases and database design.

· Experience working with cloud Data Warehouse solutions - Databricks, Apache Spark

· Experience working with data ingestion tools such as Fivetran, stitch, or Matillion.

· Working knowledge of Cloud-based solutions (e.g. AWS, Azure, GCP).

· Experience building and deploying machine learning models in production.

· Strong proficiency in object-oriented languages: Python, Java, C++, Scala.

· Strong proficiency in scripting languages like Bash.

· Strong proficiency in data pipeline and workflow management tools (e.g., Airflow).

· Strong project management and organizational skills.

· Excellent problem-solving, communication, and organizational skills.

· Proven ability to work independently and with a team.

Extra dose of awesome if you have…

· Good understanding of NoSQL databases like CrateDB, Redis, Cassandra, MongoDB, or Neo4j.

· Experience with working on large data sets and distributed computing (e.g. Hive/Hadoop/Spark/Presto/MapReduce).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all