Copy of Senior Data Engineer

Spokeo, Inc.
United States
about 1 month ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Airflow Amazon Web Services Data Analysis Big Data Information Systems Information Engineering Data Governance Extract Transform Load (ETL) Data Warehousing Distributed Systems Amazon DynamoDB Elasticsearch
+9 more
Data Flow Control Python (Programming Language) SQL Databases Apache Spark Pyspark Information Technology Performance Monitor Non-relational Database Data Pipelines

Job description

Experteer Overview As a Senior Data Engineer at Spokeo, you will build scalable data pipelines and optimize big data workloads using AWS, Spark, and Python. You will collaborate with data science and stakeholders to deliver data products, including entity resolution, and advance our data automation capabilities. You will implement ETL governance, testing, and monitoring to ensure reliable analytics. This remote-first role supports Spokeo’s mission to make data more transparent and actionable. Compensation / Benefits * Build ingestion, processing, and loading data pipelines and automate new components * Collaborate with stakeholders and data science to develop data products including entity resolution * Create unit and stress tests to monitor performance and resolve issues * Develop data analysis tools to extract insights and metrics * Research solutions and maintain technical documentation * Follow data governance, quality, cleansing, and ETL best practices Tasks * 7+ years of data engineering in production environments * Experience with large datasets (>100M records or multi-terabytes) * 5+ years in scalable distributed systems using AWS and EMR * 5+ years Python programming * 5+ years in big data ecosystems; Spark is required (PySpark preferred) * 5+ years SQL, schema design, and dimensional data modeling * 5+ years Airflow or similar dataflow orchestration tools * 2+ years with non-relational databases (e.g., DynamoDB, Elasticsearch) * Bachelor’s degree in Computer Science, Information Systems, Mathematics, or related field Key requirements * bonus program * equity plans * 401(k) * discretionary merit-based salary increases * 100% medical/dental/vision coverage * unlimited employee PTO

Requirements

engineering in production environments * Experience with large datasets (>100M records or multi-terabytes) * 5+ years in scalable distributed systems using AWS and EMR * 5+ years Python programming * 5+ years in big data ecosystems; Spark is required (PySpark preferred) * 5+ years SQL, schema design, and dimensional data modeling * 5+ years Airflow or similar dataflow orchestration tools * 2+ years with non-relational databases (e.g., DynamoDB, Elasticsearch) * Bachelor’s degree in Computer Science, Information Systems, Mathematics, or related field Key requirements * bonus program * equity plans * 401(k) * discretionary merit-based salary increases * 100% medical/dental/vision coverage * unlimited employee PTO

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all