Senior Data Engineer

Insight Global
Bentonville, AR, United States
7 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Agile Methodology Artificial Intelligence Airflow Data Analysis Application Frameworks Big Data BigQuery Cloud Computing Cloud Storage Code Review Continuous Integration
+28 more
Data Architecture Information Engineering Data Governance Extract Transform Load (ETL) Data Transformation Data Migration Data Systems DevOps Distributed Computing Environment Python (Programming Language) Machine Learning Meta-Data Management Metadata Repositories Query Optimization Cloudera SQL Databases Enterprise Data Management Google Cloud Feature Engineering Data Ingestion Sql Optimization Large Language Models Pyspark Data Lineage Deployment Automation Integration Frameworks Data Management Data Pipelines

Job description

We are seeking a highly skilled Senior Data Engineer to join a team responsible for modernizing and transforming enterprise claims data platforms. This role will focus on building scalable data solutions that support analytics, reporting, data science, and AI-driven initiatives across Employee, Finance, Legal, Claims, and other enterprise domains.

The ideal candidate will have deep expertise in Google Cloud Platform (GCP), PySpark, SQL, data modeling, and ETL development, with experience building data pipelines from scratch and migrating legacy applications into modern cloud data architectures.

This team is leading a large-scale claims transformation initiative involving decades of historical data, modernizing legacy processes, implementing new technologies, and creating a scalable foundation for advanced analytics and GenAI solutions.

Responsibilities:

Design, develop, and maintain scalable data pipelines and ETL workflows using GCP technologies.

Build and optimize data ingestion frameworks to process large-scale claims and enterprise data.

Develop data pipelines using PySpark, Airflow, BigQuery, Dataproc, and related cloud technologies.

Translate complex business requirements into scalable technical solutions.

Create logical and physical data models to support reporting, analytics, machine learning, and operational use cases.

Lead migration efforts from legacy applications and codebases into modern cloud-native architectures.

Partner with Data Scientists, Product Managers, Architects, and Business Stakeholders to deliver data-driven solutions.

Implement data quality, monitoring, lineage, governance, and compliance best practices.

Optimize SQL queries and data processing frameworks for performance and scalability.

Build reusable frameworks and components using the team’s internal platform (Helix).

Support Claims Transformation initiatives by structuring and modernizing historical claims data.

Participate in CI/CD implementation, code reviews, testing, and deployment automation.

Evaluate and recommend emerging technologies, including GenAI-enabled data solutions.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

Technical Skills:

5+ years of Data Engineering experience in enterprise-scale environments.

Strong hands-on experience with:

Python

PySpark

SQL

Data Modeling

Experience designing and developing data pipelines from scratch.

Expertise with Google Cloud Platform (GCP), including:

BigQuery

Dataproc

Cloud Storage

Cloud Composer/Airflow

Deep understanding of ETL/ELT concepts and large-scale data processing.

Experience building batch and distributed data processing solutions.

Strong query tuning and SQL optimization skills.

Experience creating logical and physical data models.

Experience implementing data quality and governance frameworks.

Familiarity with CI/CD practices and DevOps methodologies.

Ability to troubleshoot and resolve complex data issues across multiple systems.

Domain Experience:

Experience working with large-scale enterprise datasets.

Ability to interpret complex business requirements and convert them into technical solutions.

Experience supporting Analytics, BI, Data Science, or AI initiatives.

Experience working within Agile development environments. Claims, Insurance, Risk, HR, Finance, Legal, or Healthcare data domain experience.

Experience with data modernization or migration programs.

Exposure to GenAI, LLMs, vector databases, or AI-enabled data solutions.

Experience with enterprise data governance and privacy initiatives.

Knowledge of the Helix framework or similar internal data platforms.

Experience with Data Catalog, Data Lineage, and Metadata Management solutions.

Familiarity with Java development.

Experience working with historical data migration projects involving large legacy datasets.

Exposure to machine learning feature engineering pipelines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all