Data Engineer - INTL India

Insight Global
United States
17 days ago
Apply on jobs.insightglobal.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$33,280.0 - $41,600.0
Working hours
Regular working hours

Tech stack

Airflow Data Analysis Microsoft Azure Big Data BigQuery Cloud Computing Profiling Computer Programming Data Validation Extract Transform Load (ETL) Fault Tolerance Apache Hadoop
+20 more
Apache Hive Python (Programming Language) Performance Tuning Query Optimization Cloud Services Cloudera Software Construction SQL Databases Data Streaming Workflow Management Systems Data Processing Google Cloud Cloud Platform System Data Ingestion Sql Optimization System Availability Apache Spark Apache Kafka Software Version Control Data Pipelines

Job description

Must report onsite to the Bangalore Office 5 days a week**, Design, develop, and maintain batch data pipelines using Apache Spark, Hadoop, Hive, or similar frameworks in a cloud environment. Build highly optimized, fault-tolerant, and SLA-driven data pipelines that operate reliably at scale. Leverage Google Cloud Platform (GCP) services such as BigQuery, GCS, Dataproc, and Pub/Sub to support data ingestion, processing, and storage. Write and optimize SQL queries (BigQuery SQL and/or Spark SQL) for data analysis, profiling, and performance tuning. Collaborate closely with analytics, data science, and downstream consumers to ensure data availability, correctness, and usability. Monitor and troubleshoot pipeline failures; implement alerting, retries, and data quality checks. Improve pipeline performance through partitioning, clustering, resource tuning, and query optimization. Follow software engineering best practices, including version control, testing, and documentation.

Requirements

We are seeking a skilled Data Engineer to design, build, and optimize large-scale batch data pipelines in a cloud environment. This role focuses on reliability, performance, and data quality, supporting analytics and downstream consumers through well-engineered big-data solutions. The ideal candidate has strong experience with Apache Spark, cloud data platforms (GCP preferred), and writing performant SQL at scale., Strong experience with big data technologies such as Apache Spark, Hadoop, and Hive. Hands-on experience building batch data pipelines with a focus on performance, scalability, SLA adherence, and fault tolerance. Strong programming skills in Python and/or Scala, with deep experience using Spark for data processing and analytics. Experience working with GCP services including BigQuery, Google Cloud Storage (GCS), Dataproc, and Pub/Sub. Solid experience writing and optimizing SQL, preferably BigQuery SQL and/or Spark SQL. Strong understanding of data modeling, ETL/ELT patterns, and data quality best practices. Experience with Kafka or similar messaging/streaming platforms. Familiarity with workflow orchestration tools (e.g., Airflow or Cloud Composer). Experience deploying and operating data pipelines in production cloud environments (GCP preferred, Azure acceptable). Strong troubleshooting skills and ability to optimize pipelines under real-world constraints.

Benefits & conditions

Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.insightglobal.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

56 sec

Introduction to analytical data formats for software developers

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all