Data Engineer

Themesoft Inc
Chandler, IN, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$83,200.0 - $93,600.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Big Data BigQuery Cloud Database Cloud Storage Information Engineering Python (Programming Language) Metadata Software Tools Data Streaming Google Cloud
+8 more
Apache Spark Pyspark Apache Flink Apache Kafka Spark Streaming Data Management Data Lakehouse Data Pipelines

Job description

Implement and operationalize modern AI-enabled data capabilities on Google Cloud to ingest, transform, and distribute data for a variety of big data apps Leverage AI/Agentic frameworks to automate data management, governance, and data consumption capabilities - data pipelines, data quality, metadata, data compliance, etc. Demonstrable skills (recent) using AI tools such as LangChain, LangGraph/ADK, agentic frameworks, RAG, GraphRAG, and using MCP to build agent-based data capabilities 5 plus years of experience in data engineering including hands-on experience working with Cloud data solutions: creating/supporting Spark based ingestion and processing 3 plus years of experience with Data lakehouse architecture and design, including hands-on experience with Python, pySpark, Kafka, Airflow, Google Cloud Storage, BigQuery, Data Proc, Cloud Composer Hands-on experience developing data flows using Kafka, Flink, and Spark streaming, Implement and operationalize modern AI-enabled data capabilities on Google Cloud to ingest, transform, and distribute data for a variety of big data apps Leverage AI/Agentic frameworks to automate data management, governance, and data consumption capabilities - data pipelines, data quality, metadata, data compliance, etc. Work within a matrix org. with principal engineers, product managers, and data engineers to roadmap, plan, and deliver key data capabilities based on priority

Requirements

GCP - 5 to 6 years. AI exposure - 6 months to a year Demonstrable skills (recent) using AI tools such as LangChain, LangGraph/ADK, agentic frameworks, RAG, GraphRAG, and using MCP to build agent-based data capabilities 5 plus years of experience in data engineering including hands-on experience working with Cloud data solutions: creating/supporting Spark based ingestion and processing 3 plus years of experience with Data Lakehouse architecture and design, including hands-on experience with Python, pySpark, Kafka, Airflow, Google Cloud Storage, Big Query, Data Proc, Cloud Composer Hands-on experience developing data flows using Kafka, Flink, and Spark streaming

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all