> Markdown version of [/jobs/ext/609878-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/609878-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Alight, Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Big Data, Cloud Computing, Cloud Database, Cloud Engineering, Cloudera Impala, Code Review, Continuous Integration, Extract Transform Load (ETL), Data Profiling, Data Warehousing, Distributed Computing Environment, Distributed Systems, Github, Apache Hadoop, Hadoop Distributed File System, Apache Hive, Identity and Access Management, Subnetting, Python (Programming Language), Routing, Performance Tuning, Shell Script, Simple Data Format, SQL Databases, Sqoop, Data Streaming, Parquet, Data Logging, Data Ingestion, Apache Yarn, Autoscaling, Snowflake, Apache Spark, State Machines, Amazon Virtual Private Cloud (VPC), Git, Containerization, Data Lakes, Pyspark, Optimization Algorithms, AWS Data Analytics, Apache Kafka, Data Management, Interactive Whiteboards, Data Pipelines, Serverless Computing, Docker, Programming Languages, Control M - **Published:** June 12, 2026 - **Apply:** https://www.juju.com/job/00000000g7g9h9 ## About the Role Technical Skills + Strong experience from 5-8 eyars with the **Hadoop ecosystem** (HDFS, Hive, Spark, YARN, Kafka). + Strong hands-on expertise in Scala, **PySpark** , Spark optimization techniques, HiveQL, and distributed computing. + Good work experience in SQL in hive and impala + Good understanding of **AWS data stack** (S3, Glue, EMR, Lambda, Kinesis, Redshift, Step Functions). + Proficiency in at least one scripting/programming language: **Python, Shell scripting** . + Strong experience with **CI/CD** , **GitHub, Git commands.** + Expertise in ETL and Data Warehousing and cloud concepts. + Good understanding of data modelling (star/snowflake), partitioning strategies, and schema evolution. + Expertise in data profiling and decision making. + Able to understand, design and create data flow diagrams and do data modelling. (knowledge of Miro will be added advantage) + Able to understand the architecture and design end-to-end data flow. + Hands-on experience with **Airflow, Control-M** , or other orchestrators. + To monitor and support BAU and year end activities, if needed. + Well versed with security and compliance aspects in Cloud. + Good understanding of AWS networking (VPC, subnets, routing, SGs, NACLs). + Familiarity with serverless patterns and containerization (Docker, ECS/EKS). + Experience with monitoring/logging tools and incident management practices. Other Requirements + Strong logical and analytical, problem-solving, and communication skills. + Communicate effectively and concisely with multiple stakeholders and coordinate and collaborate with cross functional teams. + Ability to support both legacy Hadoop workloads and cloud-first architectures. + AWS certifications (Data Engineer, Solutions Architect, or Developer) are a plus. + Good to have health care domain knowledge. ## Description Senior Data Engineer with strong expertise across traditional big-data platforms (Hadoop ecosystem) and modern cloud-native architectures (AWS). Responsible for building scalable, secure, and high-performance data pipelines that span Hadoop clusters and AWS cloud services. Leverages deep knowledge of distributed systems, Spark optimization, cloud automation, and big-data management to support analytics, BI, ML, and AI use cases across the enterprise. Ensures reliability, governance, cost-efficiency, and operational excellence across hybrid data platforms. Associate should be self-driven, can work with minimal guidance and guide the team technically. Core Responsibilities + Design, build, and maintain high-volume **ETL/ELT** pipelines across **Hadoop (HDFS, Hive, Spark, Kafka)** and **AWS (Glue, EMR, Lambda, Step Functions, Redshift)** . + Develop distributed data processing solutions using **PySpark, Spark SQL** , and scalable cloud serverless patterns. + Implement reusable data ingestion frameworks for batch (Sqoop, Hive, Spark) and streaming (Kafka, Kinesis). + Optimize data workflows using partitioning, bucketing, compression, file formats (Parquet/ORC). + Understanding hybrid data lake architectures using **S3 + HDFS** , ensuring governance consistency (Atlas, Ranger, Lake Formation). + Understanding the reporting requirements and perform data profiling and create design for same. + Create data flow diagram and do data modelling. + Job orchestration using **Airflow, Control-M, Step Functions** , or event-driven triggers. + Understand auto-scaling, capacity planning, and performance tuning on EMR and Spark clusters. + Ensure data is protected and compliant with regulatory standards. + Work closely with business stakeholders to enable high-quality datasets. + Provide technical leadership in architecture decisions, code reviews, and best-practice adoption and provide technical guidance to peers/juniors in team. + Improve reliability, scalability, and performance through automation, autoscaling, and capacity planning. + Own deployment, incident response, and post-incident reviews for production environments, troubleshooting Spark performance issues, job failures, and cluster bottlenecks. + Understanding security best practices (IAM, KMS, security groups, WAF, parameter/secret management). + Optimize cost and usage of AWS resources and recommend architecture improvements. + Collaborate closely with developers, QA, and product teams to streamline release processes. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)