> Markdown version of [/jobs/ext/86493-lead-aws-data-engineer](https://www.wearedevelopers.com/jobs/ext/86493-lead-aws-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead AWS Data Engineer - **Company:** Fannie Mae - **Location:** Reston, VA, United States - **Experience:** Expert - **Salary:** $141,000.0 - $184,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Business Analytics Applications, Data Analysis, Big Data, Cyber Security, Continuous Delivery, Continuous Integration, Data as a Services, Data Architecture, Data Validation, Information Engineering, Extract Transform Load (ETL), Data Mapping, Data Security, Data Systems, Software Debugging, Distributed Computing Environment, Memory Management, Machine Learning, Performance Tuning, Standard Sql, Data Streaming, Data Processing, Data Ingestion, Large Language Models, Apache Spark, Electronic Medical Records, Generative AI, Data Lakes, Pyspark, Deployment Automation, AWS Glue, Data Analytics, AWS Data Analytics, Data Management, Data Pipelines - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=9bf613ae01e84c7c ## About the Role We are seeking an experienced Lead AWS Data Engineer to lead the design, development, and optimization of large-scale AWS-based Data Lake and data pipeline solutions. The ideal candidate will have deep expertise in AWS data services (EMR, Glue, S3, Athena, Lambda) and PySpark-based data processing, along with strong experience in data modeling, performance optimization, and data pipeline orchestration. This role requires strong technical leadership and stakeholder collaboration, with the ability to analyze complex datasets, perform data mapping across systems, and translate business requirements into scalable data engineering solutions., 4+ years of experience in Data Engineering / Big Data development. 4+ years of hands-on experience with AWS data services. Strong experience with AWS EMR, AWS Glue, S3 Data Lakes, Athena / Redshift / Lakehouse architectures. Expertise in PySpark and Spark-based distributed processing. Basic understanding or exposure to Generative AI concepts and AWS Bedrock services. Strong experience building large-scale data pipelines. Proven experience with EMR performance tuning and debugging. Experience with data mapping and integration across heterogeneous datasets. Strong SQL and data modeling skills. Excellent communication and stakeholder management skills. Desired Experiences: Bachelor degree or equivalent 10+ years of experience in Data Engineering / Big Data development. 5+ years of hands-on experience with AWS data services. Hands-on experience with AWS Bedrock, including working with foundation models and building GenAI-powered data solutions. Experience integrating AI/ML or Generative AI capabilities into data pipelines or analytics platforms. AWS Certification (e.g., AWS Certified Data Analytics, Machine Learning Specialty, or AI/ML-related certifications) preferred. Experience with AWS Step Functions. Experience with Data Lake governance tools (Lake Formation, Glue Catalog). Knowledge of data security and compliance frameworks. Experience implementing CI/CD pipelines for data platforms., Education: Bachelor's Level Degree (Required) The future is what you make it to be. Discover compelling opportunities at Fanniemae.com/careers. For most roles, employees are expected to work onsite on a regular basis at their designated office location. In-office work cadence is determined by your manager. Proximity within a reasonable commute to your designated office location is preferred unless the job is noted as open to remote. ## Description The Lead AWS Data Engineer role will offer you the flexibility to make each day your own, while working alongside people who care so that you can deliver on the following responsibilities: Data Engineering & Architecture Design, build, and maintain scalable AWS Data Lake architectures using services such as S3, EMR, Glue, Athena, and Lambda. Develop and optimize data pipelines and ETL/ELT workflows using PySpark, AWS Glue, and EMR. Implement high-performance distributed data processing solutions for large-scale datasets. Develop frameworks for data ingestion, transformation, validation, and publishing within the data lake ecosystem. Performance Optimization Diagnose and resolve EMR cluster performance issues including memory management, Spark job optimization, partitioning strategies, and resource allocation. Optimize Spark/PySpark workloads for cost and performance. Implement monitoring and performance tuning strategies for data processing pipelines. Data Analysis & Integration Analyze complex datasets across multiple systems to support data mapping, transformation, and integration. Define and implement data quality checks and validation frameworks. Collaborate with data architects and analysts to develop efficient data models and data flows. Leadership & Collaboration Act as a technical lead for data engineering initiatives and mentor junior engineers. Work closely with business stakeholders, product owners, and data consumers to gather requirements and translate them into technical solutions. Provide guidance on data architecture best practices and standards. Workflow & Automation Build and maintain workflow orchestration solutions using tools such as Airflow, Step Functions, or Glue Workflows. Automate deployment and management of data pipelines using CI/CD practices and infrastructure-as-code. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)