> Markdown version of [/jobs/ext/997239-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/997239-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** Toyota Motor North America - **Location:** Plano, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, Apache HTTP Server, Application Services, Code Review, Information Systems, Databases, Directed Acyclic Graph (Directed Graphs), Data Architecture, Data Validation, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Masking, Data Security, Data Sharing, Data Systems, Data Warehousing, Software Debugging, Amazon DynamoDB, Github, Apache Hive, Identity and Access Management, Python (Programming Language), Meta-Data Management, Team Foundation Server, Query Optimization, Application Data, Simple Data Format, SQL Databases, Data Streaming, Strategies of Testing, Workflow Management Systems, Parquet, Datadog, AWS Cdk, Amazon ElastiCache, Delivery Pipeline, Apache Spark, Change Data Capture, Backend, Cloudformation, Data Lakes, Integration Tests, Debezium, Information Technology, Data Lineage, Apache Flink, Avro, AWS Glue, Data Analytics, AWS Data Analytics, Apache Kafka, Spark Streaming, Presto, Event Sourcing, Cloudwatch, Amazon Simple Queue Service (SQS), Terraform, Stream Processing, Data Pipelines, Serverless Computing, Amazon Elastic Mapreduce (EMR), Amazon Redshift - **Published:** June 5, 2026 - **Apply:** https://www.dice.com/job-detail/6eac7ec3-e5b5-466b-897e-d3c1aad937e6 ## About the Role * Bachelor's degree in Computer Science, Data Engineering, Information Systems, or related field, or equivalent practical experience * 7+ years of software or data engineering experience, including 3-5 years focused specifically on data platform and pipeline engineering at scale, with a track record of operating at a principal or staff engineer level * Deep expertise in designing and building data lake and Lakehouse architectures on AWS, including: + S3 as the foundation for data lake storage, with strong opinions on partitioning, file formats (Parquet, Avro, ORC), and lifecycle management + AWS Glue for ETL/ELT jobs, crawlers, and the Data Catalog + Amazon Athena for serverless SQL analytics over the data lake + Lake Formation for fine-grained access control, governance, and cross-account data sharing + Amazon Redshift or Redshift Serverless for data warehousing and high-performance analytical queries + Amazon EMR or EMR Serverless for large-scale Spark, Hive, or Presto workloads * Production experience with real-time and streaming data architectures, including: + Amazon Kinesis (Data Streams, Data Firehose) for real-time ingestion and delivery + Amazon MSK (Managed Kafka) or self-managed Kafka for event streaming at scale + EventBridge, SQS, or SNS for event-driven integration with application services + Lambda for lightweight stream processing and event transformation + Apache Flink (via Amazon Managed Service for Apache Flink) or Spark Structured Streaming for stateful stream processing * Strong proficiency in Python and SQL - you write production-quality pipeline code, not just ad-hoc scripts, and you can optimize a complex query as fluently as you can design a DAG * Experience with workflow orchestration tools: Step Functions, Apache Airflow (via Amazon MWAA), or similar - you know how to build reliable, observable, and recoverable pipeline DAGs * Solid understanding of data modeling for both analytical and operational use cases: star schemas, slowly changing dimensions, wide tables, event sourcing, and CDC (change data capture) patterns * Experience with data quality and governance tooling and practices: Great Expectations, Deequ, or custom validation frameworks - plus data cataloging, lineage tracking, and access control * Strong understanding of Infrastructure as Code using AWS CDK, CloudFormation, or Terraform for data infrastructure * Experience with observability and monitoring for data systems: pipeline health dashboards, data freshness tracking, SLA monitoring, and alerting on failures or anomalies (CloudWatch, Datadog, or similar) * Strong understanding of security best practices for data: IAM policies, Lake Formation permissions, encryption at rest and in transit, data masking, and PII handling * Deep experience debugging complex issues across data systems - pipeline failures, data skew, schema mismatches, streaming lag, and storage cost runaway * Experience with testing strategies for data pipelines: data validation, schema contract testing, integration testing, and pipeline idempotency * Strong written and verbal communication - you can write a clear RFC, lead a design review, and explain a data architecture tradeoff to a non-technical stakeholder Added bonus if you have * Master's degree in Computer Science, Data Engineering, or related field * Experience in the financial services, banking, or insurance industry * Experience with open table formats: Apache Iceberg, Delta Lake, or Apache Hudi for ACID transactions, time travel, and schema evolution on the data lake * Experience with feature store design and implementation for ML/AI use cases (SageMaker Feature Store, Feast, or custom) * Familiarity with dbt or similar transformation frameworks for analytics engineering and data modeling * Experience with real-time analytics serving layers: Amazon OpenSearch, DynamoDB, or ElastiCache for low-latency data access * Experience designing multi-account AWS data architectures with proper governance and guardrails (AWS Organizations, Control Tower, cross-account data sharing via Lake Formation) * Hands-on experience with data mesh or data product patterns - decentralized ownership with centralized governance * Experience with CDC (change data capture) tools: AWS DMS, Debezium, or similar for streaming database changes into the data lake * Experience with cost optimization for data workloads: storage tiering, compute right-sizing, spot instances for Spark, and query optimization * Experience with GenAI data pipelines: preparing training datasets, building RAG knowledge bases, embedding generation, and vector store population * AWS certifications (Data Analytics Specialty, Solutions Architect, Database Specialty) * Experience with CI/CD pipelines for data infrastructure and pipeline deployment (CodePipeline, GitHub Actions, or similar) * Experience contributing to or maintaining open-source data engineering projects * Experience defining engineering standards, writing ADRs, or leading org-wide technical initiatives ## Description * Serve as the technical authority for data architecture across the organization, making high-impact decisions on data lake design, streaming topologies, storage formats, partitioning strategies, and data modeling patterns * Design, build, and maintain production-grade data pipelines - batch and real-time - from ingestion and transformation to serving and consumption * Own the data platform: build and evolve the foundational infrastructure that engineering, ML/AI, and analytics teams depend on for reliable, governed, and performant data access * Partner closely with ML/AI engineers to ensure training data, feature pipelines, and model serving data are accurate, fresh, and efficiently delivered - you are the upstream enabler for every model in production * Collaborate with backend and full-stack engineers to design event-driven architectures, define data contracts, and ensure application data flows cleanly into the data platform * Lead technical design reviews, architecture discussions, and RFC processes for data initiatives - driving alignment across engineering teams * Identify and resolve systemic data issues: pipeline failures, data quality degradation, schema drift, latency in streaming systems, cost inefficiencies in storage and computing, and gaps in data observability * Define and champion data engineering best practices: data modeling, schema evolution, data contracts, testing strategies, lineage tracking, cataloging, and governance * Design and implement data quality frameworks - validation rules, anomaly detection, freshness checks, and alerting - so downstream consumers can trust the data without asking * Collaborate closely with Engineering Managers, Product, Data Science, and Analytics to shape data roadmaps and ensure the platform evolves with business needs * Mentor and grow engineers at all levels through code reviews, pairing, design feedback, and technical guidance on data engineering topics * Contribute to hiring by conducting technical interviews and helping define what great looks like for data engineering at TFS * Proactively communicate technical risks, tradeoffs, and recommendations to both engineering and non-technical stakeholders ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)