> Markdown version of [/jobs/ext/1165225-data-engineer](https://www.wearedevelopers.com/jobs/ext/1165225-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Amazon.com, Inc. - **Location:** Atlanta, GA, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Microsoft Azure, Common ISDN Application Programming Interface (CAPI), Customer Data Management, Data Validation, Data Deduplication, Information Engineering, Data Governance, Extract Transform Load (ETL), Middleware, Performance Tuning, SQL Databases, Systems Integration, Web Platforms, Data Ingestion, Azure Data Factory, Snowflake, Apache Spark, Pyspark, Data Pipelines, Databricks - **Published:** July 3, 2026 - **Apply:** https://www.careerjet.com/jobad/uscd69b9e95df5420a441f6c59705ad929 ## About the Role Candidate with hands-on, recent experience in: Strong coding: PySpark + SQL (hands-on, not only orchestration) Databricks: notebooks/jobs, performance tuning fundamentals, medallion patterns Spark fundamentals: partitioning, skew/shuffle optimization, understanding failures via logs Snowflake: data modeling/usage for analytics/warehousing workloads Azure ecosystem: Azure Data Factory (ADF) (orchestration) Azure-native integrations and services exposure Data engineering reliability patterns: validation, idempotency, replay/backfills, dedup, auditability Data governance: Unity Catalog (preferred), lineage, access control patterns, PII handling Ownership mindset: can execute independently without constant approvals/check-ins Nice-to-Have Skills Event-driven/streaming ingestion exposure (even if primary is batch today) Delta/Databricks patterns such as Delta Live Tables (DLT) (some workflows exist) Experience building config-driven export frameworks for multiple downstream consumers/vendors Exposure/interest in identity resolution concepts (ML optional; ETL strength is priority) Familiarity with CAPI integrations / marketing tech data signals Experience implementing operational telemetry: dashboards, alerts, SLA monitoring What Good Looks Like (Success Criteria) Ships reliable, well-governed datasets with strong data quality practices Can scale pipelines for very large volumes (hundreds of millions of records per vendor) Prevents silent failures where quality degrades without obvious job failures Balances delivery speed with compliance, governance, and cost controls ## Description Needed for Azure-native third party data enrichment platform using Databricks/Spark + Snowflake; focus on reliable governed pipelines, strong Spark troubleshooting, privacy/governance, and cost-aware engineering; h Team / Business Context: You will join a data engineering team responsible for third party data enrichment augmenting first party datasets with external identity/attribute data to support analytics, activation, and research. The enriched datasets are consumed by multiple downstream systems and teams, including the Customer Data Platform (CDP) and other analytics/research stakeholders. The platform is Azure-native and built primarily on Databricks (processing + some ML workloads) and Snowflake (analytics/warehouse). A major focus is building reliable, governed, vendor agnostic datasets while ensuring privacy/compliance, data governance, and cost efficiency. Key Responsibilities As a Data Engineer, you will: Data Ingestion & Pipeline Development Build and enhance ingestion pipelines for large batch and event-driven paths (streaming may evolve over time). Integrate data from: Third party enrichment vendors (identity + attributes, very large volumes) Digital platforms via Conversion API (CAPI) integrations (through intermediary/middleware) Rewards/Promotions systems (e.g., TMT) for offer issuance/redemption/consumption data Data Quality, Reliability & Operations Implement strong data validation, idempotency, replay/backfill strategies, and deduplication to prevent quality drift. Own monitoring, alerting, dashboarding, and operational readiness ( wrappers around core pipelines). Troubleshoot failures with root cause analysis not just reruns: Interpret Spark logs Diagnose performance issues (shuffle, skew, partitioning) Improve stability and SLA adherence Governance & Compliance (First-class NFR) Apply privacy, compliance, and governance requirements across pipelines and datasets. Support governance standards such as: Unity Catalog, lineage, access controls Managing PII vs non PII access Documentation of tables, schemas, catalogs, and cluster usage Cost Governance & Performance Optimization Design pipelines with cost awareness from day one: Cluster sizing, workload tuning, efficient compute/storage usage Trade-off decisions balancing cost vs quality vs SLA Collaboration & Ownership Work in a small, fast-moving team; be self-driven and ownership-oriented. Raise and manage data quality escalations when issues are detected. Contribute to evolving architecture (product is early-stage; first live month was recent). Must-Have Skills (Screening Keywords) ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)