> Markdown version of [/jobs/ext/3541844-data-engineer](https://www.wearedevelopers.com/jobs/ext/3541844-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Cricut, Inc. - **Location:** South Jordan, UT, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, A/B Testing, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, Code Review, Continuous Integration, Data as a Services, Data Architecture, Information Engineering, Data Sharing, Software Design Documents, Amazon DynamoDB, Identity and Access Management, Python (Programming Language), Machine Learning, Recommender Systems, Power BI, Standard Sql, Data Streaming, Management of Software Versions, Usage Analysis, Feature Store, Feature Engineering, Large Language Models, Apache Spark, Data Lakes, Pyspark, Templating, Apache Kafka, Machine Learning Operations, Functional Programming, Amazon Simple Queue Service (SQS), Data Pipelines, Amazon Redshift - **Published:** September 30, 2026 - **Apply:** https://www.dice.com/job-detail/471960d1-91e5-48a6-87fe-868683420500 ## About the Role * 8+ years of experience in data engineering, including building and owning production pipelines at scale * Strong SQL and Python skills; experience with PySpark or Spark is a plus * Deep hands-on experience with AWS data services (S3, Glue, Redshift, DynamoDB, Lambda, IAM) * Production experience with Apache Airflow, including DAG design, dependency management, templating, alerting, and backfills * Experience with streaming and event ingestion (Kafka, Kinesis, SQS, or similar) and clickstream or product analytics data * Strong data modeling skills (dimensional and event modeling) and a clear sense of how data design affects downstream metrics * A track record of building data quality and observability frameworks, not just pipelines * Experience supporting ML systems in production: feature engineering, training datasets, feature stores, or model-serving data flows * Ability to lead cross-functional technical work, write clear design documents, and turn ambiguous business needs into sound architecture * Clear communication with engineers, data scientists, and business partners alike, * Experience with recommendation systems or personalization data (interaction logs, embeddings, candidate generation, ranking features) * Familiarity with SageMaker, AWS Batch, or MLOps tooling (model registries, experiment tracking, pipeline orchestration) * Experience with analytics instrumentation tooling or tracking-plan governance * Exposure to LLM or generative AI applications, such as vector stores, retrieval pipelines, or evaluation and feedback data * Experience with A/B testing platforms and experiment metric pipelines * Experience with data lake table formats (Iceberg, Delta, Hudi) or Redshift data sharing * A passion for making, crafting, or creative tools ## Description We're hiring a Senior Data Engineer to shape the data foundation behind Cricut's product analytics, personalization, and AI initiatives. You'll own and evolve our event data platform, from app instrumentation through streaming and batch ingestion to the warehouse. You'll also build the pipelines and feature infrastructure that feed our recommendation systems and machine learning models. This role sits where data engineering meets ML. You'll work closely with product engineering, data science, ML engineering, analytics, and experimentation teams. Together you'll make sure our data is reliable, well-modeled, and ready for both decision-making and production models. What You'll Do * Design, build, and operate scalable batch and streaming pipelines on AWS using Airflow (MWAA), Glue, Kafka, S3, and Redshift * Lead the evolution of our product event platform, including schema design, event taxonomy, versioning, and migrations to next-generation event architecture * Build event data quality and observability: schema validation, instrumentation testing, anomaly detection, freshness and completeness monitoring, and lineage across pipelines * Handle late-arriving and out-of-order data correctly, using lookback reprocessing and idempotent, incremental loads that keep business metrics accurate * Build and maintain the data foundations for personalization and recommendation systems, including the interaction, content, and project datasets used for model training and inference * Partner with ML engineers to build feature pipelines and a batch plus low-latency feature store (for example, Redshift or S3 to DynamoDB) for real-time serving * Support ML workflows on AWS Batch and SageMaker, including training data generation, offline evaluation datasets, and delivery of model outputs to downstream APIs * Instrument and model data for new AI-powered product experiences, including interaction, generation, and feedback events that close the loop on model improvement * Develop well-designed fact and dimension models that power BI, ad-hoc exploration, and experimentation platforms * Tune warehouse performance and cost through distribution and sort strategies, workload management, and efficient unload and serving patterns * Set engineering standards for code review, testing, CI/CD, documentation, and on-call practices, and mentor other engineers ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Beyond Dashboards: Fixing Text-to-SQL with Semantic RAG](https://www.wearedevelopers.com/videos/2036-beyond-dashboards-fixing-text-to-sql-with-semantic-rag) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)