Senior Data Engineer

Daybreak AI, Inc.
San Francisco, CA, United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Cloud Computing Databases Information Engineering Software Debugging DevOps Distributed Data Store Distributed Systems Hadoop Distributed File System
+14 more
Apache Hive Python (Programming Language) PostgreSQL Machine Learning Operational Databases Sql Optimization Apache Spark Containerization Kubernetes Information Technology Deployment Automation Machine Learning Operations Data Pipelines Docker

Job description

As a Senior Data Engineer on our India engineering team, you will own the design and delivery of scalable, reliable data pipelines that process real-world supply chain data at scale. You will work closely with Data Science, DevOps, and Customer Success teams to deploy and evolve our core platform - and your hands-on experience with production data will directly shape product decisions.

This is a high-ownership role. You will be expected to solve ambiguous problems, mentor junior engineers, and contribute to the technical direction of the team.

What You’ll Do

  • Design, build, and maintain production-grade data pipelines handling large-scale supply chain datasets
  • Own end-to-end data modeling, warehousing architecture, and pipeline reliability across multiple customer environments
  • Collaborate with Data Science, DevOps, Infrastructure, and Customer Success teams to deliver and iterate on product deployments
  • Debug complex pipeline failures and performance bottlenecks across distributed systems
  • Drive improvements to engineering practices, tooling, and deployment automation
  • Mentor and support junior data engineers, fostering technical growth within the team
  • Contribute to Daybreak’s culture of continuous learning and operational excellence

Requirements

  • 5+ years of experience in data engineering, with demonstrated ownership of production pipelines
  • Strong proficiency in Python for data engineering workloads
  • Advanced SQL skills across multiple databases (PostgreSQL and others)
  • Hands-on experience with pipeline orchestration tools such as Apache Airflow or Dagster
  • Experience working with distributed data systems - Spark, Hive, or HDFS
  • Solid understanding of containers and Docker in a development/deployment context
  • Strong debugging skills and systematic approach to resolving data quality issues
  • Bachelor’s or advanced degree in Computer Science, Engineering, or a related field
  • Ability to thrive in a fast-paced, evolving environment and pick up new technologies quickly

Helpful to Have

  • Cloud experience, preferably AWS (S3, EMR, or equivalent services)
  • Experience processing large-scale time series data
  • Familiarity with Kubernetes for containerized workloads
  • Understanding of ML model lifecycle and MLOps pipelines
  • Experience working in interdisciplinary teams across engineering and data science

About the company

Daybreak AI is building next-generation supply chain intelligence for the world’s most complex manufacturing and logistics operations. Our Data Engineers are the backbone of that mission - designing, building, and operating the data infrastructure that powers our AI products., Daybreak AI is an AI-first supply chain technology company backed by TPG Growth and Dell Technologies Capital. Our mission is to eliminate waste in the world’s most complex supply chains through intelligent, enterprise-grade AI software. We are a diverse, intellectually curious team that values collaboration, first-principles thinking, and real-world impact.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all