> Markdown version of [/jobs/ext/3464023-data-engineer-foundations](https://www.wearedevelopers.com/jobs/ext/3464023-data-engineer-foundations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer, Foundations - **Company:** Superhuman LLC - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $157,000.0 - $220,500.0 - **Contract:** Permanent contract - **Skills:** Sql Data Warehouse, Airflow, Batch Processing, Big Data, BigQuery, Cloud Engineering, Continuous Integration, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Systems, Data Warehousing, Software Design Patterns, Distributed Systems, Python (Programming Language), Machine Learning, Standard Sql, Data Streaming, Workflow Management Systems, Data Processing, Data Storage Technologies, Snowflake, Apache Spark, Git, Data Lakes, Apache Flink, Apache Kafka, Spark Streaming, Terraform, Data Pipelines, Databricks - **Published:** September 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=db4b2ff29c4a064c ## About the Role * You have 3+ years of experience running live production environments, including high-load or data-intensive workflows, with a focus on uptime and reliability. * You're proficient in SQL and Python, with deep hands-on experience in Spark and a modern lakehouse or cloud data warehouse (Databricks, Delta Lake, dbt, Snowflake, or similar). * You have strong knowledge of ETL/ELT design patterns, orchestration tools (e.g., Airflow, dbt, or Dagster), data quality frameworks, and CI/CD for data with Git-based deployments. * You have data modeling and data warehouse design skills, along with a rigorous approach to data quality and observability. * You have hands-on experience with modern data storage technologies (for example, Delta Lake, Snowflake, BigQuery, or Redshift). * You communicate clearly and collaborate well across diverse teams and stakeholders, turning business needs into robust data solutions and trustworthy metrics. * You bring a strong ownership mindset, taking end-to-end responsibility for the systems you build, from strategy and architecture through production and ongoing reliability. * You have a strong acumen for scalable, highly complex data processing, think holistically about problems, and raise the bar on your team's craft. * You're a self-starting problem-solver who thinks from first principles, manages priorities across multiple projects, and thrives in a fast-paced, results-driven environment., * You have experience operating large-scale distributed systems, as well as in data engineering and data modeling. * Experience with high-throughput, real-time streaming systems (for example, Kafka, Flink, or Spark Structured Streaming) at the scale of billions of events per day. * Comfort with data lake and lakehouse technologies (for example, Delta Lake, Iceberg, or Hudi) and managing cloud infrastructure as code (for example, Terraform). * A track record of building self-serve data products, tools, or frameworks that other teams rely on. * Experience partnering with analytics, data science, or machine learning teams as a strategic data partner to productionize data and models. ## Description As a Data Engineer on the Data Foundations team, you will design and implement robust, scalable, and reliable data pipelines and systems that handle large volumes of data, empowering both product features and data-driven decision-making across the company. Your work will span areas such as ETL data pipelines, data lakes, performance-oriented data processing, and ETL framework development. * Architect, build, and own large-scale data pipelines and data lakes (Spark/Databricks) that ingest and process billions of daily events into reliable, decision-grade datasets. * Design and implement solutions that keep data available, secure, and scalable across the platform, enabling both real-time and batch processing. * Own data quality, freshness, and reliability for the foundational datasets other teams depend on, with automated checks, monitoring, and alerting. * Build the ETL frameworks and tooling that let other Data teams and data scientists self-serve and model core business entities into clean, well-documented, reusable datasets. * Partner with product teams, back-end engineers, ML engineers, and Data Science to turn business questions into high-impact data products behind business-critical features, research, and experimentation. * Collaborate with leadership to shape the team's charter and technical roadmap.