Data Engineer - Senior/Lead

Salesforce.com, Inc.
San Francisco, CA, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$172,500.0 - $260,100.0
Working hours
Regular working hours

Tech stack

A/B Testing Application Programming Interfaces (APIs) Airflow Amazon Web Services Data Analysis Big Data Databases Continuous Integration Customer Data Management Information Engineering Data Governance Web Scraping
+30 more
Data Infrastructure Extract Transform Load (ETL) Data Profiling Data Structures Data Warehousing Dimensional Modeling Distributed Data Store Github Apache Hadoop Python (Programming Language) Operational Databases Performance Tuning Software Reliability Testing Standard Sql Salesforce.Com Shell Script SQL Databases Data Streaming Subversion Tableau (Software) Scripting Snowflake Apache Spark Data Layers Data Lakes Deployment Automation Machine Learning Operations Software Version Control Data Pipelines Databricks

Job description

We are looking for a Data Engineer who thrives at the intersection of data engineering and analytics. In this role you will partner closely with data analysts and strategy experts to turn raw, distributed data into trusted, well-modeled datasets that power product strategy. Our teams are made up of data scientists, engineers, and strategy lead who drive product strategy with data-driven insights. We work alongside executives, product managers, customer strategy, and sales strategy partners to discover new opportunities for growth, experiment with data, drive adoption, and surface insights that shape what we build next.

Responsibilities:

  • Sit alongside data scientists and strategy leads in planning, design reviews, and roadmap discussions - treating their questions and hypotheses as first-class inputs to architecture decisions.
  • Translate analytical and statistical requirements into well-performing SQL and scalable pipelines, and coach partners on patterns that scale (windowing, partitioning, incremental loads, idempotency).
  • Own the technical solution design and architecture of data acquisition and integration projects (batch and real-time), implementing a layered stack - raw cleansed curated semantic - that ensures high data quality, predictable freshness, and timely insights.
  • Craft design artifacts (functional design documents, data flow diagrams, data models, schema contracts) that the broader team can review, extend, and rely on.
  • Build the data pipelines, curated marts, semantic layers, and feature stores that let analysts answer business questions independently and let data scientists iterate on features and models without re-engineering raw sources.
  • Design tailored data structures (fact/dimension models, wide analytical tables) and end-to-end infrastructure for data science work: feature pipelines, model-ready training datasets, experimentation data, and the plumbing required for reliable ML and statistical workflows.
  • Take exploratory analyst/DS notebooks and prototypes and reinvent them as production-ready, monitored data flows.
  • Co-own data quality, lineage, and trust with your partners, and establish shared conventions - naming, documentation, testing, and review - that make handoffs low-friction.
  • Proactively identify gaps in data quality and performance, integrate data from disparate sources, and advocate for architectural and code improvements that improve execution speed and reliability.
  • Perform data profiling, sophisticated sampling, statistical testing, and reliability testing on data.
  • Serve as a domain expert and mentor for ETL/ELT design, dimensional modeling, and big data patterns; evaluate technology trade-offs and run proofs of concept to inform tooling decisions.
  • Bring strong SQL optimization and performance tuning expertise in high-volume, parallel-processing environments, working with the team’s stack: SQL, Python, Airflow, AWS, Spark, Tableau, Hadoop (and adjacent tools like dbt, Snowflake, and Databricks where they fit).
  • Participate in the team’s on-call rotation to address production data issues in real time and keep services operational and highly available for analytics and ML consumers.

Requirements

  • 8+ years of experience in data engineering.
  • Demonstrated experience working closely with data scientists and strategy analysts - not just shipping pipelines, but understanding how the data will be modeled, sampled, and analyzed downstream.
  • Build programmatic ETL/ELT pipelines with SQL-based technologies and platforms.
  • Solid understanding of databases and working with sophisticated datasets in Salesforce environment.
  • Data governance, verification, and data documentation using current and emerging tools and platforms.
  • Comfort across multiple technologies (Python, shell scripts) and the ability to translate business and analytical logic into well-performing SQL.
  • Comfort with tasks such as writing scripts, web scraping, and pulling data from APIs.
  • Experience automating data pipelines using scheduling and orchestration tools like Airflow.
  • Ability to adapt to changes in business direction and recognize when designs need to evolve.
  • Experience writing production-level SQL and a strong grasp of data engineering pipelines end-to-end.
  • Experience with the Hadoop ecosystem and similar frameworks.
  • Previous projects that show technical leadership across data lake, data warehouse, business intelligence, big data analytics, and enterprise-scale custom data products - ideally where analysts and data scientists were primary consumers.
  • Strong knowledge of data modeling techniques and high-volume ETL/ELT design.
  • Experience with version control systems (GitHub, Subversion) and deployment tools (e.g., CI/CD).
  • Ability to work effectively in an unstructured, fast-paced environment, both independently and as part of a cross-functional team, with a high degree of self-management, clear communication, and commitment to delivery timelines.
  • A technical degree is required.

Preferred Qualification:

  • Experience building feature stores, ML pipelines, or model-serving infrastructure in partnership with data scientists.
  • Experience designing semantic layers or metrics layers (e.g., dbt metrics, LookML, Cube) that empower analyst self-service.
  • Experience with experimentation platforms or A/B testing data infrastructure.

Benefits & conditions

In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: https://www.salesforcebenefits.com.Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records. At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions. The typical base salary range for this position is, $172,500 -

About the company

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword - it’s a way of life. The world of work as we know it is changing and we’re looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce’s core values at the heart of it all.

Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all