Senior Data Engineer

Procter & Gamble
Cincinnati, OH, United States
15 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$110,000.0 - $165,300.0
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Application Programming Interfaces (APIs) Artificial Intelligence Application Integration Architecture Automation of Tests Microsoft Azure Big Data BigQuery Cloud Database Cloud Storage Software Documentation Code Review
+58 more
Continuous Integration Data as a Services Data Architecture Data Transmissions Information Engineering Data Governance Data Infrastructure Data Integration Extract Transform Load (ETL) Data Systems Data Vault Modeling Data Warehousing Software Design Patterns DevOps Programming Tools Distributed Computing Environment Data Flow Control Graph Database Python (Programming Language) Machine Learning MongoDB Neo4j NoSQL Performance Tuning Queueing Systems Cloud Services DataOps Software Engineering SQL Databases Data Streaming Azure Service Bus Data Processing Google Cloud Real Time Systems Azure Data Factory Apache Spark Generative AI Data Strategy Git Data Lakes Pyspark Information Technology Data Lineage Apache Flink Dask Apache Kafka Cosmos DB Spark Streaming Data Management Machine Learning Operations Terraform Stream Processing Azure Synapse Analytics Stream Analytics Data Pipelines Api Management Databricks Microservices

Job description

As a Senior Data Engineer, you will be a technical leader and expert, responsible for architecting, designing, and implementing highly scalable and robust cloud-based data and analytics platforms (DAP) and complex data pipelines. You will drive the strategy for acquiring, cleansing, transforming, and publishing critical data assets from diverse enterprise and external sources. Beyond building cutting-edge data solutions, you will act as a principal liaison, partnering deeply with senior business stakeholders, solution architects, and analytics leaders to define technical roadmaps, significantly influence data architecture, and establish enterprise-wide engineering standards and best practices. We seek individuals who are not only masters of current technologies but are also visionary, continuously exploring and integrating emerging data engineering paradigms and tools to push the boundaries of what’s possible., * Strategic Technical Leadership:

  • Lead the architectural design and implementation of complex, large-scale data solutions, ensuring scalability, performance, security, and cost-efficiency.
  • Partner with senior business stakeholders and product owners to deeply understand strategic business objectives and translate them into architectural blueprints and technical roadmaps for data platforms.
  • Influence and drive the overall data strategy, architecture, and technology choices across multiple teams or domains.
  • Advanced Data Platform & Pipeline Development:
  • Architect, build, and optimize highly resilient, performant, and secure ETL/ELT pipelines on modern cloud data platforms, handling petabyte-scale data volumes and real-time processing requirements.
  • Design and implement advanced data integration patterns, connecting complex enterprise systems, third-party services, streaming sources, and APIs, ensuring high data availability and reliability.
  • Drive the adoption of advanced data processing techniques (e.g., stream processing, graph databases, data mesh principles).
  • Mentorship & Community Building:
  • Serve as a primary technical mentor and subject matter expert for a team of data engineers, providing guidance on complex technical challenges, architectural decisions, and career development.
  • Lead code reviews, design discussions, and technical workshops, fostering a culture of excellence and continuous improvement.
  • Champion and evolve our enterprise-wide engineering standards, best practices, and governance for data (e.g., data quality frameworks, testing automation, CI/CD pipelines, security protocols, documentation standards, data observability).
  • End-to-End Ownership & Operational Excellence:
  • Take ultimate end-to-end ownership for critical data solutions, from strategic inception and architectural design through implementation, deployment, advanced monitoring, performance tuning, and incident response for production systems.
  • Implement robust data quality frameworks, observability solutions, and anomaly detection to ensure the highest integrity and reliability of data assets.
  • Innovation & AI Integration:
  • Proactively evaluate, prototype, and integrate cutting-edge technologies, including advanced Generative AI models and sophisticated agentic systems, to dramatically enhance developer productivity, automate complex tasks, and create novel data solutions.
  • Act as a thought leader in the responsible and ethical application of AI in data engineering, ensuring best practices for security, privacy, and bias mitigation.
  • Lead initiatives for continuous learning and knowledge sharing across the broader engineering organization.
  • Modern Development Practices:
  • Master modern development tools and practices, including advanced IDE features, sophisticated Git strategies (e.g., monorepos, gitflow), infrastructure as code (IaC), and advanced CI/CD pipelines tailored for data platforms.

Requirements

  • Education: Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a closely related quantitative field.
  • Experience: 5+ years of progressive experience in data engineering, with a significant track record of designing and delivering large-scale, complex data platforms and pipelines.
  • Technical Leadership: Proven experience leading technical projects, mentoring senior and junior engineers, and influencing architectural decisions across multiple teams.
  • Advanced Python & SQL: Expert-level proficiency in Python and SQL for complex data manipulation, optimization, performance tuning, and advanced analytics.
  • Deep Cloud Expertise: Expert-level understanding and hands-on experience with at least one major modern cloud platform (Azure preferred, and/or GCP), including deep knowledge of their data services (e.g., Azure Synapse, Databricks, Data Factory, Event Hubs, Data Lake Storage; or GCP BigQuery, Dataflow, Pub/Sub, Cloud Storage).
  • Distributed Processing Mastery: Extensive hands-on experience and deep understanding of distributed data processing technologies (e.g., Spark, PySpark, Dask), including performance optimization, cluster management, and resource allocation for petabyte-scale data.
  • Data Modeling & Architecture: Expert-level knowledge of advanced data modeling techniques (dimensional, Kimball, Inmon, data vault, data mesh concepts), data warehousing principles, and data lake architectures. Ability to design highly optimized and flexible data schemas.
  • API & Integration Expertise: Proven ability to architect and implement complex data integrations with a wide array of systems, including advanced API integrations, message queues (e.g., Kafka, Azure Event Hubs), and enterprise-grade data transfer protocols.
  • DevOps & MLOps for Data: Extensive experience with modern development tools, CI/CD pipelines, infrastructure as code (Terraform, ARM templates), and best practices for deploying, monitoring, and managing data and machine learning pipelines in production.
  • AI Integration & Responsible AI: Demonstrated practical experience and leadership in leveraging Generative AI tools and agentic systems to accelerate development and solve complex data problems. Deep understanding of responsible AI principles, including data privacy, security, and ethical considerations.
  • Communication & Influence: Exceptional communication, presentation, and interpersonal skills, with the ability to articulate complex technical concepts to both technical and non-technical senior stakeholders and influence strategic decisions.
  • Strategic Ownership: Demonstrated ability to drive initiatives from conception to completion, taking full architectural and operational responsibility for critical data assets.
  • Continuous Innovation: A profound curiosity and passion for continuous learning, staying abreast of industry trends, and proactively evaluating and adopting emerging data technologies.

Preferred:

  • Azure Specialization: Deep expertise and certifications in Azure data services (e.g., Azure Databricks, Azure Synapse Analytics, Azure Data Factory, Azure Stream Analytics).
  • Advanced Data Governance: Experience implementing robust data governance, master data management (MDM), and data lineage solutions.
  • Real-time Processing: Hands-on experience with real-time data streaming and processing frameworks (e.g., Kafka, Spark Streaming, Flink).
  • Software Engineering Background: Strong software engineering fundamentals (design patterns, clean code principles, microservices architecture) applied to data platforms.
  • NoSQL/Graph Databases: Experience with NoSQL databases (e.g., Cosmos DB, MongoDB) or graph databases (e.g., Neo4j) for specialized data use cases.
  • Advanced Certifications: Professional or Expert-level certifications (e.g., Azure Data Engineer Expert, Databricks Certified Data Engineer Professional, Google Cloud Professional Data Engineer).
  • Machine Learning/MLOps: Experience collaborating with or supporting MLOps initiatives and integrating data pipelines with ML models.

Compensation for roles at P&G varies depending on a wide array of non-discriminatory factors including but not limited to the specific office location, role, degree/credentials, relevant skill set, and level of relevant experience. At P&G compensation decisions are dependent on the facts and circumstances of each case. Total rewards at P&G include salary + bonus (if applicable) + benefits. Your recruiter may be able to share more about our total rewards offerings and the specific salary range for the relevant location(s) during the hiring process.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:19 min

Scaling performance across multiple GPUs using specialized frameworks

Paul Graham Paul Graham · World Congress 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

1:31 min

Selecting the ideal data stack for distinct workloads

Alan Mazankiewicz · LIVE

Videos

See all

Related articles

See all