Principal Data Engineer

Cg Infinity, Inc.
Houston, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Business Analytics Applications Automation of Tests Microsoft Azure BigQuery C Sharp (Programming Language) Cloud Computing Cloud Computing Security
+69 more
Software Quality Information Systems Continuous Integration Data Architecture Data Validation Information Engineering Data Governance Data Infrastructure Data Integration Extract Transform Load (ETL) Data Transformation Data Systems Data Vault Modeling Data Warehousing DevOps Dimensional Modeling Distributed Data Store Apache Hadoop Identity and Access Management Python (Programming Language) Key Management Machine Learning Metadata Meta-Data Management Operational Data Store Power BI DataOps Salesforce.Com SAP (Applications) Scala (Programming Language) SQL Databases Data Streaming Tableau (Software) Pulumi Google Cloud Data Ingestion Azure Data Factory Retrieval-Augmented Generation Delivery Pipeline Snowflake Apache Spark Generative AI Data Strategy Git Cloudformation Data Layers Microsoft Fabric Containerization Data Lakes Kubernetes Infrastructure Automation Frameworks Information Technology Data Lineage Collibra Apache Flink Deployment Automation AWS Glue Bicep Apache Kafka Data Management Terraform Azure Synapse Analytics Looker Analytics Software Version Control Data Pipelines Docker Amazon Redshift Databricks Programming Languages

Job description

We are seeking a Principal Data Engineer to lead the architecture, design, and delivery of modern data platforms and analytics solutions for clients. This is a hands-on technical leadership role for someone who can move comfortably between executive-level client conversations, solution architecture, engineering delivery, and mentoring high-performing data teams. The ideal candidate combines deep data-engineering expertise with strong consulting instincts: they can translate business objectives into scalable technical solutions, communicate complex concepts through clear presentations, and guide teams from discovery through implementation and operationalization., * Lead the end-to-end architecture, design, and delivery of enterprise data engineering, data platform, and analytics solutions for client engagements.

  • Serve as a trusted technical advisor to client stakeholders, including technology leaders, data leaders, architects, and business partners.
  • Facilitate discovery sessions, requirements workshops, technical assessments, architecture reviews, and solution-design discussions.
  • Translate business goals, data challenges, and operating-model requirements into practical data strategies, roadmaps, reference architectures, and implementation plans.
  • Define scalable, secure, reliable, and cost-effective architectures for data ingestion, transformation, storage, governance, orchestration, analytics, and data consumption.
  • Remain hands-on in engineering work, including designing data pipelines, reviewing code, building proof of concepts, resolving complex technical issues, and establishing engineering patterns.
  • Architect and implement batch, real-time, streaming, and event-driven data solutions as appropriate for client needs.
  • Design modern cloud data platforms using technologies such as Snowflake, Databricks, Microsoft Fabric, Azure Data Factory, AWS Glue, Amazon Redshift, BigQuery, or equivalent platforms.
  • Lead the development of robust ETL/ELT pipelines, data models, data APIs, semantic layers, and data products.
  • Establish data engineering standards for code quality, testing, CI/CD, observability, metadata management, data lineage, documentation, security, and production support.
  • Drive adoption of DataOps practices, infrastructure as code, automated testing, deployment pipelines, monitoring, and incident-management processes.
  • Partner with data architects, data scientists, analysts, application architects, security teams, and business stakeholders to ensure solutions are aligned to enterprise architecture and business outcomes.
  • Present technical recommendations, architecture options, delivery status, risks, and strategic roadmaps to both technical and executive audiences.
  • Lead client-facing demonstrations, workshops, steering-committee updates, and technical presentations.
  • Mentor senior, mid-level, and junior data engineers; provide technical coaching, career guidance, design feedback, and hands-on support.
  • Lead and influence cross-functional delivery teams, including onshore/offshore engineers, architects, analysts, and client technical resources.
  • Participate in project estimation, staffing, delivery planning, risk management, solution scoping, and proposal development.
  • Support pre-sales and business-development activities, including solutioning, client presentations, technical discovery, RFP responses, estimates, and statement-of-work development.
  • Stay current on emerging data, cloud, AI, governance, and analytics technologies; evaluate where they provide meaningful business value for clients.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field; equivalent professional experience may be considered.
  • 10+ years of progressive experience in data engineering, data warehousing, data integration, data architecture, or related technical disciplines.
  • 3+ years of experience in a technical leadership, lead engineer, solution architect, staff engineer, principal engineer, or consulting leadership capacity.
  • Demonstrated experience architecting and delivering enterprise-scale data platforms and data integration solutions.
  • Strong hands-on expertise in SQL and at least one modern programming language, preferably Python, Scala, Java, or C#.
  • Strong experience with ETL/ELT design, data pipeline development, data transformation frameworks, orchestration, and workflow automation.
  • Experience with one or more cloud providers: Microsoft Azure, AWS, or Google Cloud Platform.
  • Experience with modern cloud data and analytics platforms such as Databricks, Snowflake, Microsoft Fabric, Synapse Analytics, BigQuery, Redshift, or similar technologies.
  • Experience with orchestration and pipeline technologies such as Apache Airflow, Azure Data Factory, AWS Step Functions, dbt, Dagster, Prefect, or equivalent tools.
  • Knowledge of distributed data-processing technologies such as Apache Spark, Kafka, Flink, Hadoop ecosystems, or comparable platforms.
  • Experience designing both batch and near-real-time or streaming data solutions.
  • Strong understanding of dimensional modeling, data vault, normalized data models, lakehouse architectures, data lake architectures, and data warehouse design principles.
  • Experience implementing data quality, data validation, monitoring, lineage, metadata, governance, and security controls.
  • Working knowledge of DevOps and DataOps practices, including Git-based source control, CI/CD, automated testing, deployment automation, and infrastructure as code.
  • Strong understanding of cloud security principles, identity and access management, encryption, secrets management, and role-based access controls.
  • Proven ability to lead architecture discussions and make well-reasoned technical tradeoffs involving performance, scalability, reliability, maintainability, security, and cost.
  • Excellent verbal, written, and presentation skills, with the ability to explain technical concepts clearly to business stakeholders and executive audiences.
  • Experience in a consulting, professional services, systems integrator, or client-facing delivery environment., * Experience designing data platforms that support AI, machine learning, generative AI, retrieval-augmented generation, feature stores, vector databases, or advanced analytics workloads.
  • Certifications in AWS, Azure, Google Cloud, Databricks, Snowflake, Microsoft Fabric, or other relevant data technologies.
  • Experience with master data management, data cataloging, data governance, privacy, regulatory compliance, or data stewardship programs.
  • Experience with tools such as Collibra, Alation, Microsoft Purview, Unity Catalog, Informatica, Monte Carlo, Great Expectations, or similar governance and observability platforms.
  • Experience with Salesforce, SAP, ERP, CRM, finance, supply-chain, healthcare, retail, manufacturing, or other enterprise operational data domains.
  • Familiarity with BI and semantic-layer technologies such as Power BI, Tableau, Looker, ThoughtSpot, or similar platforms.
  • Experience with containerization and cloud-native technologies such as Docker, Kubernetes, Terraform, CloudFormation, Bicep, or Pulumi.
  • Prior experience contributing to proposals, estimates, statements of work, or technical sales pursuits.
  • Experience managing or leading distributed onshore/offshore delivery teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

Videos

See all

Related articles

See all