Principal Engineer 1, Data Platforms

Halozyme, Inc.
San Diego, CA, United States
8 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$175,000.0 - $185,000.0
Working hours
Regular working hours

Tech stack

Query Performance Artificial Intelligence Audit Trail Microsoft Azure Clinical Trial Management Systems Code Review Cyber Security Information Systems Continuous Integration Information Engineering Data Governance Data Infrastructure
+27 more
Extract Transform Load (ETL) DevOps Human Resources Information System (HRIS) Github Python (Programming Language) Laboratory Information Management Systems Role-Based Access Control Power BI SQL Databases Transact-SQL Data Processing Data Classification Delivery Pipeline Apache Spark HR Software Build Management Microsoft Fabric Pyspark Git Flow Storage Technologies Information Technology Data Management Veeva Azure Synapse Analytics Data Pipelines GXP Databricks

Job description

The Principal Engineer 1, Data Platforms serves as the technical lead accountable for the data platform: architecting the solution, building alongside the team, holding the quality bar, and driving delivery with urgency and accountability. They work directly with business partners across the organization to understand their data needs, translate requirements into platform capabilities, and ensure the UDP delivers real, measurable value to every function it serves, and build the operational foundation that keeps the platform healthy as adoption scales: supporting models, SLA frameworks, and feedback loops that make the platform responsive and reliable. This role helps define and track technical KPIs like pipeline reliability, data freshness, query performance, onboarding velocity, incident resolution and use them to drive continuous improvement.

In this role, you’ll have the opportunity to:

  • Serve as the technical lead for Halozyme’s Microsoft Fabric-based data lakehouse - workspace topology, OneLake storage design, medallion-layer standards (bronze/silver/gold), Spark and SQL compute configuration, and CI/CD pipeline design
  • Provide technical leadership and mentorship to data and analytics engineers through code reviews, architectural guidance, and hands-on problem solving - personally contributing code and architectural artifacts alongside the team
  • Design and build production-grade data ingestion pipelines across 50+ source systems (ERP, Veeva, LIMS, ELN, CTMS, HRIS, and others) using Data Factory, dataflows, shortcuts, and mirroring patterns
  • Serve as the hands-on technical counterpart to external implementation partners - reviewing deliverables line by line, enforcing engineering standards, and catching quality issues before they reach production
  • Partner directly with business stakeholders across R&D, Commercial, Manufacturing, Business Development, Supply Chain and Corporate functions to understand data requirements, prioritize platform capabilities, and ensure the UDP delivers actionable value aligned to each function’s strategic needs
  • Design and operationalize the platform support model - including intake workflows, tiered support processes, issue triage, and escalation paths - to ensure the data platform remains responsive, reliable, and well-governed as enterprise adoption scales
  • Help define, instrument, and report on key technical KPIs - pipeline reliability, data freshness, query performance, onboarding velocity, and incident resolution time - to drive continuous improvement and demonstrate platform health to leadership
  • Implement and maintain data quality monitoring, alerting frameworks, incident response procedures, and platform SLAs - owning reliability as a commitment, not just a team metric
  • Build and maintain the semantic layer (Power BI datasets, gold-layer views, and self-service analytics capabilities) to ensure business consumers get trusted, performant data products
  • Execute data classification (GxP vs. non-GxP), role-based access control, sensitivity labeling, and audit logging in coordination with IT Security and the Data Governance Lead
  • Technically scope, sequence, and deliver new source system integrations from discovery through production, managing dependencies and communicating timelines with precision
  • Establish and enforce platform engineering standards: naming conventions, branching strategies, testing patterns, documentation requirements, and operational runbooks
  • Other duties as assigned

Requirements

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or a related technical field with 10+ years experience of progressive experience in data engineering, data platform architecture, or analytics infrastructure roles
  • Master’s degree preferred
  • An equivalent combination of experience and education may be considered
  • 2+ years technically leading data engineers or analytics engineers in a platform or lakehouse environment
  • Deep, hands-on expertise with Microsoft Fabric (lakehouses, notebooks, Data Factory, deployment pipelines); equivalent depth in Databricks or Azure Synapse considered with demonstrated willingness to go deep on Fabric
  • Strong production-level proficiency in PySpark, T-SQL, Python, and modern ELT/ETL design patterns
  • Proven experience designing and operating medallion-architecture (bronze/silver/gold) or equivalent layered data platforms at enterprise scale
  • Demonstrated ability to manage external implementation vendors at a technical level - reviewing code, challenging architecture decisions, and holding delivery quality, not just tracking timelines
  • Experience with DevOps and CI/CD for data platforms (Azure DevOps, GitHub Actions, Fabric deployment pipelines, infrastructure-as-code)
  • Familiarity with data governance tooling (Microsoft Purview, Unity Catalog, or comparable) and data quality frameworks
  • Biopharmaceutical, life sciences, or regulated industry experience strongly preferred - including familiarity with GxP data handling, validation requirements, and audit readiness
  • Experience integrating pharma-specific source systems such as Master Control, Veeva, LIMS, ELN such as Benchling, CTMS, EDC, or HRIS platforms is a significant plus
  • Exposure to AI/ML workloads on a Lakehouse (feature stores, model training pipelines, vector databases) is a plus
  • Microsoft Fabric Data Engineering Certificate is a plus

Benefits & conditions

In return, we offer you:

  • Full and comprehensive benefit program, including an Employee Stock Purchase Program and 401(k) matching.
  • Opportunities to grow in a culture that prioritizes learning, development and progression through in-house programs and tuition reimbursement.
  • A collaborative, innovative team that works as one to amplify your impact-on your career, the work you do and patients’ lives.

The most likely base pay range for this position is $175K - $185K per year. Several factors, such as experience, tenure, skills, and particular business needs, will determine an individual’s exact level of compensation. Base salary is only one element of employee compensation at Halozyme. Total compensation could include bonuses, sales incentives, and equity awards.

About the company

At Halozyme, we are reinventing the patient experience and building the future of drug delivery. We are passionate about the important work we do and constantly strive to do more. We embrace transformation and work hard to innovate for the future. We do this together, as One Team - we rise by lifting others up and believe in the power of working together for the collective win. That’s why we need you-to help us make a significant impact by taking on increasingly complex challenges, leaping beyond the status quo, advancing our mission and making our One Team culture thrive.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:13 min

Core components of the internal Optimize ecosystem

Dominik Schneider Dominik Schneider · World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

2:43 min

Deploying a web application through an AI agent workflow

Mike Mike · World Congress 2025

Videos

See all

Related articles

See all