Principal / Staff Data Platform Engineer

SWEET UNION BAPTIST CHURCH
United States
19 days ago
Apply on career.sigma.software
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Data Analysis Apache HTTP Server Software as a Service Cloud Computing Continuous Integration Data Architecture Data Governance Data Infrastructure Distributed Computing Environment
+18 more
Distributed Data Store Distributed Systems Python (Programming Language) Operational Data Store Operational Databases Performance Tuning Scala (Programming Language) Data Streaming Management of Software Versions Real Time Systems Sql Optimization Apache Spark Event Driven Architecture Infrastructure Automation Frameworks Apache Flink Deployment Automation Data Management Data Pipelines

Job description

Join a greenfield initiative focused on building a next-generation AI-first data platform for a global AdTech ecosystem. We are looking for a Principal/Staff Data Platform Engineer who will drive architectural decisions, establish engineering standards, and shape the foundation for scalable analytics and AI capabilities.

As part of Sigma Software, you will collaborate with experienced engineering teams and contribute to a platform designed around immutable event-driven architecture, governed data contracts, and modern lakehouse principles. This role is ideal for a Principal-level engineer who enjoys solving complex distributed systems challenges and influencing platform strategy from day one.

We offer the opportunity to work on large-scale international products, collaborate with highly skilled professionals, and contribute to a technically ambitious environment with long-term growth potential.

Customer

Our Customer is a Sweden-based AdTech company specializing in advanced self-serve advertising platforms that automate direct transactions between advertisers and major global publishers. Their solutions improve transparency and operational efficiency in digital advertising and are trusted by globally recognized brands including TripAdvisor, Bloomberg, The Washington Post, Opera, and Dow Jones. The company processes millions of advertising transactions worldwide and is actively investing in AI-driven data capabilities.

Project

The project focuses on building a modern data platform from the ground up with an emphasis on immutable event streams, governed canonical models, semantic layers, and scalable analytics infrastructure. The platform will support analytical workloads, operational data products, AI applications, and future customer-facing data experiences.

As a Principal / Staff Data Platform Engineer, you will own key architectural decisions, define scalable engineering standards, and help deliver the first production-ready version of the platform in a high-scale AdTech environment.

Key Technologies: Apache Iceberg, Spark, Flink, Trino, AWS, Python, Scala, Java, CI/CD, Infrastructure as Code, * The first production data foundation is running.

  • Clear architectural decisions have been made and documented.
  • Events entering the data platform have enforceable contracts.
  • Core canonical entities exist and are tested.
  • Data quality and lineage are observable.
  • Other engineers can contribute without needing to understand every implementation detail.

Responsibilities

Responsibilities

  • Lead the technical design and architecture of a modern enterprise-scale data platform
  • Define scalable data flows from Iceberg-based event storage into analytical and operational data products
  • Evaluate and select technologies for orchestration, transformation, querying, serving, and storage
  • Design reusable canonical entities and data models across advertising, campaigns, inventory, billing, and customer domains
  • Establish scalable engineering patterns for batch and near-real-time processing
  • Build reliable and observable data pipelines for high-volume AdTech workloads
  • Implement data quality, lineage, observability, and reconciliation capabilities
  • Collaborate with Platform Engineering teams to introduce CI-enforced data contracts and governance standards
  • Define tenant isolation, access control, and regional data boundary strategies
  • Establish standards for testing, deployment automation, schema evolution, and versioning
  • Optimize platform performance, scalability, and infrastructure costs
  • Mentor engineers and contribute to engineering excellence across the team
  • Partner closely with leadership and cross-functional stakeholders

Requirements

Apache Iceberg / strong Distributed Data Systems / expert AWS / expert Python / strong Spark & Flink / strong, * 8+ years of experience designing and operating production-grade data platforms

  • Strong expertise in distributed data systems and high-volume event-driven architectures
  • Deep understanding of modern lakehouse architectures and Apache Iceberg
  • Advanced SQL skills and production experience with Python, Java, Scala, or similar languages
  • Experience with distributed processing and query technologies such as Spark, Flink, or Trino
  • Strong knowledge of data modeling, partitioning strategies, and performance optimization
  • Proven experience building batch and near-real-time data pipelines
  • Hands-on experience with AWS cloud infrastructure
  • Strong understanding of CI/CD pipelines, Infrastructure as Code, and production observability
  • Experience implementing data contracts, schema evolution, and data quality frameworks
  • Ability to make pragmatic architectural decisions in greenfield environments
  • Upper-Intermediate or higher English level
  • Strong communication and technical leadership skills

WILL BE A PLUS

  • Experience in AdTech or other high-volume event-processing domains
  • Experience designing multi-tenant SaaS data architectures
  • Familiarity with semantic layer technologies
  • Experience supporting both analytics and ML/AI workloads
  • Understanding of privacy regulations and data residency requirements

About the company

  • Diversity of Domains & Businesses
  • Variety of technology
  • Health & Legal support
  • Active professional community
  • Continuous education and growing
  • Flexible schedule
  • Remote work
  • Outstanding offices (if you choose it)
  • Sports and community activities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on career.sigma.software
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:48 min

Automating exploratory data analysis within training pipelines

Dora Petrella · World Congress 2023

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:55 min

Infrastructure challenges when combining Kafka with Apache Flink

Bobur Umurzokov · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · World Congress 2024

2:04 min

Comparing offline data analytics with online stream processing

Artem Volk Artem Volk +1 · World Congress 2024

Videos

See all

Related articles

See all