Data Engineer (Temporal & Apache Kafka required)

Infinitive
McLean, VA, United States
12 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$90,000.0
Working hours
Regular working hours

Tech stack

Clean Code Principles Query Performance Java (Programming Language) Amazon Web Services Data Analysis Apache HTTP Server Automation of Tests Big Data BigQuery Computer Programming Continuous Integration Data Validation
+37 more
Information Engineering Data Governance Extract Transform Load (ETL) Data Warehousing Relational Databases Software Design Patterns Distributed Data Store Distributed Systems Fault Tolerance Protocol Buffers Apache Hive JSON Python (Programming Language) Machine Learning Open Source Technology Software Engineering SQL Databases Data Streaming Management of Software Versions Workflow Management Systems Sql Optimization Snowflake Apache Spark Pyspark Integration Tests Avro AWS Glue Star Schema Apache Kafka Data Management Stream Processing Grpc Data Pipelines Confluent Amazon Redshift Databricks Microservices

Job description

Infinitive is a data and AI consultancy that helps clients modernize, monetize, and operationalize their data to generate lasting value. They pride themselves on their deep industry and technology expertise, ensuring that they drive and sustain the adoption of new capabilities. Infinitive is committed to aligning their team with their clients’ culture, ensuring a successful partnership by bringing the right mix of talent and skills for high return on investment. Infinitive has earned recognition as one of the “Best Small Firms to Work For” by Consulting Magazine, receiving this accolade nine times, most recently in 2026. They have also been honored as a “Top Workplace” by the Washington Post, “Best Places to Work” by the Washington Business Journal, and “Best Places to Work” by Virginia Business., We are seeking an experienced Data Engineer to help design, build, and scale our next-generation event-driven data platforms. In this role, you will be instrumental in bridging high-throughput distributed streaming with complex, fault-tolerant workflow orchestration and strict data governance. You will work extensively with Apache Kafka for real-time event streaming and Temporal (the open-source, durable execution engine originating from Uber/Cadence) to build resilient, distributed stateful workflows and data pipelines. A core focus of this position is establishing robust data schema design and automated validation to ensure strong data contracts across distributed systems. Alongside these technologies, you will design robust batch and streaming ETL/ELT pipelines leveraging Python, Apache Spark, and modern cloud data warehouses/lakehouses., Stream Processing & Messaging: Architect, deploy, and maintain high-volume distributed data streams using Apache Kafka (producers, consumers, Kafka Connect, Schema Registry). Data Schema Design & Validation: Establish and enforce schema design standards, versioning strategies, and automated schema validation (e.g., Avro, Protobuf, JSON Schema) to maintain strict data contracts across microservices, streaming consumers, and lakehouse storage. Resilient Workflow Orchestration: Design and implement durable execution workflows using Temporal to coordinate long-running distributed pipelines, compensate transactions (Saga pattern), and manage cross-system ETL tasks. Pipeline Development: Build end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark / PySpark. Data Modeling & Warehousing: Design and optimize analytical data models (dimensional/star schema) in modern cloud data warehouses/lakehouses (e.g., Snowflake, BigQuery, Databricks, Redshift). Reliability & Data Quality: Implement automated testing, continuous schema validation, data drift detection, and observability across streaming and batch workflows. * Cross-Functional Collaboration: Partner with software engineers, machine learning engineers, and analysts to define standard schema definitions, data contracts, and production-grade CI/CD release patterns., Lead Data Engineer (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI) Do you love building and pioneering in the technology space? Do you enjoy solving complex busine…

  • 13 days ago, Lead Data Engineer (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI) Do you love building and pioneering in the technology space? Do you enjoy solving complex busine…
  • 16 days ago, Lead Data Engineer (Python, AWS, Spark, Kafka, SQL, Snowflake, Databricks, GenAI) Do you love building and pioneering in the technology space? Do you enjoy solving complex busine…
  • 1 month ago +

Requirements

4+ years of professional experience in data engineering, backend distributed systems, or software engineering. Hands-on experience with Temporal (or Cadence): Proven understanding of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration. Deep expertise with Apache Kafka: Practical experience with message partitioning, consumer groups, offset management, and topic design. Strong background in Data Schema Design & Validation: Demonstrated proficiency with schema definition frameworks (Apache Avro, Protocol Buffers/gRPC, or JSON Schema). Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry). Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests). Strong programming proficiency in Python (Go or Java is a plus) with clean code, design patterns, and unit/integration testing standards. Distributed computing experience: Hands-on development with Apache Spark (PySpark/Spark SQL) processing large-scale datasets. Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning.

Benefits & conditions

Infinitive is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

1:11 min

Evaluating architectural trade-offs between REST and gRPC

Sakshi Nasha Sakshi Nasha · Europe 2026 Virtual

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all