Data Pipeline Engineer

Kforce Inc.
Washington, DC, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work

Tech stack

Java (Programming Language) Big Data C++ (Programming Language) Computer Networks Information Engineering Extract Transform Load (ETL) Software Debugging Distributed Computing Environment Distributed Systems Python (Programming Language) Operational Data Store Operational Databases
+8 more
Data Streaming Data Processing Data Ingestion Delivery Pipeline Pyspark Api Design Code Restructuring Data Pipelines

Job description

This role is heavily focused on maintaining and stabilizing large-scale data pipelines in a production environment. The majority of time is spent troubleshooting and resolving issues across existing data workflows rather than building new systems. Early success in this position looks like gaining enough familiarity with the platform, data flows, and key stakeholders to independently diagnose and resolve pipeline failures across multiple environments., Investigate and resolve data pipeline failures across multiple production environments Perform root cause analysis on data quality and pipeline performance issues Apply targeted code fixes and adjustments to restore pipeline functionality Monitor pipeline health and respond to alerts within defined SLAs Support and maintain existing ETL processes rather than developing new ones Refactor pipelines to resolve performance issues such as memory constraints or inefficient processing Coordinate with upstream data providers and internal teams to resolve data ingestion issues Escalate issues when access, ownership, or dependencies fall outside immediate control

Day-to-Day Breakdown

~85-90%: Debugging, incident response, and pipeline issue resolution ~5-10%: Monitoring, validation, and health checks ~5-10%: Minor code updates, optimizations, and pipeline adjustments

Work is centered on fixing and stabilizing existing pipelines, not building new ones from scratch., Engineers support a large number of pipelines across multiple environments simultaneously Work is highly reactive, driven by incoming alerts and data incidents Engineers are expected to quickly assess and troubleshoot pipelines they have not previously worked on High alert volume, with multiple issues often tied to common root causes

Collaboration

Frequent interaction with data providers to resolve source data issues Regular coordination with cross-functional technical teams on pipeline failures Occasional engagement with end users reporting data discrepancies

On-Call & Incident Response

Rotating on-call schedule supporting different pipeline groups Some rotations may include off-hours alerts tied to overnight pipeline processing Majority of incidents handled during business hours, with occasional escalation scenarios Engineers are expected to own resolution when possible and coordinate when dependencies exist

Requirements

Strong experience with large-scale data engineering and ETL/ELT workflows Proficiency in Python and distributed data processing frameworks (PySpark preferred) Solid understanding of dataframes and data manipulation at scale Experience troubleshooting production data pipelines and debugging failures Knowledge of relational databases and SQL fundamentals Familiarity with distributed computing concepts

Additional Technical Exposure

Experience with Java or similar languages (C++ acceptable alternative) Ability to diagnose and resolve memory/performance issues in distributed jobs Exposure to visual pipeline tools or data workflow platforms is helpful Basic understanding of networking concepts and API-based data ingestion

Operational Environment, Strong foundation in data engineering within production environments Experience supporting operational data systems rather than purely building new solutions Comfortable working in high-volume, incident-driven environments Able to quickly understand and troubleshoot unfamiliar systems Hands-on experience with distributed data processing and large datasets

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on clearancejobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy Ā· LIVE

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff Ā· World Congress 2024

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy Ā· LIVE

3:10 min

Understanding the core concepts of API design

Alen Pokos Ā· LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph Ā· LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy Ā· LIVE

Videos

See all

Related articles

See all