Big Data Engineer

Resourcesys Inc.
United States
25 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Build Automation Big Data Cluster Analysis Continuous Integration Data Dictionary Information Engineering Data Security Python (Programming Language) Laboratory Information Management Systems Machine Learning
+25 more
SQL Azure Message Queuing Telemetry Transport (MQTT) Performance Tuning Role-Based Access Control Power BI Standard Sql OPC Unified Architecture SAP (Applications) Simple Data Format SQL Databases Systems Integration Management of Software Versions Data Ingestion Azure Data Factory Mttr Caching Git Data Layers Pyspark Integration Tests SAP S/4HANA Operational Systems Spark Streaming Key Vault Databricks

Job description

Client is seeking a Big Data Engineer to build and operate our enterprise Unified Data Layer (UDL) - spanning IT and OT - to deliver trustworthy, performant data products that power Finance, Operations, Supply Chain & Logistics, HSE, Commercial, and corporate analytics. You’ll engineer batch/CDC/streaming pipelines, model curated/semantic layers, and harden run-state with testing, CI/CD, security, and observability. You’ll partner closely with the data team and larger IT organization.

Mission

Design and deliver scalable, secure data pipelines and data models that safely connect operational systems to analytics, ensure trusted and well governed data, and enable repeatable delivery of BI, ML, AI, and automation solutions.

Data Engineering & Modeling

Build ingestion pipelines (batch, CDC, streaming) from S/4HANA/DataSphere, PHD/historian, LIMS, TMS, HSE, and other sources into landing curated semantic layers.

Implement data contracts, schema/versioning, SCD handling, partitioning, and performance tuning (file formats, clustering, caching).

Develop dimensional/semantic models that back certified Power BI datasets and APIs for apps/agents.

OT/IT Integration & Safety

Integrate OT data via OPC UA/MQTT, broker/DMZ patterns, read-only historian feeds, and event/batch frames-no control-net reads.

Collaborate with plant controls on change control, signal quality, and downtime windows.

Quality, Security & Observability

Embed data quality rules, unit/integration tests, and validation checks (freshness, completeness, drift/PSI).

Instrument lineage and end-to-end monitoring; build alerting and on-call runbooks to minimize MTTR.

Enforce RBAC, secrets management, PII/HSE classifications, and retention aligned to Governance/MDM policies.

CI/CD, Cost & Reliability

Automate build/test/deploy with Git-based CI/CD (environments, approvals, blue/green).

Track and optimize cost/performance (cluster sizing, autoscaling, cache strategy); contribute to FinOps reviews.

Collaboration & Documentation

Partner with Reporting & BI on semantic model contracts, RLS, and performance SLAs; avoid direct system scraping.

Produce ā€œreadmeā€ docs, data dictionaries, runbooks, and post-incident reviews; support knowledge transfer with vendors.

Requirements

Minimum 5 years’ in data engineering building production pipelines at scale (batch/CDC/streaming).

Hands-on with Azure data stack: Databricks or Fabric/Synapse, ADF/Pipelines, ADLS/OneLake, Azure SQL/SQL MI, Key Vault.

Strong SQL and Python/PySpark; comfort with Spark Structured Streaming and performance tuning.

Experience implementing tests/observability (freshness, schema, expectations), and Git-based CI/CD.

Familiarity with SAP S/4HANA structures and SAP DataSphere semantic modeling.

OT concepts: historians (PHD/PI), OPC UA/MQTT, event/batch frames, ISA-95/99 basics.

Understanding of Power BI consumption (semantic models, RLS) and APIs for downstream AI/ML apps/agents.

Preferred Qualifications:

Time-series/data-quality tooling (e.g., Great Expectations or equivalent patterns), feature/metric stores.

MDM concepts (keys, survivorship), lineage/catalog tooling.

TMS/WMS, LIMS, Historian, HSE domain exposure; Lean/Six Sigma mindset; FinOps awareness.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon Ā· WWC Europe 2026

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn Ā· WWC Europe 2026

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley Ā· WWC 2021

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer Ā· Coffee With Developers

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser Ā· Europe 2026 Virtual

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz Ā· LIVE

Videos

See all

Related articles

See all