Software Data Engineer

EPAM Systems, Inc.
United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Java (Programming Language) Amazon Web Services Data Analysis Apache HTTP Server Batch Processing Big Data Software Quality Extract Transform Load (ETL) Data Warehousing Apache Hadoop Hadoop Distributed File System Apache Hive
+22 more
Issue Tracking Systems Oracle (Applications) Mesos Software Construction Software Engineering SQL Databases Tableau (Software) Teradata SQL Web Services Data Processing Apache Yarn Snowflake Apache Spark Druid Cassandra Apache Kafka Spark Streaming Presto Vertica Functional Programming Splunk Data Pipelines

Job description

We are seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This is not a data-science or analytics position centered on ad-hoc data exploration; instead, the role focuses on building software, data processing jobs, and data pipelines consumed by internal and external partners. Data is our main product and first-class citizen, and we value correct, high-quality data as much as clean and maintainable code. Responsibilities Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages Fix bugs in code and correct data caused by incorrect logic or implementation Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions Monitor and troubleshoot production issues with pipelines owned by the team Develop and adopt data

Requirements

quality checks to monitor data issues in the systems Scope and plan new development, including assessing level of effort and providing timelines Maintain tickets hygiene in Radar (ticketing system) Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines Requirements 3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3 Excellent knowledge of Scala or Java Understanding of batch processing and ETL principles in Data Warehouses Familiarity with data completeness signals and orchestration Knowledge of approaches for historical reprocessing and data correction Skills in handling bad data and late data in inputs and outputs Understanding of schema migrations and datasets evolution Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more Nice to have Understanding of functional programming ideas and principles Experience in building and using web services Familiarity with any of Teradata, Vertica, Oracle, Tableau Skills in Spark Streaming and Kafka Knowledge of Apache Iceberg, Trino (Presto), Druid, Cassandra, or Blob storage like AWS Experience with Splunk Experience with Snowflake

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:29 min

Evaluating orchestration tools for distributed container deployments

Aleksandr Kalikov · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

4:19 min

Introduction to network security and endpoint monitoring architectures

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all