Data Engineer - Hybrid

SmartIMS Inc.
Arlington, VA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$83,200.0 - $166,400.0
Working hours
Regular working hours

Tech stack

Clean Code Principles Big Data Code Review Extract Transform Load (ETL) Data Systems Database Design Database Testing Distributed Computing Environment Apache Hadoop Hadoop Distributed File System Apache Hive SQL Databases
+12 more
Enterprise Data Management Software Organization Data Processing Apache Yarn Sql Optimization Apache Spark Pyspark Information Technology Spark Streaming Data Management Software Version Control Data Pipelines

Job description

As a Data Engineer, you will support the design, development, and maintenance of enterprise data platforms and large-scale data processing solutions. You will be responsible for building and optimizing data pipelines, developing ETL processes, and ensuring the availability, reliability, and quality of data across business-critical systems. This role requires expertise in big data technologies, SQL optimization, and distributed data processing, along with the ability to collaborate with cross-functional teams to deliver scalable and efficient data solutions., * Design, implement, and maintain enterprise ETL processes and data pipelines

  • Develop scalable and efficient code to process, transform, and deliver large datasets
  • Build and optimize data pipelines using Apache Spark, Hadoop, and related big data technologies
  • Collaborate with engineering and analytics teams to solve complex data challenges and maintain data quality
  • Support the delivery of accurate and actionable data solutions for business stakeholders
  • Design and manage distributed data processing workflows and orchestration processes
  • Develop and optimize SQL queries for large-scale data retrieval, transformation, and analysis
  • Participate in data modeling and database design initiatives to support scalable solutions
  • Monitor, troubleshoot, and resolve data processing and pipeline issues
  • Automate routine data management tasks and improve operational efficiency
  • Apply testing and validation practices to ensure data accuracy, consistency, and reliability
  • Participate in code reviews and follow development best practices and version control standards
  • Build strong working relationships with internal teams and business stakeholders
  • Ensure compliance with organizational policies, standards, and regulatory requirements

Requirements

  • Experience as a Data Engineer or in a similar data-focused engineering role
  • Strong expertise in writing and optimizing SQL queries for large datasets
  • Hands-on experience with Apache Spark, including PySpark, Spark SQL, and Spark Streaming
  • Experience working with Hadoop ecosystem technologies such as HDFS, Hive, and YARN
  • Strong understanding of ETL frameworks and data pipeline development
  • Knowledge of distributed data processing and big data architectures
  • Understanding of data modeling concepts and database design principles
  • Experience working with Python for data engineering and automation tasks
  • Ability to analyze, troubleshoot, and resolve complex data issues independently
  • Knowledge of data testing, validation, and quality assurance practices
  • Strong verbal and written communication skills with the ability to collaborate with technical and non-technical stakeholders
  • Experience with version control, code reviews, and software development best practices
  • Ability to work effectively in a collaborative, fast-paced environment
  • Bachelor’s degree in Engineering, Mathematics, Finance, Business, Computer Science, or a related quantitative field, or equivalent practical experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all