Remote | Data Engineer

Mag Llc.
New York, NY, United States
1 day ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$140,000.0 - $180,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Big Data Cloud Computing Cloud Engineering Databases Data as a Services Data Architecture Information Engineering Data Files Data Infrastructure Data Systems
+23 more
Data Visualization Distributed Computing Environment Distributed Data Store Distributed Systems Systems Analysis Python (Programming Language) Machine Learning NoSQL Operational Databases Performance Tuning Standard Sql Requirements Management Software Engineering SQL Databases Unstructured Data Data Processing Scripting Large Language Models Apache Spark Matplotlib Plotly Data Management Data Pipelines

Job description

We are sharing a full-time opportunity for an experienced Data Engineer with strong expertise in Python, SQL, Apache Spark, AWS, distributed data processing, and scalable data architecture to build and operate infrastructure supporting AI-driven products and research initiatives.

The role will focus on designing and scaling distributed data pipelines, managing large datasets across cloud environments, and building reliable systems for analytics, experimentation, and model development., Data Pipelines & Distributed Processing

  • Design, build, and maintain large-scale pipelines for structured and unstructured data
  • Develop distributed processing workflows using Apache Spark or comparable frameworks
  • Optimise transformations, partitioning strategies, and computational workloads
  • Identify and resolve performance bottlenecks across high-volume data systems
  • Support downstream analytics, experimentation, and model-development requirements

Cloud Architecture & Data Engineering

  • Design scalable AWS-based data architectures across SQL and NoSQL systems
  • Build reliable ingestion, transformation, storage, and distribution workflows
  • Write efficient Python and SQL for production data processing
  • Evaluate storage and database technologies against workload requirements
  • Improve scalability, maintainability, accessibility, and operational efficiency

Data Quality, Reliability & AI Support

  • Implement monitoring, validation, and automation across data workflows
  • Identify failures, anomalies, and data-quality issues
  • Maintain integrity and reliability throughout pipelines and storage layers
  • Collaborate with AI researchers, data scientists, and engineering teams
  • Support data infrastructure for AI/ML training, evaluation, and experimentation

Requirements

  • Strong professional experience in data engineering or distributed data systems
  • Advanced proficiency in Python and SQL
  • Hands-on experience with Apache Spark or comparable distributed-processing frameworks
  • Strong experience with AWS data services and cloud-native architecture
  • Experience with SQL and NoSQL databases
  • Demonstrated experience processing large-scale datasets
  • Strong understanding of partitioning, performance optimisation, and scalable architecture
  • Familiarity with orchestration, automation, monitoring, and data-quality workflows
  • Exposure to AI/ML or research environments is advantageous
  • Familiarity with LLM training, evaluation, or experimentation datasets is beneficial
  • Experience with data-visualisation tools such as Matplotlib, Seaborn, or Plotly is a plus, Amazon Web Services (AWS), Apache Spark, Artificial Intelligence (AI), Automation, Cloud Architecture, Cloud Computing, Data Management, Data Processing, Data Quality, Data Science, Data Sets, Data Visualization Tools, Database Technology, Distributed Computing, NoSQL, Performance Tuning/Optimization, Product Programs, Project Evaluation, Python Programming/Scripting Language, Requirements Management, SQL (Structured Query Language), SQL Databases, Sales Pipeline, Scientific Research, Software Engineering, Structured Data, Systems Analysis, Technical Consulting, Unstructured Data

Benefits & conditions

Engagement Details

  • Full-time engagement
  • Fully remote
  • Base compensation: $140,000-$180,000/year
  • Work will involve Python, SQL, Apache Spark, AWS, distributed data processing, and scalable data architecture
  • Responsibilities will span data ingestion, transformation, storage, monitoring, and operational reliability
  • The role may support AI/ML experimentation, model-development workflows, and LLM-related data infrastructure
  • Data volumes, infrastructure requirements, and technical priorities may evolve as products and research initiatives scale

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:28 min

Identifying root causes through global and local SHAP plots

Bernhard Bernhard +1 · World Congress 2025

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all