Remote

MAG 24 LLC
New York, NY, United States
1 day ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$140,000.0 - $180,000.0
Working hours
Regular working hours

Tech stack

Query Performance Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis Databases Extract Transform Load (ETL) Relational Databases Database Design Database Schema Python (Programming Language) PostgreSQL Machine Learning
+12 more
MySQL NumPy Standard Sql SQL Databases Unstructured Data Jupyter Notebook Data Processing Pandas Scikit Learn HuggingFace Machine Learning Operations Data Pipelines

Job description

We are sharing a full-time opportunity for an experienced Data Engineer with strong expertise in Python, SQL, ETL pipelines, data modelling, exploratory data analysis, and scalable data-processing workflows to support research, analytics, and AI/ML initiatives. The role focuses on transforming structured and unstructured information into reliable, production-quality datasets. The successful candidate will build scalable data pipelines, improve data quality and validation, optimise SQL and processing workflows, and collaborate with researchers, data scientists, and engineers., Data Pipelines & Processing

  • Design, build, and maintain scalable ETL pipelines
  • Collect, clean, normalise, and transform structured and unstructured data
  • Automate recurring processing workflows
  • Monitor pipeline performance and troubleshoot failures
  • Improve efficiency, scalability, and maintainability

SQL, Databases & Data Modelling

  • Write and optimise SQL queries for extraction and transformation
  • Work with relational databases such as PostgreSQL and MySQL
  • Design and maintain schemas and data models
  • Improve query performance and data-access patterns
  • Support scalable and reliable data-storage architectures

Data Quality & Analysis

  • Conduct exploratory data analysis to identify patterns, anomalies, and quality issues
  • Build automated validation and quality-control workflows
  • Monitor accuracy, completeness, integrity, and consistency
  • Investigate data issues and implement corrective actions
  • Produce clear analytical summaries for technical stakeholders

Python & AI/ML Data Support

  • Build processing workflows using Python, Pandas, and NumPy
  • Develop reusable components for transformation and analysis
  • Prepare datasets for AI and machine-learning initiatives
  • Collaborate with researchers and data scientists on data requirements
  • Support reliable training, evaluation, and experimentation workflows, * Work will involve Python, SQL, ETL, exploratory data analysis, database design, data modelling, validation, automation, and AI/ML dataset preparation
  • Responsibilities may involve both structured and unstructured data
  • Regular collaboration with researchers, data scientists, and engineering teams is expected
  • Project scope, data sources, and technical requirements may evolve based on business and research needs
  • Work must be completed without using confidential, proprietary, regulated, client-identifiable, restricted-access, or otherwise protected information belonging to any employer, client, research institution, data provider, or other third party

About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy

Requirements

  • Strong proficiency in Python and SQL
  • Hands-on experience designing and maintaining ETL pipelines
  • Experience with exploratory data analysis
  • Proficiency with Pandas and NumPy
  • Experience with PostgreSQL, MySQL, or comparable relational databases
  • Strong understanding of data modelling and database schemas
  • Experience with structured and unstructured datasets
  • Strong focus on data quality, integrity, and reliability
  • Strong analytical and problem-solving skills
  • Experience collaborating with researchers, data scientists, or engineering teams
  • Familiarity with Jupyter Notebook, VS Code, PyCharm, or similar tools
  • Exposure to AI/ML workflows is advantageous
  • Familiarity with scikit-learn, Hugging Face Transformers, or AI APIs is beneficial

Engagement Details

  • Full-time engagement
  • Fully remote

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:01 min

Executing remote data exploration and model training

Mingshen Sun Mingshen Sun · World Congress 2024

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all