Machine Learning Engineer

DESMATA INC.
New Paltz, NY, United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$114,400.0
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Data Analysis Big Data Data Cleansing Information Engineering Data Governance Data Mining Data Visualization Relational Databases Machine Learning Natural Language Processing NumPy
+18 more
Performance Tuning Power BI Standard Sql Unstructured Data Google Cloud Cloud Platform System Feature Engineering Large Language Models Apache Spark Model Validation Pandas Pyspark Scikit Learn Information Technology Data Analytics Machine Learning Operations Data Pipelines Databricks

Job description

Advanced Analytics & Machine Learning

  • Design, develop, and optimize machine learning models (forecasting, classification, clustering).
  • Apply data mining techniques to uncover patterns and insights in large datasets.
  • Perform feature engineering, model validation, and performance tuning.
  • Explore and deploy modern AI and ML approaches to enhance automation and analytics.

Data Preparation & Quality

  • Prepare structured and unstructured data for modeling and advanced analysis.
  • Develop scripts and tools for data cleansing, validation, and enrichment.
  • Collaborate with Data Engineering to maintain efficient data pipelines.
  • Identify data quality issues and propose remediation.

Analytics, Insights & Reporting

  • Conduct deep-dive analyses to identify trends and improvement opportunities.
  • Communicate complex findings in clear, concise ways to technical and non-technical stakeholders.
  • Support the development of dashboards, metrics, and analytical solutions.

Cross-Team Collaboration

  • Work with architects, engineers, and analysts to define analytical requirements.
  • Contribute to conceptual data model design and workflow optimization.
  • Promote best practices in machine learning, analytics, and data governance.

Requirements

Do you have experience in Spark?, Do you have a Bachelor’s degree?, * Bachelor’s or Master’s degree in Computer Science, Data Science, Machine Learning, Statistics, Mathematics, or related field.

  • Strong experience in machine learning algorithms, predictive modeling, and data mining.
  • Proficiency in PySpark, Python Pandas (required) for data science workloads.
  • Strong SQL (required) knowledge and experience with relational databases.
  • Minimum 3 years of experience with data visualization tools such as Power BI, DAX Queries, and best practices.
  • Experience with Azure Databricks, Google Cloud, and modern data science libraries (e.g., scikit-learn, pandas, NumPy).
  • Experience with GenAI and large language models.
  • Ability to interpret complex datasets and produce actionable insights.
  • Must know how to analyze the root cause of dashboard errors.
  • Have experience in ML Ops and have a strong coding background.
  • Have experience with Natural Language Processing (NLP).
  • Knowledge or experience with A/B Testing.
  • Working knowledge of designing, training, and implementing machine learning models.
  • Familiarity with cloud-based infrastructure.
  • Excellent communication and problem-solving skills.
  • 7 or more years of experience in data science and machine learning engineering.
  • Additional Skills (Skills that are a plus, but not required)
  • Knowledge of statistical methods and experimental design.

Benefits & conditions

From $55 an hour - Full-time, Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

6:58 min

Analyzing production code coverage data using pandas

Markus Harrer Markus Harrer · World Congress 2021

Videos

See all

Related articles

See all