Data scientist machine, learning, AWS and data bricks

Elevate IQ Inc
Plano, TX, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$76,736.0 - $92,413.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Amazon Web Services Amazon S3 Data Analysis Big Data Cloud Database Cluster Analysis Computer Programming Databases Continuous Integration Data Cleansing Data Infrastructure
+38 more
Database Queries Revision Control Systems Identity and Access Management Python (Programming Language) Linear Regression Machine Learning Natural Language Processing NumPy Scrum Methodology SQL Databases Unstructured Data Data Processing Cloud Platform System Feature Engineering Random Forest Apache Spark State Machines Deep Learning Model Validation AWS Lambda Git Pandas Pyspark Scikit Learn Infrastructure Automation Frameworks Information Technology Xgboost Feature Selection Machine Learning Operations Functional Programming Cloudwatch Api Gateway Restful APIs Software Version Control Serverless Computing Docker Unsupervised Learning Databricks

Job description

We are seeking an experienced Data Scientist with strong expertise in machine learning, Python, AWS, and Databricks. The ideal candidate will be responsible for analyzing complex datasets, developing predictive models, and deploying scalable machine-learning solutions in cloud environments., Collect, clean, transform, and analyze structured and unstructured datasets.

  • Perform exploratory data analysis to identify patterns, trends, correlations, and business insights.

  • Develop, train, test, and optimize machine-learning models using algorithms such as:

  • XGBoost

  • Random Forest

  • Decision Trees

  • Logistic and Linear Regression

  • Gradient Boosting

  • Clustering and other supervised or unsupervised learning techniques

  • Perform feature engineering, feature selection, and hyperparameter tuning.

  • Evaluate models using appropriate metrics such as accuracy, precision, recall, F1-score, ROC-AUC, RMSE, and MAE.

  • Build reusable data-processing and machine-learning pipelines using Python, Pandas, and NumPy.

  • Work with large-scale datasets using Databricks, Apache Spark, and PySpark.

  • Develop and deploy cloud-based data science solutions on AWS.

  • Build serverless workflows and model-processing components using AWS Lambda.

  • Work with AWS services such as S3, SageMaker, Glue, Step Functions, CloudWatch, IAM, and API Gateway.

  • Deploy, monitor, maintain, and retrain machine-learning models in production environments.

  • Collaborate with data engineers, cloud engineers, software developers, analysts, and business stakeholders.

  • Translate business requirements into analytical and machine-learning solutions.

  • Document model assumptions, methodologies, performance, limitations, and results.

  • Follow machine-learning development, testing, version control, security, and governance best practices.

Requirements

The candidate should have hands-on experience with algorithms such as XGBoost and Random Forest, along with strong proficiency in Python libraries including Pandas and NumPy., Bachelor’s or Master’s degree in Computer Science, Data Science, Statistics, Mathematics, Engineering, or a related field.

  • Strong programming experience with Python.

  • Strong experience with Pandas, NumPy, and Scikit-learn.

  • Hands-on experience building machine-learning models using XGBoost and Random Forest.

  • Strong understanding of supervised and unsupervised machine-learning techniques.

  • Experience with data preprocessing, feature engineering, model validation, and hyperparameter optimization.

  • Hands-on experience with Databricks, Spark, or PySpark.

  • Experience working with AWS services, particularly AWS Lambda and Amazon S3.

  • Strong SQL skills and experience working with relational or cloud-based databases.

  • Understanding of model deployment, monitoring, scalability, and production support.

  • Experience using Git or another version-control system.

  • Strong analytical, problem-solving, and communication skills.

Preferred Qualifications

  • Experience with Amazon SageMaker and MLflow.

  • Experience developing end-to-end machine-learning pipelines.

  • Knowledge of MLOps, CI/CD, Docker, and Infrastructure as Code.

  • Experience building REST APIs for machine-learning model inference.

  • Knowledge of time-series forecasting, natural language processing, or deep learning.

  • Experience working in Agile or Scrum development environments.

  • AWS or Databricks certifications are a plus.

Key Technical Skills

Programming: Python, SQL

Data Processing: Pandas, NumPy, PySpark

Machine Learning: XGBoost, Random Forest, Scikit-learn, Gradient Boosting

Cloud: AWS, Lambda, S3, SageMaker, Glue, Step Functions, CloudWatch

Big Data Platform: Databricks, Apache Spark

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all