Data Scientist - Feature Engineering and Machine Learning

Sierra Business Solution LLC
Plano, TX, United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Business Analytics Applications Data Analysis Big Data Information Engineering Data Transformation Data Profiling Data Reduction Database Queries Text Processing Statistical Hypothesis Testing Python (Programming Language) Machine Learning
+14 more
Regression Analysis Natural Language Processing NumPy Operational Databases Performance Tuning SQL Databases Feature Engineering Pandas Scikit Learn Statistics Packages Feature Selection Machine Learning Operations Feature Extraction Data Pipelines

Job description

We are looking for a highly analytical Data Scientist with strong expertise in feature selection, statistical modeling, Machine Learning, NLP, Data Engineering, Python, and SQL. The ideal candidate will focus on identifying, engineering, and validating the most impactful features from structured and unstructured datasets to improve predictive analytics and business outcomes., Lead feature selection initiatives using advanced statistical techniques such as correlation analysis, hypothesis testing, regression analysis, Information Value (IV), Weight of Evidence (WoE), PCA, and feature importance methods.

Analyze large and complex datasets to identify, evaluate, and prioritize key variables that significantly influence business outcomes and predictive performance.

Design, develop, and maintain scalable feature engineering frameworks for both structured and unstructured data sources.

Perform exploratory data analysis (EDA), data profiling, and statistical validation to uncover meaningful patterns, relationships, and predictive features.

Utilize Python and SQL to extract, transform, analyze, and validate data while ensuring data quality and consistency across analytical workflows.

Build, train, validate, and optimize Machine Learning models using appropriate algorithms and techniques to solve business problems and improve predictive accuracy.

Apply Machine Learning algorithms to assess feature effectiveness, validate feature sets, and improve model accuracy, robustness, and interpretability.

Leverage NLP techniques to extract, engineer, and optimize features from textual data for downstream analytics and predictive modeling.

Collaborate closely with business stakeholders, product teams, and data engineers to translate business requirements into meaningful analytical features, predictive models, and actionable insights.

Build, optimize, and productionize data pipelines and feature datasets, working with Data Engineering teams to ensure scalability, efficiency, governance, and operational readiness.

Conduct Data Engineering handover activities by documenting data pipelines, feature engineering logic, model inputsoutputs, transformation rules, and deployment requirements to ensure seamless transition to engineering and operations teams.

Support model deployment, monitoring, performance tracking, and continuous improvement by partnering with Data Engineering and MLOps teams.

Document feature selection methodologies, model development processes, statistical findings, assumptions, and recommendations while establishing best practices for reusable and automated analytics solutions.

Requirements

Feature Selection, Feature Engineering, Statistical Modeling, Model Building, Machine Learning, Python, SQL, NLP, Data Engineering, ETLELT, Data Pipeline Development.

Other Requirements

8+ years of experience in Data Science, Analytics, or Machine Learning.

Strong expertise in Feature Selection, Feature Engineering, Statistical Modeling, and Predictive Modeling.

Deep understanding of Statistics, Hypothesis Testing, Regression Analysis, Feature Importance Techniques, and Dimensionality Reduction.

Hands-on experience in Machine Learning model development, validation, tuning, and evaluation.

Strong proficiency in Python (Pandas, NumPy, Scikit-learn, Statsmodels).

Strong SQL skills with experience writing complex queries and performing large-scale data analysis.

Experience with NLP techniques, text processing, and feature extraction.

Knowledge of Data Engineering concepts, ETLELT pipelines, data transformation, and production data workflows.

Experience with model deployment, MLOps concepts, and collaboration with Data Engineering teams., Excellent analytical, problem-solving, documentation, and stakeholder communication skills.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

Videos

See all

Related articles

See all