> Markdown version of [/jobs/ext/2284299-data-scientist](https://www.wearedevelopers.com/jobs/ext/2284299-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DATA SCIENTIST - **Company:** Sally Beauty Supply - **Location:** Plano, TX, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Microsoft Excel, A/B Testing, Agile Methodology, Amazon Web Services, Data Analysis, Automation of Tests, Microsoft Azure, Cloud Computing, Code Review, Continuous Integration, Data Validation, Information Engineering, Data Integrity, Data Reduction, Serialization, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, NumPy, Operational Data Store, Microsoft PowerPoint, Recommender Systems, SQL Databases, Feature Engineering, Apache Spark, Deep Learning, Git, Pandas, Containerization, Data Lakes, Pyspark, Scikit Learn, Information Technology, Production Code, Performance Monitor, Feature Selection, Machine Learning Operations, Restful APIs, Software Version Control, Data Pipelines, K Means, Databricks - **Published:** August 28, 2026 - **Apply:** https://eigx.fa.us6.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_13/requisitions/preview/53847?utm_medium=jobshare ## About the Role * Education: Master's degree in mathematics / Statistics/Data Science and Analytics, Computer Science, Economics, Physics, or a related field required. Master's degree preferred. * Experience: 4+ years of hands-on experience in data science, applied machine learning, or customer analytics. * Industry: Experience in Retail, CPG, e-commerce, or subscription/loyalty-driven businesses is highly desirable. Technical Skills * Core Languages: Advanced proficiency in Python (pandas, NumPy, scikit-learn) and SQL. Experience with R is a plus. * Big Data & Cloud: Hands-on experience with Databricks, Spark/PySpark, Delta Lake, and a major cloud environment (Azure preferred; AWS/GCP acceptable). * Machine Learning: Regression, classification, clustering (K-Means), tree-based and boosting methods, survival analysis, recommender systems, and dimensionality reduction (PCA). Exposure to deep learning frameworks is a plus. * Statistical Knowledge: Solid grounding in hypothesis testing, confidence intervals, experimental design, cross-validation, and basic probability and linear algebra. * MLOps & Engineering: Experience with model deployment and monitoring (e.g., MLflow), model re-training automation, drift detection, version control (Git), and code review practices. Exposure to REST APIs, containerization, or orchestration tooling is a plus. * Presentation: Strong PowerPoint and Excel skills, with the ability to build an executive-ready narrative., Ability to "connect the dots" across disparate data points to form a cohesive business recommendation. A self-starter who operates independently, seeks opportunities to scale and automate, is comfortable challenging the status quo, and can present findings credibly to non-technical senior leadership. Outstanding organizational skills and dedication to quality and integrity, with the ability to contribute effectively in a fast-paced environment. * Passionate Learner - inquisitive about the business; open to feedback and coaching, applies learning quickly; applies learning to improve processes and procedures, proactively shares learning with colleagues and leaders; realigning and reshaping projects * Flexible & Agile Adapter - responsive and open to change; works well with ambiguity; adapts to new plans or directions; keeps calm under pressure; perseveres to achieve the plan/task; doesn't dwell on the past * Effective Communicator - articulates in an appropriate and accurate manner; emotionally astute while remaining authentic to own style/self; encourages others to express views and opinions; demonstrates active listening and uses probing questions; is concise and relevant with data/info * Team Builder - references the importance of teamwork and actively demonstrates collaboration and sharing; builds and/or participates in effective teams; values the importance of inclusion and various sources of thought/input; humble when operating within a team * Customer Focused Partner - understands internal and external customer needs; contributes to plans and actions to improve the associate and customer journey/experience; holds self and team accountable for improving the customer experience; is an advocate for the customer * Strategic Thinker - progressive thinking with the ability to bring new ideas to life; works with others to develop progressive and cost-effective strategies; provides suggestions to improve upon continuous improvement and scalability within department; uses a broad range of data sources * Big Picture Thinker - understands own department and how other key departments operate; adopts an inclusive approach; seeks feedback, reviews progress, and adapts plans as needed; understands interdependencies with other departments * Results Driver - effective at driving and delivering on plans; holds self accountable to a high standard of delivery; suggests opportunities for innovation and continuous improvement; focuses on the right priorities and uses resources/time wisely; demonstrates grit and determination * Problem Solver & Decision Maker - able to consume department/operational data to identify business issues; identifies, gathers, and examines the relevant information; makes recommendations and takes action to solve challenges; considers importance/impact of decisions against relevant factors ## Description We are seeking a hands-on Data Scientist to build the models and data products that power how SBH understands and activates its customers. This is a technical individual-contributor role for a data scientist who is equally comfortable engineering a feature pipeline in Databricks, defending a model's validation approach, and explaining the resulting business recommendation to a non-technical stakeholder. Your mission is to increase Customer Lifetime Value (CLV) by developing the predictive assets behind hyper-personalized CRM campaigns, retention strategies, and proactive churn reduction. You will own models end to end, from problem framing and feature engineering through training, validation, deployment, and ongoing performance monitoring, and ensure they run reliably in production rather than living in a notebook. This role sits at the intersection of advanced analytics and commercial strategy and is critical to advancing SBH's Understand & Activate Customer agenda. You will work closely with Data Engineering, CRM, Marketing, and Merchandising partners, and will help mature our Databricks environment from a place where analysis happens into a platform that delivers repeatable, monitored, production-grade customer intelligence., Model Development & Advanced Analytics (50%) * Predictive Modeling: Build machine learning models through all phases of development, design, feature engineering, training, evaluation, validation, and implementation, using Python. Core use cases include churn and retention modeling, propensity-to-buy, CLV prediction, customer persona/segmentation, and next-best-action recommendations. * Segmentation Strategy: Move beyond basic demographics to build behavioral, psychographic, and value-based segments (e.g., RFM, K-Means clustering, propensity tiers) that CRM and Marketing can activate directly. * Advanced Methodologies: Apply Market Basket Analysis (MBA), survival analysis, uplift/incrementality modeling, and recommender approaches to uncover cross-sell, up-sell, and hidden revenue opportunities. * Statistical Rigor: Produce summary statistics, distribution and correlation studies, and disciplined feature selection; apply confidence intervals, cross-validation, and appropriate model performance validation to every deliverable. * Ad Hoc Agility: Prioritize and deliver critical ad hoc analysis supporting the SALLY plan and forecast, balancing speed with statistical accuracy. Data Engineering for Analytics & Productionization (30%) * Data Pipelines: Develop and maintain the customer 360 view and supporting data pipelines in Databricks using Python, PySpark, and SQL. You should be comfortable building your own pipelines to enable your analysis rather than waiting on others. * Production-Ready Code: Convert proof-of-concept code into production-ready, re-usable components integrated into products and services, with attention to scalability (compute, memory, I/O, model serialization, caching). * MLOps: Automate ML processes such as scheduled scoring and model re-training, and implement monitoring for data drift, model drift, and accuracy degradation. Contribute to impact assessment, back-testing, explainability, reproducibility, and data quality checks. * Engineering Discipline: Follow source control, peer code review, and automated testing practices (Git/Azure DevOps); contribute to CI/CD for analytics assets. Experimentation, Test & Learn (20%) * Test Design: Design and analyze A/B and multivariate tests for email, SMS, push, and in-app campaigns to optimize engagement, conversion, and incremental lift. * Audience Construction: Build statistically sound test and control audiences in partnership with CRM, ensuring balance, mutual exclusivity, and reproducibility. * Rigorous Measurement: Execute the Test vs. Control measurement framework and apply guardrails for attribution, incrementality, and performance readouts. * Standardization: Apply and help maintain Standard Operating Procedures (SOPs) for customer campaign measurement, reporting hygiene, and data integrity. * Lifecycle Optimization: Support data-driven customer journeys (Onboarding, Growth, Retention, Reactivation) with the models and audiences that make them work. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)