Data Scientist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Job description
We are seeking a collaborative, hands-on Data Scientist to develop, enhance, and deploy machine-learning and advanced analytics solutions. This is a broad, generalist role supporting work across risk, personalization, and other business-focused analytical initiatives., * Enhance, maintain, validate, and monitor existing machine-learning models, representing approximately 80% of the role.
- Design and develop new predictive models and advanced analytics solutions for business needs, representing approximately 20% of the role.
- Perform exploratory data analysis to identify opportunities, assess data quality, and guide modeling approaches.
- Independently engineer training datasets, including target definition, sampling strategy, class-balance considerations, and temporal data integrity.
- Select and evaluate appropriate modeling approaches based on the business problem, available data, explainability needs, inference requirements, and operational complexity.
- Validate features and model performance beyond in-sample metrics by assessing stability across time periods and cohorts, inference-time availability, potential leakage, and real-world plausibility.
- Use time-aware validation approaches when working with temporal data and investigate distribution shift, seasonality, and performance degradation.
- Take models from exploratory development through production deployment using the team’s established, low-overhead deployment process.
- Refactor notebook-based work into maintainable production code with version control, testing, and appropriate deployment practices.
- Contribute to development and release processes across DEV, QA, and PROD environments.
- Partner closely with technical and business stakeholders to clarify requirements, gather feedback, and iterate on model solutions.
- Work effectively in a two-week sprint cadence, delivering usable improvements incrementally.
- Manage multiple workstreams and maintain context across concurrent projects and stakeholder conversations.
- Create visualizations or dashboards as needed to communicate analysis and model insights.
Requirements
The ideal candidate combines strong Python and SQL skills with practical judgment: they can build reliable training datasets, validate model behavior rigorously, and ship maintainable solutions without waiting for a model to be “perfect.”, * 3-5 years of professional experience in data science, machine learning, advanced analytics, or a related field.
- Strong proficiency in Python and SQL.
- Hands-on experience with Snowflake; this is strongly preferred.
- Experience developing and evaluating a range of model types, such as linear regression, classification models, collaborative filtering/recommendation approaches, ranking models, or other predictive techniques.
- Demonstrated ability to structure high-quality model training datasets and assess target sampling strategies.
- Strong understanding of:
- Feature engineering, feature validation, and feature importance.
- Data leakage and ensuring features are available at prediction time.
- Temporal train/validation/test splits.
- Distribution shift, cohort-based validation, and model generalization.
- Class imbalance, negative sampling, and evaluation methods aligned to real-world inference scenarios.
- Experience moving machine-learning models from Jupyter notebooks into production environments.
- Working knowledge of GitHub, version control, code review practices, testing, and environment separation across DEV, QA, and PROD.
- Experience working in VS Code or comparable modern development environments.
- Ability to interpret technical blueprints or specifications, ask effective clarifying questions, and independently execute against defined requirements.
- Strong communication skills and a collaborative, iterative working style.
Preferred Qualifications
- Experience with model monitoring, containerization, and production model lifecycle practices.
- Experience supporting analytics or machine-learning use cases in personalization, risk, behavioral prediction, or related domains.
- Experience building dashboards, visualizations, or self-service analytical tools.
- A GitHub portfolio or examples of production-quality data science work.
Skills: Analysis Skills, Blueprints, Business Analysis, Cadence, Code Reviews, Communication Skills, Computer Security, Data Analysis, Data Quality, Data Science, GitHub, Machine Learning, Metrics, Model Validation, Performance Modeling, Predictive Modeling, Production Systems, Python Programming/Scripting Language, Quality Assurance, Refactoring, Reporting Dashboards, Requirements Management, Risk, SQL (Structured Query Language), Source Code/Configuration Management (SCM), Stability Analysis, Team Player, Testing, Training Data Sets, Use Cases, Validation Testing
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Résumé-Driven Development: How IT trends affect the job market for software developers
Top Big Data Technologies That You Need to Know
Making Data Warehouses Fast: A Developer’s Story
Data Engineer Salary UK