Databricks Data Scientist
Role details
Job location
Tech stack
Job description
- Develop, train, and evaluate machine learning and statistical models to support business and mission needs using the Databricks platform.
- Prepare, clean, and maintain datasets for modeling, experimentation, and analysis.
- Write, optimize, and maintain Python and SQL workflows for data exploration, feature engineering, and model development.
- Work with large-scale datasets using Databricks, Spark, and Delta Lake platforms.
- Design reusable feature engineering workflows and model training pipelines using Databricks notebooks, workflows, and MLflow.
- Register, version, promote, and document models using MLflow Model Registry and Unity Catalog-based model governance practices.
- Monitor deployed models for performance, drift, data quality, usage patterns, and operational issues; recommend retraining, tuning, or retirement actions as needed.
- Analyze data to identify trends, patterns, and insights to support business decisions.
- Translate business requirements into analytical approaches, models, and data science solutions.
- Perform data validation, quality checks, and issue resolution to ensure accuracy and consistency.
- Collaborate with cross-functional teams including data engineers, analysts, and business stakeholders.
- Communicate model outputs, analytical findings, and recommendations to both technical and non-technical audiences.
- Document models, datasets, and methodologies to support reproducibility, transparency, and reuse.
- Follow data governance, security, and compliance standards within the platform.
Requirements
Guidehouse is seeking a Databricks Data Scientist to join our AI & Data team to support client projects involving advanced analytics, machine learning, and data science solutions. This role focuses on working with data to develop models, generate insights, and support data-driven decision-making across teams. The role requires strong Python and SQL skills, analytical thinking, and the ability to collaborate with clients and stakeholders to deliver scalable data science solutions., * Bachelor's degree in computer science, engineering, mathematics, statistics, or another relevant field.
- 3-8 years of relevant experience in data science, machine learning, or advanced analytics.
- Strong experience with Python and SQL for data analysis, modeling, and transformation.
- Experience with Databricks, Spark, Delta Lake, or similar cloud-native data platforms.
- Hands-on experience designing, building, evaluating, and deploying machine learning models, including experience moving models from prototype to production or production-like environments.
- Experience with ML lifecycle practices including experiment tracking, model evaluation, model registry, version control, deployment workflows, monitoring, and retraining approaches.
- Familiarity with model serving patterns, API-based inference, scheduled batch scoring, and integration of model outputs into dashboards, applications, or operational workflows.
- Experience with data preparation, feature engineering, and model development.
- Ability to analyze data and communicate insights clearly.
- Ability to troubleshoot technical issues, communicate recommendations clearly, and work effectively in team-based delivery environments.
- Experience supporting AI governance practices, including model documentation, validation, monitoring, version control, and responsible AI considerations.
- Ability to work across data science, data engineering, cloud, security, and client stakeholder teams to translate analytical prototypes into scalable, maintainable solutions.
What Would Be Nice to Have:
- 2+ years of hands-on experience with the Databricks platform.
- Active Databricks Machine Learning Engineer, GenAI Engineer, Data Analyst, or related certification.
- Experience with Databricks MLflow, Feature Engineering, Feature Store, Model Serving, Workflows, Unity Catalog, Mosaic AI, Vector Search, AI Gateway, or related Databricks AI/ML capabilities.
- Experience developing GenAI, LLM, RAG, agentic AI, or prompt evaluation workflows using Databricks Mosaic AI, MLflow, open-source frameworks, or cloud-native AI services.
- Experience with CI/CD, automated testing, code packaging, environment promotion, and source control practices for data science and machine learning workloads.
- Experience with machine learning frameworks and statistical modeling techniques.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Experience working in project-based or consulting delivery environments.
- Familiarity with data modeling, data warehousing, and large-scale data processing concepts.
Benefits & conditions
The annual salary range for this position is $113,000.00-$188,000.00. Compensation decisions depend on a wide range of factors, including but not limited to skill sets, experience and training, security clearances, licensure and certifications, and other business and organizational needs.
What We Offer:
Guidehouse offers a comprehensive, total rewards package that includes competitive compensation and a flexible benefits package that reflects our commitment to creating a diverse and supportive workplace.
About Guidehouse
Guidehouse is an Equal Opportunity Employer-Protected Veterans, Individuals with Disabilities or any other basis protected by law, ordinance, or regulation.
Guidehouse will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of applicable law or ordinance including the Fair Chance Ordinance of Los Angeles and San Francisco.