Principal Data Scientist

Ford Motor Company
Dearborn, MI, United States
14 days ago
Apply on jobs.mitalent.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Big Data BigQuery Data Cleansing Information Engineering Data Infrastructure Data Integrity Extract Transform Load (ETL) Data Structures Data Systems Data Warehousing Monitoring of Systems Python (Programming Language)
+28 more
Machine Learning Meta-Data Management Multiprocessing Neo4j NLTK (NLP Analysis) NumPy Query Optimization Cloud Services Standard Sql SciPy Software Deployment Data Streaming Data Processing Scripting Data Storage Technologies Large Language Models Generative AI Pandas Matplotlib Data Lakes Pyspark Scikit Learn Information Technology Data Lineage Data Analytics Data Management Data Pipelines Docker

Job description

Principal Data Scientist - positions offered by Ford Motor Company (Dearborn, Michigan). Note, this is a hybrid position whereby the employee will work both from home and from the aforementioned worksite. Hence, the employee must live within a reasonable commuting distance from the worksite. Design, build, and deploy end-to-end machine learning algorithms to tackle complex data challenges unique to platform. Developing identity resolution algorithms from scratch, refining for accuracy and performance at scale, and supporting production deployment. Lead efforts to identify, analyze, and resolve complex data quality issues across large-scale data warehouses and data lakes. Design and implement ML-driven monitoring tools for anomaly detection, data cleansing, standardization, and validation, ensuring high data integrity, reliability, and consistency from all sources. Contribute to developing internal tools, libraries, and automation scripts focused on enhancing data lineage, metadata management, and automated data quality checks. Building and managing Docker images and automating algorithms for efficient orchestration and scheduling. Apply advanced data science (statistical analysis, predictive modeling) to understand the performance, efficiency, and cost-effectiveness of large-scale data pipelines, ETL/ELT processes, and data storage solutions. Propose and implement data-driven improvements to optimize data flow, achieve sub-second latency, and enhance resource utilization. Use advanced analytical techniques for deep-dive root cause analysis on data discrepancies, performance bottlenecks, or system failures. Drive the strategic development of data science capabilities within the Product group, leading R&D efforts. Leverage Generative AI (LLMs) to identify problems and develop solutions within the data platform and data warehouse environment. Partner closely with data engineers and architects to design, optimize, and maintain scalable data structures and cloud-native data services. Manage the ingestion of data from various sources and formats into BigQuery and other data platforms. Implement robust transformation processes necessary for downstream analytical consumption. Act as a key liaison and project leader, mentoring data engineers and collaborating with global stakeholders. Translate complex business requirements into precise technical specifications for data solutions. Serve as a technical conduit between central privacy, product, and engineering teams for end-to-end system design. Create comprehensive documentation for data structures, data quality rules, and analytical findings related to the data platform. Share expertise, mentor junior team members, and foster best practices within the team and across the organization.

Requirements

Ph.D. in Computer Science, Statistics, Mathematics or a related field and 5 years of experience in the job offered or a related occupation. 5 years of experience with each of the following skills is required: 1. Data Science or advanced data engineering. 2. Python development (including Pandas, NumPy, and Scikit-learn) for data manipulation, analysis, and scripting. 3. Using Machine Learning frameworks and libraries including at least 5 of the following: Pyspark, BigQuery ML, Pandas, Scipy-Weave, Multiprocessing, Graph ML libraries, Neo4j, NLTK, or MatplotLib. 4. Designing, implementing, and deploying machine learning models for complex data problems including anomaly detection or predictive maintenance within data systems, including algorithm fine-tuning for performance and accuracy at scale. 5. Researching, evaluating, and integrating new ML capabilities, covering both foundational and cutting-edge approaches. 3 years of experience with the following skill is required: 1. Using SQL skills for querying, manipulating, and optimizing complex datasets, including query tuning., Ph.D. in Computer Science, Statistics, Mathematics or a related field and 5 years of experience in the job offered or a related occupation. 5 years of experience with each of the following skills is required: 1. Data Science or advanced data engineering. 2. Python development (including Pandas, NumPy, and Scikit-learn) for data manipulation, analysis, and scripting. 3. Using Machine Learning frameworks and libraries including at least 5 of the following: Pyspark, BigQuery ML, Pandas, Scipy-Weave, Multiprocessing, Graph ML libraries, Neo4j, NLTK, or MatplotLib. 4. Designing, implementing, and deploying machine learning models for complex data problems including anomaly detection or predictive maintenance within data systems, including algorithm fine-tuning for performance and accuracy at scale. 5. Researching, evaluating, and integrating new ML capabilities, covering both foundational and cutting-edge approaches. 3 years of experience with the following skill is required: 1. Using SQL skills for querying, manipulating, and optimizing complex datasets, including query tuning.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.mitalent.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:54 min

Development history of scientific computation libraries and PyViz tools

Radovan Kavický · LIVE

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:30 min

Introduction to Neo4j and remote developer relations work

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

Videos

See all

Related articles

See all