> Markdown version of [/jobs/ext/2728362-advanced-data-scientist](https://www.wearedevelopers.com/jobs/ext/2728362-advanced-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Advanced Data Scientist - **Company:** Zebra Technologies - **Location:** Lincolnshire, IL, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Artificial Neural Networks, Microsoft Azure, Cloud Computing, Data Architecture, Data Validation, Information Engineering, Extract Transform Load (ETL), Data Transformation, Github, Python (Programming Language), Machine Learning, NoSQL, NumPy, Software Engineering, SQL Databases, Cloud Platform System, Data Ingestion, Prophet, Large Language Models, Snowflake, Git, Pandas, Pyspark, Information Technology, Production Code, Xgboost, Machine Learning Operations, Virtual Agents, Databricks - **Published:** September 5, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/phpyf50pp3 ## About the Role The Data Scientist Sr will have experience working across the full lifecycle of a Data Science project. The ideal candidate will demonstrate proficiency in the design, building, and optimization of data ingestion and data transformation pipelines; the design, enhancement and tuning of ML/AI models, the operationalization of resultant models; and communicating root cause analysis and model explainability to business stakeholders. The role will entail interfacing with various source systems, and proficiency in PySpark, SQL, Databricks, and related cloud resource management services. Prior experience with advanced analytics, ML/AI, and optimization solutions in the Retail/CPG domain is strongly preferred. A working knowledge of GenAI/LLMs and Agentic-AI in the context of workflow automation will be valuable but is not the primary requirement for this position., * Minimum Education: * Master's degree in engineering, computer science, data science, operations research, statistics, mathematics, quantitative sciences or relevant work experience * Minimum Work Experience (years): * 6+ years of experience in Data Science/Data Engineering with emphasis on the full lifecycle of Data Science-ML/AI projects. * Within that timeframe, experience is expected in: Python/PySpark, SQL, and relational or NoSQL databases, and cloud resource management. * Key Skills and Competencies: * Experience working with AWS, Azure, or GCP cloud environments. * Experience implementing advanced analytics, ML/AI algorithms (such as: , * Statistical Time Series: Exponential Smoothing Models, S/ARIMA; * Machine Learning: Random Forests, Gradient Boosting Methods; * Neural Networks: TiDE (Google), DenseNet & Prophet (FB-Meta); * Foundational Time Series Models: TimesFM (Google), Chronos (AWS) * Mathematical (constrained linear, non-linear and network) optimization models * Proven experience building end-to-end production grade Data & ML/AI pipelines using PySpark, Python (Pandas/NumPy) and SQL. * Experience working with Git (or similar code management repositories) as a collaboration tool. production-grade * Experience with orchestration tools like Databricks, Airflow (or similar tools like Snowflake, Dagster, etc.). * Understanding of Retail/CPG industry business challenges with an emphasis on Supply Chain, Pricing/Promotions/Markdowns, Inventory Allocation, Assortment Mix Planning, Size & Pack Optimization, Retail Shrink/Fraud & Anomaly Detection, and Workforce optimization applications are highly desirable. * Excellent verbal and written communication skills, especially as it relates to technical communications. Ability to present technical analysis to business stakeholders. * Demonstrated ability to learn new technologies quickly and independently, particularly as technology in this domain rapidly advances. In this context, a working knowledge of GenAI/LLMs, Agentic-AI and related frameworks (e.g. LangChain) will be a plus. * Ability to work independently with minimal supervision and achieve stretch goals in a highly innovative and a fast-paced environment. ## Description * Design, optimize, and maintain scalable ETL pipelines using PySpark and Databricks on cloud platforms (Azure/GCP). * Develop automated data validation process to proactively perform data quality checks. * Employ key Databricks modules (DeltaLiveTables, Unity Catalog, MLFlow) to facilitate creating, running experiments, automating, and scheduling jobs on Databricks. * Optimize the allocation of cloud resources and Databricks DBUs to manage and control cloud and Databricks consumption cost. * Employ GitHub repositories to ensure that production code management best practices are being strictly adhered to. * Build and tune ML/AI and optimization models, identify algorithmic performance improvement opportunities, and perform experiments to demonstrate incremental value delivered. * Have frequent conversations with Business Stakeholders to understand their requirements and concerns. Explain data deficiencies, model performance/root cause analysis, and model output. * Follow best practices in Data Architecture, Coding, and Project Management operations. * Collaborate with cross-functional teams, such as Customers' Stakeholders, Engagement Managers, Data Ops/Job Monitoring, Product Management & Software Engineering. * Expand the use of analytics, ML/AI, mathematical optimization, Gen-AI/LLMs and Agentic-AI in the context of Retail/CPG business use cases such as anomaly detection, demand forecasting, price elasticity modeling, promotions features & strategy simulation, product cannibalization and halo modeling, markdown optimization, product allocation, reorder/replenishment, size and pack optimization, workforce scheduling & task optimization ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)