Forecasting Lead Data Scientist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
The candidate must be able to name the industry and the outcome variable for each such engagement. Retail same-store analysis is the classic form; the analogue here is comparing similar schools and events rather than following one trend line.
-
Presents to non-statisticians: business outcome first, method second; confidence stated in plain language; explicitly states what the forecast cannot do; never opens with an undefined statistical term.
-
Can teach the method to a client team, not only execute it.
-
Participate actively in stand-ups and backlog refinement, engage business stakeholders directly, understand why the business is asking a question, and challenge or refine the request when it is wrong.
-
Strategic recommendations are expected alongside hands-on delivery.
Requirements
Required:
-
Must be able to work EST hours
-
8+ years of applied forecasting.
-
Two or more comparable forecasting engagements led start to finish.
-
Comparable-unit / same-store forecasting experience.
-
Executive communication.
-
Thought leadership.
-
Multivariable regression, plus collinearity analysis and VIF interpretation.
-
Forecast model development, tuning, selection and holdout validation.
-
Metric fluency: R², WAPE, MAPE, p-values - and why WAPE is used at event grain (many events sell zero, which breaks MAPE).
-
Sparse and zero-inflated data. Many variables populate on under 25% of events, some as low as 10%. Nulls must never be silently treated as zeros.
-
Data-leakage discipline and point-in-time correctness: every feature must exist before the event starts.
-
Python and SQL; reproducible notebooks.
-
Snowflake, including Snowflake ML Model Registry (model versions carry metrics and training-dataset references).
-
Git and pull-request workflow; all code merged to the client repository, no private forks.
Preferred:
-
Architecture Decision Records (ADRs) and written process documentation.
-
Categorical encoding at scale (~30-35 source variables expand to ~70 columns).
-
Sports, streaming, ticketing or subscription-business domain exposure.
-
Hierarchical or mixed-effects models for low-volume segments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Making Data Warehouses Fast: A Developer’s Story
Data Engineer Salary UK
The Prompt Engineer ✍️