> Markdown version of [/jobs/ext/2702702-machine-learning-scientist](https://www.wearedevelopers.com/jobs/ext/2702702-machine-learning-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Scientist - **Company:** Spotter, Inc - **Location:** Culver City, United States - **Experience:** Expert - **Salary:** $167,000.0 - $185,000.0 - **Contract:** Franchise - **Skills:** A/B Testing, Artificial Intelligence, Data Analysis, Artificial Neural Networks, Python (Programming Language), Machine Learning, Recommender Systems, Standard Sql, SQL Databases, Web Platforms, Reinforcement Learning, Data Logging, Deep Learning, Model Validation, Information Technology, Low Latency, Machine Learning Operations - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/machine-learning-scientist-spotter-8287139 ## About the Role * Master's degree or PhD in Computer Science, Statistics, Applied Mathematics, Electrical Engineering, Physics, or another quantitative field. * 5+ years building, evaluating, and deploying machine learning models in production environments. * Experience with reinforcement learning or contextual bandit systems gained through graduate coursework, academic research, or hands-on industry experience. Candidates with experience building and deploying these systems in production, from problem formulation through offline evaluation to live deployment, are strongly preferred. * Solid grasp of core RL training objectives and loss functions, including temporal-difference and Bellman error losses (Q-learning, DQN), policy gradient objectives (REINFORCE, actor-critic advantage estimation), and clipped surrogate objectives (PPO, TRPO), with an understanding of when each applies and how they behave in training. * Practical experience with bandit and reinforcement learning methods such as Thompson sampling, UCB or LinUCB, neural bandits, non-stationary bandits, policy gradients, actor-critic methods, or Q-learning. * Ability to design reward functions and objective trade-offs for systems optimizing long-horizon outcomes, including diagnosing and mitigating reward hacking and feedback loops. * Knowledge of off-policy and counterfactual evaluation, such as inverse propensity scoring (IPS), self-normalized IPS, doubly robust estimators, and replay evaluation, and with counterfactual learning from logged bandit feedback, including propensity logging. * Experience working with logged interaction data, behavioral data, or feedback signals to train, evaluate, and improve models. * Track record of designing experiments and using data to improve model performance in real-world product environments, including A/B testing and causal inference. * Strong experience with modern deep learning frameworks and production ML workflows. * Expertise in training, evaluating, tuning, and deploying machine learning models across deep learning and traditional ML approaches. * Strong understanding of embeddings, representation learning, neural networks, sequence modeling, and modern deep learning architectures. * Strong Python and SQL skills. * Excellent communication skills and the ability to work cross-functionally with Product, Engineering, Analytics, and other stakeholders. * Curiosity, ownership, and a passion for building products that customers love. Nice to Have * Hands-on work building large-scale recommendation, ranking, or personalization systems. * Understanding of offline reinforcement learning methods, such as CQL or IQL, for training policies from logged data. * Knowledge of constrained or safe reinforcement learning and guardrailed deployment, including offline evaluation gates ahead of live A/B tests. * Familiarity with ad recommendation, ad ranking, or campaign optimization systems used by large-scale platforms, such as YouTube, Google, Meta, TikTok, Amazon, or similar consumer marketplace platforms. * Experience serving large-scale ML models in production. * Background building machine learning systems for large-scale digital platforms, such as Creator platforms, consumer apps, recommendation systems, ad recommendation systems, campaign optimization systems, or workflow automation tools. ## Description You'll develop machine learning models that move beyond experimentation and into production, where they directly improve Creator workflows and product experiences. Working alongside Analytics, Product, and Engineering, you'll help develop intelligent systems that improve how Creators discover insights, make decisions, and create content. Your work may include: * Designing, training, evaluating, optimizing, and deploying production reinforcement learning, contextual bandit, and online learning systems that improve product outcomes. * Creating systems that balance exploration and exploitation, short-term performance and long-term value, and multiple competing product objectives. * Developing reward models, feedback models, and objective functions that translate noisy, sparse, delayed, or implicit signals into reliable model training and evaluation targets, and diagnosing and mitigating reward hacking and feedback loops in deployed systems. * Applying offline policy evaluation and counterfactual techniques, such as inverse propensity scoring, doubly robust estimation, and replay evaluation, to reason about model changes before and after deployment. * Working with logged interaction data to understand user behavior, evaluate model performance, improve decision quality, and reduce bias in model evaluation. * Designing experiments to evaluate model performance, measure product impact, and continuously improve production systems. * Building scalable model training, evaluation, deployment, and inference pipelines. * Optimizing models for accuracy, latency, scalability, reliability, and production maintainability. * Working with structured and unstructured datasets using Python and SQL. * Collaborating closely with Product and Engineering to translate customer problems into machine learning solutions. * Staying current with advances in reinforcement learning, bandits, recommendation systems, ranking, personalization, deep learning, experimentation, and production ML, and thoughtfully applying new techniques where they create measurable value. ## Related Videos - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [How We Built a Machine Learning-Based Recommendation System (And Survived to Tell the Tale)](https://www.wearedevelopers.com/videos/752-how-we-built-a-machine-learning-based-recommendation-system-and-survived-to-tell-the-tale) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Beyond Autocomplete: Local AI Code Completion Demystified](https://www.wearedevelopers.com/videos/961-beyond-autocomplete-local-ai-code-completion-demystified) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)