> Markdown version of [/jobs/ext/2113129-data-scientist-programmatic-algorithms](https://www.wearedevelopers.com/jobs/ext/2113129-data-scientist-programmatic-algorithms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist, Programmatic Algorithms - **Company:** Impact - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $165,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Big Data, BigQuery, Cloud Computing, Data Architecture, Data Infrastructure, Software Debugging, Distributed Systems, Data Flow Control, Monitoring of Systems, Python (Programming Language), Machine Learning, Automation of Marketing, Tensorflow, Azure Machine Learning, SQL Databases, Management of Software Versions, Reinforcement Learning, Feature Engineering, Pytorch, Apache Spark, Yield Optimization, Scikit Learn, Low Latency, Xgboost, Ripple (payment Protocol), Apache Kafka, Machine Learning Operations, Restful APIs, Marketplace, Software Version Control, Data Pipelines, Databricks - **Published:** August 19, 2026 - **Apply:** https://job-boards.greenhouse.io/impact/jobs/8427973002 ## About the Role * Experience: 5+ years in data science, ML engineering, or quantitative research, with at least 2+ years building and deploying ML models in programmatic advertising, ad tech, marketplace optimization, or a closely related domain (e.g., real-time bidding, dynamic pricing, auction systems). * Programmatic & marketplace depth: Demonstrated understanding of programmatic auction mechanics (RTB, header bidding, floor pricing, deal types, bid shading) and how ML can be applied to optimize outcomes across the supply-demand stack. * Production ML engineering: Proven ability to take models from prototype to production independently - including real-time inference, monitoring, retraining pipelines, and SLO ownership. * Data architecture: Experience designing and building data pipelines, feature stores, and training infrastructure for high-volume, low-latency ML systems. * Technical skills: * Strong Python and SQL; proficiency with ML libraries (scikit-learn, XGBoost, LightGBM, PyTorch/TensorFlow) and large-scale data tools (Spark, Kafka, or equivalent streaming/batch frameworks). * Experience with real-time feature serving and low-latency model deployment (REST APIs, gRPC, or streaming inference). * Familiarity with production ML workflows: model versioning, drift monitoring, A/B testing, evaluation, and retraining. * Experience processing and modeling at programmatic data scale: high-cardinality auction logs, bid stream data, impression and click events. Experimentation rigor: Strong grasp of causal inference and experiment design in online, delayed-feedback environments (auction holdouts, switchback tests, variance reduction techniques). Communication: Ability to explain complex modeling decisions and tradeoffs to Product and business stakeholders; comfortable presenting in cross-functional forums. Education: Bachelor's in a quantitative field (CS, Statistics, Math, Engineering, Economics, or similar); Master's/PhD preferred. Preferred / Nice to Have * Direct experience with SSP, DSP, or exchange-side yield optimization - particularly floor price optimization, bid landscape modeling, or deal matching algorithms. * Familiarity with auction theory (first-price vs. second-price dynamics, optimal reserve pricing, revenue equivalence) and its practical implications for programmatic ML. * Experience with contextual bandits, multi-armed bandits, or reinforcement learning applied to real-time decisioning problems. * Knowledge of online learning and adaptive algorithms in production environments with non-stationary data distributions. * Familiarity with privacy-preserving ML techniques relevant to programmatic (differential privacy, federated learning, cookieless attribution modeling). * Experience with GCP tools (BigQuery, Vertex AI, Dataflow, Pub/Sub) and/or Databricks/Spark for large-scale event processing and model training. * Exposure to supply forecasting, inventory management, or capacity planning in programmatic or marketplace contexts. * Familiarity with Impact's affiliate and partnership ecosystem, or prior experience at the intersection of performance marketing and programmatic delivery., How many years of experience do you have specifically building models for Programmatic Advertising, Real-Time Bidding (RTB), or Marketplace Optimization? * ## Description We're seeking a Senior Data Scientist to serve as an embedded Data Scientist within our Programmatic Experience Group. You'll own the design and deployment of machine learning models that optimize yield, pricing, and inventory allocation at scale - sitting at the intersection of data science, platform engineering, and marketplace economics. This is a high-craft, high-ownership individual contributor role. You'll work end-to-end: architecting the data pipelines that feed your models, engineering the features that drive performance, and deploying real-time inference systems that make decisions at speed. Your work directly determines how effectively Impact's programmatic marketplace balances advertiser performance with publisher monetization - making this one of the highest-leverage technical roles in the business. You'll collaborate closely with Product, Data Science, and Programmatic Delivery Engine Engineering, but you operate with significant autonomy. You're expected to bring both the modeling rigor of a data scientist and the production instincts of an ML engineer - and to be genuinely excited about both. What You'll Do: Yield Optimization & Pricing Models * Design and deploy ML models that optimize auction pricing, bid shading, floor price setting, and yield across Impact's programmatic inventory. * Build and iterate on real-time pricing algorithms that balance short-term revenue efficiency with long-term publisher and advertiser health. * Develop and maintain feedback loops that allow pricing models to adapt to shifting market conditions, inventory mix, and demand patterns. * Quantify the revenue impact of pricing model improvements; communicate tradeoffs between yield maximization, fill rate, and partner ROI to stakeholders. Inventory Allocation & Supply Optimization * Own ML-driven inventory allocation logic: routing, pacing, and matching supply to demand across partner segments, deal types, and campaign objectives. * Build models that forecast inventory availability, demand curves, and clearing prices to support proactive allocation decisions. * Identify and address inefficiencies in inventory utilization - including unsold inventory, suboptimal deal matching, and allocation imbalances across the publisher base. Data Architecture & Feature Engineering * Design and own the data infrastructure that feeds programmatic models: event pipelines, feature stores, training datasets, and real-time feature serving. * Engineer high-signal features from auction logs, bid stream data, user signals, contextual attributes, and historical performance - at the scale of programmatic data volumes. * Build robust data pipelines with production-grade standards: reliability, observability, versioning, and efficient reprocessing. Real-Time Inference & Production ML * Deploy models to production real-time inference environments; own latency, reliability, and throughput requirements for auction-time decision-making. * Build monitoring systems that track model performance, data drift, and system health in production; define alerting thresholds and retraining triggers. * Partner with MLOps and Platform Engineering to ensure scalable, low-latency serving infrastructure meets SLOs under high-volume auction traffic. * Own the full model lifecycle: training, evaluation, deployment, A/B testing, and iteration. Experimentation & Performance Measurement * Design and execute rigorous A/B and holdout experiments to measure the causal impact of model changes on yield, fill rate, advertiser performance, and publisher revenue. * Build evaluation frameworks that go beyond offline metrics - validating model behavior in live auction environments where feedback signals are delayed or noisy. * Translate experimental results into clear business narratives; present findings and recommendations to Product and business stakeholders. Self-Learning Systems & Feedback Loops * Research and implement adaptive, self-learning components within the programmatic stack - including contextual bandits, reinforcement learning signals, and online learning approaches where appropriate. * Design feedback mechanisms that close the loop between auction outcomes, model updates, and system behavior; reduce reliance on manual tuning and rule-based overrides. * Stay current with advances in programmatic ML, auction theory, and online optimization; evaluate applicability to Impact's specific marketplace dynamics. Cross-Functional Collaboration * Serve as the primary ML technical partner for the Rubicon product and engineering teams; translate business requirements into modeling approaches and communicate technical tradeoffs clearly. * Collaborate with Data Science peers on shared infrastructure, modeling standards, and cross-domain feature reuse. * Document models, architectures, and experimental findings to a standard that enables review, replication, and knowledge transfer across teams., * Marketplace intuition. You understand programmatic auctions not just as an engineer but as an economist - you think about incentive structures, equilibrium dynamics, and how model decisions ripple through the supply-demand stack. * Full-stack ML ownership. You're as comfortable designing a feature store schema as you are tuning a gradient boosting model or debugging a latency spike in production. You own the whole chain. * Feedback loop thinking. You don't just deploy models - you design the systems that make them smarter over time. You think about how today's decisions become tomorrow's training signal. * Rigor under real-world constraints. You know how to run clean experiments in environments where feedback is delayed, data is noisy, and business pressures create tradeoffs. You don't let imperfect conditions become an excuse for imprecise thinking. * Pragmatic delivery. You ship. You balance the perfect with the production-ready, iterate fast, and know when an MVP outperforms a six-month research project. * Collaborative depth. You build genuine technical trust with engineering and product partners - not just by having good ideas, but by following through, communicating clearly, and making the integration easy., Are you comfortable working in a role that requires "Full-Stack" ownership, including building your own data pipelines and feature engineering (not just modeling)?* Select... Have you managed the deployment of machine learning models into real-time inference environments where you were responsible for meeting latency and throughput SLOs?* Select... Which of the following have you worked with at a high-volume scale (e.g., processing millions of events per day)? (Select all that apply) * Auction/Bid Logs Real-time Feature Stores Distributed computing (Spark/Kafka/Dataflow) None of the above Which of these tools are you proficient in using for production-grade ML workflows? (Select all that apply) * Python (Scikit-learn, XGBoost, or LightGBM) SQL (Complex joins & window functions) Cloud ML Platforms (GCP Vertex AI, Databricks, or similar) A/B Testing & Causal Inference frameworks Will you now or in the future require employer sponsorship for employment visa status (e.g., H-1B)?* Select... What is your salary expectation?* Voluntary Self-Identification For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file. As set forth in Impact.com's Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Gender Select... Are you Hispanic/Latino? Select... Race & Ethnicity Definitions If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows: A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability. A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service. An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense. An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985. Veteran Status Select... ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [TikTok's Privacy Innovation](https://www.wearedevelopers.com/videos/1036-tiktok-s-privacy-innovation) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career)