> Markdown version of [/jobs/ext/2822028-ml-engineer-tabular-data-experimentation](https://www.wearedevelopers.com/jobs/ext/2822028-ml-engineer-tabular-data-experimentation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Engineer - Tabular Data & Experimentation - **Company:** TANGO ASSOC INC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** A/B Testing, Python (Programming Language), Recommender Systems, SQL Databases, Feature Engineering, Large Language Models, Kaggle, Power Analysis (Cryptography), Xgboost, Machine Learning Operations - **Published:** September 10, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pj26fg9fkd ## About the Role * 5+ years in applied ML with real product impact, not Kaggle-only. * Deep working knowledge of tabular and panel data: temporal leakage, non-stationarity, the ways these problems differ from i.i.d. * Solid statistics and experimentation - comfortable designing a test from scratch, and equally comfortable pushing back when someone wants to "just look at the p-value." * Causal inference at the depth where you can pick a method, justify it, and explain what it doesn't tell you. * Production-grade Python - maintainable code, not just research notebooks. * SQL at production-analytics level; able to get to the right sample independently, including non-trivial joins, window functions, and reasoning about query plans. * A clear view of when an LLM is the right tool for a problem and when a GBDT or classical method is, and the ability to defend either choice with evaluations and cost. ## Description The core of this role is building ML on tabular and panel data, and designing the experiments that tell you whether it works. Most of the interesting questions are about experimentation: how to design an A/B that's worth running, how to extract a credible answer from offline data when an online experiment isn't feasible, and how to know when offline analysis is enough and when it isn't. This applies whether the underlying model is a gradient boosting model, a classical recommender, or an LLM-based component. LLMs are one tool among several here - chosen on merits against the alternatives, and held to the same experimental rigor as anything else the team ships. * End-to-end ML on tabular and panel data: feature engineering, validation strategy including time-aware splits, gradient boosting and related methods, calibration, drift monitoring. * Designing A/B tests that answer real questions: hypothesis design, power and MDE, handling peeking and multiple testing, variance reduction (CUPED and similar), interference and network effects. * Extracting credible answers from offline data when online experimentation isn't feasible - DiD, synthetic control, IV, uplift modeling - and choosing the right method for the situation rather than the most fashionable one. * Knowing when offline evaluation is sufficient and when it isn't, and being willing to defend that judgment. * Designing evaluations for LLM-based components the team owns: offline metrics, online proxies, drift detection, power analysis, guardrail metrics - with the same rigor as classical ML. * Translating business questions into ML formulations: metrics, loss, constraints, the trade-offs that actually matter to the product.