> Markdown version of [/jobs/ext/2160615-machine-learning-engineers](https://www.wearedevelopers.com/jobs/ext/2160615-machine-learning-engineers). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineers - **Company:** jazzhr - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automated Storage and Retrieval Systems, Cloud Computing, Encodings, Data Deduplication, Intrusion Detection Systems, Python (Programming Language), Machine Learning, Tensorflow, Data Logging, Pytorch, Large Language Models, Snowflake, Multi-Agent Systems, Deep Learning, Kaggle, Git, Free and Open-Source Software, Machine Learning Operations - **Published:** August 21, 2026 - **Apply:** http://talentwwinc.applytojob.com/apply/jobs/details/qO5zKGT3hz ## About the Role * 5+ years shipping ML systems into production - and you can name the system, the metric before and after, and how you knew the model caused the change. * Depth in both classical ML and deep learning (PyTorch or TensorFlow) applied to live products, not notebooks and Kaggle sets. * Working fluency with LLMs in production - retrieval, evals, prompt and context engineering, and the judgment to recognize when an LLM is the wrong tool. * You already ship with agentic coding tools - Claude Code, Claude Design, or close equivalents - and can point to work you built with them. * Software engineering fundamentals strong enough to own your own deploys - Python, Git, cloud (we run AWS), containers, and the patience for genuinely messy, human-authored, self-reported data. THE NICE-TO-HAVES * Entity resolution, record linkage, or taxonomy design at scale * Ranking, recommendation, or two-tower retrieval systems * Sequence models on longitudinal or event-stream data * Embedding and vector retrieval systems in production * Experiment design, causal inference, or off-policy evaluation * Warehouse-native ML (dbt, Snowflake, or similar) * Labor market, HR tech, or people-data domain experience * Open-source contributions or publications ## Description We're growing our machine learning team. We're looking for Machine Learning Engineers who own products end to end - from the problem, to production, to the metric that proves it worked. This role exists because of how we build. A small product strategy team sets direction and priorities; engineers own the work end to end - discovery, design, build, ship, and the result. You'll have the autonomy of a founder inside your domain and the accountability that comes with it. That accountability includes the unglamorous half of ML. You own the experiment that doesn't pan out and the call to kill it, not just the launch. We'd rather you run four honest experiments and ship the one that works than ship four things that all look fine on a dashboard., Depending on area of focus: Canonical data and entity resolution * Canonical datasets for titles, companies, skills, and industries - the layer every application depends on. Content-addressed IDs, faceted taxonomies, alias graphs accumulated across tens of millions of rows. * Rules-based resolution pipelines with LLM escalation, where the accumulated alias graph is the durable asset and escalation volume should fall over time. * Nightly agent loops that adjudicate ambiguous entities, propose structural changes, and get gated by invariant checks and blast-radius limits before anything commits. * Job ingestion at scale: multi-source feeds, deduplication, freshness, and the indexing economics underneath. Retrieval, ranking, and matching * Job matching v2: two-tower retrieval with cross-encoder reranking, trained on outcome labels rather than clicks. Hard-negative mining, propensity weighting, impression-time logging. * Mobility embeddings learned from observed career sequences - the similarity a text encoder can't recover, where Claims Adjuster and Underwriting Assistant are substitutable despite sharing no vocabulary. * Pivot feasibility: given where someone is, what moves are realistic, what's missing, and which intermediate roles actually worked for peers. Applied LLMs and agents * Fine-tuning where it earns its cost - against outcome labels, not for tasks a well-prompted frontier model already handles. * Agentic systems in production with human approval gates: agents that analyze, propose changes as reviewable artifacts, and execute only after a human signs off. We have this pattern running against tens of millions of customer touchpoints a year and want to push it much further. * Continuous skills inference from work artifacts rather than static documents - a problem several of our enterprise customers are currently solving for themselves, badly. * New product surfaces where the right answer genuinely requires an LLM, and the discipline to notice when it doesn't. Across all of it * Evaluation infrastructure you'd defend in a design review: time-forward splits, calibration, offline-to-online agreement, and honest handling of feedback-loop degeneration and survivorship bias. * Building inside real constraints: GDPR, EU AI Act high-risk classification for employment AI, and client data commitments are design inputs here, not someone else's problem. ## Related Videos - [Machine learning 101: Where to begin?](https://www.wearedevelopers.com/videos/1014-machine-learning-101-where-to-begin) - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction)