> Markdown version of [/jobs/ext/2072810-data-scientist](https://www.wearedevelopers.com/jobs/ext/2072810-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Applaudo Studios SA de CV - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Aliasing, Artificial Neural Networks, Big Data, BigQuery, Data Deduplication, Python (Programming Language), Machine Learning, Tensorflow, Standard Sql, Pytorch, Large Language Models, Snowflake, Apache Spark, Model Validation, Unsupervised Learning, Databricks - **Published:** August 15, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/n4b8nnli0e ## About the Role You are an experienced Data Scientist with strong applied Machine Learning expertise and a track record of building and evaluating models using real-world, messy, large-scale data. You are comfortable working with embeddings, semantic similarity, LLMs, NLP, classification, and both supervised and unsupervised learning. You approach ambiguous problems through structured experimentation, clearly defined hypotheses, baselines, metrics, and error analysis. You are highly autonomous, intellectually honest about experimental results, and able to clearly communicate technical recommendations and trade-offs to engineering and business stakeholders. You Bring to Applaudo the Following Competencies * 5+ years of professional Data Science / Machine Learning experience. * Strong applied Machine Learning fundamentals. * Excellent Python and SQL skills. * Hands-on experience with embeddings and semantic similarity. * Practical experience applying LLMs to real-world problems. * Experience with supervised and unsupervised learning. * Strong experience with classification and NLP. * Working knowledge of neural networks and transformer architectures. * Hands-on experience with TensorFlow, PyTorch, PyCaret, or equivalent ML frameworks. * Experience retraining or maintaining classification models in production. * Strong experimental design and model evaluation skills. * Experience defining baselines, metrics, test sets, and error-analysis processes. * Ability to evaluate model quality and demonstrate measurable improvements. * Strong understanding of scalability and ML inference costs. * Strong English communication skills. Nice-to-Have * Entity resolution, record linkage, or deduplication experience. * Ranking and similarity scoring. * Retrieval, clustering, or candidate-generation techniques. * LLM/embedding solutions designed for cost and scale constraints. * Spark, Snowflake, Databricks, or BigQuery. * Experience with company, domain, website, or firmographic data. * Experience working with multilingual datasets., * Strong analytical and experimental mindset. * Intellectual honesty and willingness to communicate negative results. * Strong autonomy and self-direction. * Excellent written and verbal communication. * Ability to defend technical recommendations with stakeholders. * Strong problem-solving skills. * Comfort working with ambiguity and large-scale datasets. * Ability to balance model quality, cost, and scalability. ## Description * Build and evaluate ML approaches for company/entity matching. * Develop embedding and LLM-based matching approaches. * Develop scoring and ranking methodologies to identify true matches and distinguish them from duplicates, lookalikes, and unrelated entities. * Work with messy data, including names, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies. * Define benchmark datasets, metrics, baselines, and error-analysis processes. * Design and execute experiments to validate hypotheses. * Compare LLM-assisted approaches against lower-cost alternatives. * Analyze model behavior, edge cases, and trade-offs. * Consider inference economics and scalability from the beginning. * Communicate experimental findings and recommendations to engineering and business stakeholders. * Independently establish experimental pipelines and research approaches. * Clearly document both successful and unsuccessful experiments. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Pointers? In My Python? It's More Likely Than You Think](https://www.wearedevelopers.com/videos/358-pointers-in-my-python-it-s-more-likely-than-you-think) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career)