> Markdown version of [/jobs/ext/3431422-research-data-scientist](https://www.wearedevelopers.com/jobs/ext/3431422-research-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Data Scientist - **Company:** Innodata Inc - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $160,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Data Analysis, Microsoft Azure, Information Engineering, Data Files, Data Governance, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Natural Language Processing, NumPy, Tensorflow, SQL Databases, Unstructured Data, Feature Engineering, Pytorch, Large Language Models, Prompt Engineering, Model Validation, Generative AI, Git, Pandas, Scikit Learn, Information Technology, Free and Open-Source Software - **Published:** September 16, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6a5704c18b25e948 ## About the Role We are looking for a highly skilled Research Data Scientist - GenAI/LLM to join our AI/LLM Delivery Unit and work on research-driven AI/ML initiatives involving Generative AI, Large Language Models (LLMs), NLP, multimodal AI, model evaluation, and AI data. The role combines strong research and analytical capabilities with hands-on AI/ML expertise, requiring the candidate to design experiments, develop evaluation methodologies, analyze complex datasets, build research prototypes, and translate research findings into practical AI/ML solutions. The ideal candidate will have a strong research orientation, excellent statistical and analytical skills, and the ability to work collaboratively with researchers, data scientists, AI/ML engineers, domain experts, and client-facing teams., * Master's or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or a related discipline. * Bachelor's/Master's degree from IITs, NITs, or other premier engineering/research institutions is strongly preferred. * 4-7 years of hands-on research experience in AI/ML, Data Science, NLP, Generative AI, LLMs, or related areas. * Strong demonstrated research experience with the ability to independently formulate research questions, design experiments, analyze results, and communicate findings. * Demonstrated research track record through research publications, patents, conference presentations, open-source contributions, or significant AI/ML research projects. * Candidates with publications in reputed conferences/journals and a strong academic/research profile will be preferred * Strong proficiency in Python and SQL. * Strong hands-on experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch/TensorFlow. * Strong understanding of: + Machine learning algorithms + Statistics and experimentation + Data analysis and feature engineering + Model evaluation and performance metrics + Hypothesis testing and statistical inference * Hands-on exposure to LLMs, NLP, Generative AI, and multimodal AI. * Experience with one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF/DPO, embeddings, or model benchmarking. * Experience working with large-scale structured and unstructured datasets. * Familiarity with Git and cloud platforms such as AWS, Azure, or GCP is desirable. ## Description * Conduct independent and collaborative research in Generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation, and AI data. * Formulate research questions and translate complex AI/ML problems into structured research methodologies and experiments. * Design, execute, and analyze experiments to evaluate and improve AI/ML models and solutions. * Build analytical models, prototypes, and research pipelines using Python and relevant ML frameworks. * Stay current with emerging research, methodologies, papers, and developments in GenAI, LLMs, NLP, multimodal models, and AI evaluation. LLM & Model Evaluation: * Develop and implement LLM evaluation frameworks, benchmarks, datasets, and evaluation criteria. * Evaluate models for accuracy, robustness, bias, hallucination, reasoning, relevance, response quality, and other performance dimensions. * Conduct model benchmarking, error analysis, comparative analysis, and performance evaluation. * Work on areas such as RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and LLM optimization, as applicable. * Identify model and data gaps and recommend improvements to enhance model performance and reliability. Data Science & Statistical Research: * Collect, clean, analyze, and interpret large and complex structured and unstructured datasets. * Perform EDA, statistical analysis, hypothesis testing, significance testing, correlation analysis, sampling, and error analysis. * Develop data-driven insights and identify patterns, trends, and relationships relevant to AI/ML research. * Apply appropriate statistical and quantitative methodologies to validate research findings AI Data & Dataset Development: * Develop and evaluate datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks, and evaluation criteria for AI/ML models. * Analyze data quality and identify issues affecting model performance. * Collaborate with annotation, data engineering, and AI/ML teams to improve AI training and evaluation data. * Translate data and research findings into actionable recommendations for improving AI system Research & Innovation: * Contribute to research papers, technical reports, whitepapers, patents, benchmarks, internal publications, and other research outputs, where applicable. * Identify opportunities to apply emerging research and technologies to real-world AI and data challenges. * Explore new methodologies, models, datasets, and evaluation approaches to improve AI capabilities. * Contribute to capability building and innovation within the AI/LLM practice Collaboration & Stakeholder Engagement: * Work closely with researchers, data scientists, AI/ML engineers, data/annotation teams, domain experts, and delivery teams. * Present research findings, analytical insights, and technical recommendations to senior technical stakeholders. * Translate complex research and technical concepts into clear, actionable recommendations. * Where required, participate in client-facing technical discussions and presentations and help translate business requirements into AI/ML solutions. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)