> Markdown version of [/jobs/ext/2706494-technical-solutions-architect-for-evals-fine-tuning](https://www.wearedevelopers.com/jobs/ext/2706494-technical-solutions-architect-for-evals-fine-tuning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Solutions Architect for Evals & Fine-Tuning - **Company:** Innodata Inc - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $140,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computational Linguistics, Python (Programming Language), Machine Learning, Pytorch, Large Language Models, Information Technology, HuggingFace, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/technical-solutions-architect-evals-fine-tuning-innodata-inc-company-8705164 ## About the Role * 7+ years of experience in applied ML, ML engineering, ML research, or technical solutions roles, with at least 2+ years focused specifically on LLM evaluation and/or post-training. * Hands-on experience fine-tuning LLMs (SFT at minimum; preference optimization methods like RLHF, DPO, or KTO strongly preferred) and designing the data pipelines that feed them. * Deep familiarity with LLM evaluation methodology: public benchmarks and their limitations, custom benchmark construction, LLM-as-judge design and its failure modes, inter-annotator agreement, and human eval workflow design. * Strong fluency in Python and the modern LLM toolchain (Hugging Face, PyTorch, vLLM, evaluation frameworks such as lm-evaluation-harness, lighteval, or equivalents). * Excellent technical communication. You can hold your own in a room with research scientists at a frontier lab and, an hour later, brief a non-technical executive on the same engagement. * A consultative mindset: you ask sharp questions, you push back when a customer's stated request won't actually solve their problem, and you are comfortable owning a recommendation. * Bachelor's or advanced degree in computer science, machine learning, computational linguistics, or related field - or equivalent demonstrated experience. ## Description * Lead technical discovery with prospective and existing customers - foundation model labs, frontier AI teams, and large enterprises - to understand model objectives, gaps, and constraints. * Design end-to-end solutions across the post-training stack: SFT data curation, preference data collection for RLHF/DPO, golden datasets, custom benchmarks, LLM-as-judge pipelines, human-in-the-loop evaluation, red teaming, and multimodal eval (text, image, audio, video, long-context). * Architect engagements that combine Innodata's platforms (GenAI Test & Evaluation Platform, Annotation Platform, GenAI Workbench) with our global SME workforce across 85+ languages and domains. * Author technical proposals, SOWs, solution diagrams, and pricing models in partnership with sales, delivery, and finance. * Run technical workshops, POCs, and pilot designs that de-risk larger programs and prove value quickly. * Serve as the ongoing technical advisor during delivery, partnering with applied research scientists, AI/ML research engineers, language data scientists, and program managers to keep solutions aligned with the original intent. * Feed customer signal back into Innodata's R&D and product roadmap - what benchmarks customers actually want, where eval methodology is breaking, what new fine-tuning paradigms are gaining traction. * Stay current on the state of the art in evals (e.g., dynamic and agentic benchmarks, capability vs. safety evals, long-context and tool-use evaluation) and post-training (SFT, RLHF, DPO, RLAIF, rejection sampling, distillation). * Represent Innodata externally - at customer reviews, conferences, and in technical content. ## Related Videos - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)