> Markdown version of [/jobs/ext/731838-software-engineer-model-evaluation-benchmarking](https://www.wearedevelopers.com/jobs/ext/731838-software-engineer-model-evaluation-benchmarking). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer (Model Evaluation & Benchmarking) - **Company:** SpreeAI Corporation - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Data Analysis, Automation of Tests, C++ (Programming Language), Computer Programming, Continuous Delivery, Continuous Integration, Data Structures, Python (Programming Language), Machine Learning, NumPy, Object-Oriented Software Development, Open Source Technology, Visual Systems, Large Language Models, Model Validation, Generative AI, Pandas, Information Technology, HuggingFace, Stable Diffusion - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=bc4972d2e887975e ## About the Role Do you have experience in Research?, * LLM, VLM, or Stable Diffusion model evals * Image/Video benchmarking techniques * Multimodal evaluation frameworks * dataset-driven testing workflows * research experiment validation pipelines, * Degree in Computer Science, AI, Engineering, or comparable combination of education and practical experience. * Strong programming skills in Python. * Familiarity with object-oriented programming (C++, Java, Python, or similar). * Strong data structures and algorithms fundamentals. * Understanding of machine learning experimentation workflows., * Experience evaluating vision or generative models. * Familiarity with HuggingFace ecosystem or open-source ML toolkits. * Experience building automated test frameworks or benchmarking tools. * Knowledge of diffusion models or multimodal architectures. Experience with data analysis tools (NumPy, Pandas, visualization libraries). ## Description We are hiring Engineers focused on AI Model Evaluation to build the systems that ensure multimodal AI behaves reliably, consistently, and predictably as it moves from research into production. This position focuses on evaluating generative and vision-based models through automated benchmarking, dataset-driven testing, and performance validation pipelines. You will work at the intersection of applied science, infrastructure, and product - helping define how we measure realism, consistency, and quality across image, video, and multimodal AI systems. Why This Role Exists Modern AI evaluation extends beyond pass/fail testing. Multimodal generative systems require: * benchmarking across visual realism, pose consistency, and identity preservation, * automated regression detection across model checkpoints, * scalable evaluation pipelines integrated into continuous deployment workflows. We are building evaluation systems where research velocity and product reliability must coexist. This role is for engineers interested in defining how quality is measured in generative AI systems. What you'll do * Build automated evaluation pipelines for multimodal AI models. * Benchmark diffusion models, vision systems, and generative workflows. * Validate model checkpoints and detect regressions across versions. * Develop evaluation metrics for realism, consistency, and performance. * Integrate evaluation tooling into CI/CD workflows. * Collaborate with ML researchers and infrastructure teams to ensure production readiness. * Analyze failure modes and propose evaluation strategies. ## Related Videos - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Evals vs. Evil - AI and Package Security - Laurie Voss](https://www.wearedevelopers.com/videos/2131-evals-vs-evil-ai-and-package-security-laurie-voss) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)