> Markdown version of [/jobs/ext/2210012-deep-learning-engineer-accuracy-evaluation](https://www.wearedevelopers.com/jobs/ext/2210012-deep-learning-engineer-accuracy-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Deep Learning Engineer, Accuracy Evaluation - **Company:** Nvidia - **Location:** Spain - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Code Generation, Statistical Hypothesis Testing, Machine Learning, Regression Analysis, Open Source Technology, Large Language Models, Information Technology, Slurm, Machine Learning Operations - **Published:** August 24, 2026 - **Apply:** https://www.jobleads.com/es/job/eb258d40ce21c8dff8f38ce8e61057841 ## About the Role * BS, MS, or PhD in Computer Science, Machine Learning, Statistics, or a related field * 6+ years of hands-on experience with LLMs, designing and running evaluations for large language models or multimodal AI systems, including experience with agentic, multi-turn, or reasoning-heavy settings. * Strong statistical foundations: experimental design, significance testing, regression analysis, and the ability to distinguish signal from noise in benchmark results at scale. * Proven experience building evaluation infrastructure (pipelines, benchmark harnesses, reproducible CI systems) not just consuming existing benchmarks. * Clear, precise communicator who can translate quantitative evaluation results into decisions for researchers, product teams, and senior leadership. Ways To Stand Out From The Crowd * Deep familiarity with open-source evaluation frameworks * Experience designing evaluations for agentic systems: tool use, multi-turn reasoning, environment-based benchmarks (SWE-bench, GAIA, WebArena-style), or interactive evaluation settings. * Track record of publishing or contributing to evaluation research, new benchmark design, methodology papers, or reproducibility analyses that shaped how the field measures model capability. * Experience measuring model accuracy under low-precision inference (FP8, INT4, quantization-aware settings) and understanding how calibration and sparsity interact with benchmark results. * Comfort running large-scale workloads on HPC/Slurm clusters, including reproducible experiment management (MLflow, W&B) and compute cost optimization across hundreds of benchmark runs. ## Description * Design and build decision-grade evaluation environments for NVIDIA's frontier models spanning reasoning, multimodal, long-context, and agentic systems, producing auditable accuracy signals that gate every major model release. * Research and develop novel evaluation methodologies for emerging model families and capability domains (low-precision numerics, multi-turn agentic tasks, code generation) where established benchmarks don't yet exist or don't generalize. * Build and operate the evaluation infrastructure and pipelines including benchmark environments, regression CI systems, and statistical analysis tooling, used by model, product, and applied research teams across NVIDIA. * Partner with model research, training, and customer teams to translate evaluation signals into concrete decisions: release go/no-go, training iteration direction, and competitive positioning against external frontier models. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Leapter: The Reinvention of Software Development? A Future Built On AI Generated Code.](https://www.wearedevelopers.com/videos/1663-leapter-the-reinvention-of-software-development-a-future-built-on-ai-generated-code) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [What non-automotive Machine Learning projects can learn from automotive Machine Learning projects](https://www.wearedevelopers.com/videos/397-what-non-automotive-machine-learning-projects-can-learn-from-automotive-machine-learning-projects) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)