> Markdown version of [/jobs/ext/2622317-senior-software-engineer-scientific-evaluation](https://www.wearedevelopers.com/jobs/ext/2622317-senior-software-engineer-scientific-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - Scientific Evaluation - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $82,735.0 - $97,335.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), CMake, Nvidia CUDA, Computer Engineering, Linux, Distributed Systems, Github, Python (Programming Language), Software Engineering, Pytorch, Large Language Models, Caching, Gitlab, Git, Kubernetes, Information Technology, Slurm, Data Pipelines, Docker - **Published:** August 9, 2026 - **Apply:** https://www.jofdav.com/jobs/59184331-senior-software-engineer-scientific-evaluation ## About the Role * A BS or MS, or equivalent experience, in Computer Science, Computer Engineering, or a related field. * 5 plus years of relevant industry experience. * Strong CS fundamentals and production C++ and Python skills, with fluency in Linux, Git, GitHub/GitLab pipeline orchestration, CMake, Docker, Python packaging, and containers. * A record of creating test, benchmark, evaluation, or distributed execution systems that deliver versioned, reproducible results across repositories. * Hands-on operation of shared GPU compute with Slurm, Kubernetes, or a similar scheduler, including monitoring, isolation, and failure recovery. Sound measurement practices cover correctness, variance, flakiness, scorer calibration, and regression detection. Ways to stand out from the crowd: * Background in agentic or LLM evaluation is valuable, especially tool-use tasks, sandboxed execution, trace analysis, model-assisted scoring, and calibration. Multi-GPU or multi-node systems, CUDA software development, statistical benchmarking, and Nsight profiling can also distinguish an application. * Creative, collaborative engineers who care about reliable science are encouraged to apply! ## Description We are seeking a Software Engineer - Scientific Evaluation to own a shared platform for classical testing, scientific benchmarking, and agentic evaluation. The portfolio spans CUDA and C++ libraries, Python packages, PyTorch integrations, scientific models, and AI agents. This hands-on role combines production software engineering, rigorous measurement, distributed systems, and large GPU fleets. What you will be doing: * Own the architecture and roadmap for evaluation and benchmarking of scientific software, serving agentic and numerical evaluations. * Design and implement infrastructure, benchmarks and test levels for classical software and scientific agents, balancing rapid turnaround time and thorough scientific coverage. * Establish trusted references, numerical tolerances, calibrated scorers, regression thresholds, and human-review hooks for nondeterministic workloads. * Operate heterogeneous GPU capacity using multiple control plane technologies, self-hosted runners, schedulers, queues, containers, caching, observability, and automated recovery. * We partner with applied scientists, kernel and framework engineers, product teams, and release owners to create reproducible quality signals and production-readiness gates. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Code to Road in < 12 hours](https://www.wearedevelopers.com/videos/1082-code-to-road-in-12-hours) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)