> Markdown version of [/jobs/ext/2691726-ai-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2691726-ai-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Performance Engineer - **Company:** Greenhouse Software, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $180,000.0 - $255,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Program Optimization, Nvidia CUDA, Data Structures, Distributed Systems, Data Flow Control, Python (Programming Language), Machine Learning, OpenCL, Performance Tuning, Tensorflow, Pytorch, Large Language Models, Deep Learning, Gpu Programming, Numerical Computing, Information Technology, Low Latency, Machine Learning Operations, TensorRT, GPT - **Published:** September 3, 2026 - **Apply:** https://boards.greenhouse.io/embed/job_app?for=sambanovasystems&token=5719052004 ## About the Role * Bachelor's or higher degree in computer science, electrical engineering, or a related field (e.g., applied mathematics, physics, or statistics). * 3+ years of experience in one or more of the following areas: * Deep learning model development and performance optimization * Compiler, runtime, or kernel-level optimization * Software-hardware co-design or systems performance tuning * Proficiency in Python or C++, with strong foundations in algorithms, data structures, and numerical computing. * Experience with at least one major ML framework - PyTorch, TensorFlow, or JAX. * Demonstrated ability to analyze and optimize performance in real-world ML pipelines., * Hands-on experience with LLM or multimodal model training and inference. * Background in large-scale distributed training, continuous batching, and high-throughput inference systems. * Familiarity with quantization, graph optimization, kernel fusion, and model partitioning. * Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT. * Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus. * Knowledge of memory hierarchy optimization, caching, and scheduling for large-scale model execution. * Publication record or open-source contributions in ML systems or performance optimization is a plus. ## Description We are seeking a talented and driven ML performance engineer to optimize and scale state-of-the-art foundation models on SambaNova's reconfigurable dataflow platform. You'll work hands-on with some of the most advanced models in the world - such as DeepSeek R1, GPT OSS, and other frontier architectures - to push the limits of throughput, latency, and efficiency. In this role, you'll bridge the gap between deep learning and systems performance, collaborating across compiler, runtime, and hardware layers to deliver world-record performance for large-scale AI inference., * Bring up and optimize cutting-edge foundation models (e.g., DeepSeek, Llama, Qwen, and others) on the SambaNova platform through the SambaNova software stack. * Profile and enhance model performance across compiler, runtime, and hardware layers to achieve SOTA throughput and latency. * Collaborate with machine learning, compiler, runtime, and hardware teams to deliver co-designed, high-performance AI applications. * Integrate the latest advances in model architecture, quantization, scheduling, and memory optimization from both academia and industry. * Develop robust, scalable, and efficient end-to-end inference solutions aligned with customer needs. * Identify performance bottlenecks and propose dataflow or scheduling optimizations for both single-node and distributed systems. ## Related Videos - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)