> Markdown version of [/jobs/ext/1800937-senior-performance-engineer](https://www.wearedevelopers.com/jobs/ext/1800937-senior-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Performance Engineer - **Company:** Samsung - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $138,000.0 - $206,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Application Frameworks, Application Performance Management, Systems Engineering, C++ (Programming Language), Profiling, Computer Engineering, Distributed Systems, Python (Programming Language), AI Infrastructure, Graphics Processing Unit (GPU), High Performance Computing, Pytorch, Large Language Models, AI Platforms, Information Technology, TensorRT, Virtual Agents - **Published:** July 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17453845?backUrl=%2Fcareer%2F17453845%2FSenior-Performance-Engineer-California-San-Jose ## About the Role * MS or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field. * B.S with 5+ years of experience in performance engineering, AI systems, distributed systems, high-performance computing, or a related area.MS in Computer/Electrical Engineering or Computer Science with 3+ years of relevant working experience or PhD and 0+ years of relevant working experience preferred. * Strong understanding of LLM inference and training systems. * Strong understanding of NVIDIA GPU architecture and performance characteristics, including compute, memory hierarchy, communication, and system-level bottlenecks. * Hands-on experience profiling and optimizing AI workloads on NVIDIA GPU platforms using tools such as Nsight Systems, Nsight Compute, and related performance analysis frameworks. * Experience analyzing performance of large-scale distributed AI workloads. * Proficiency in Python and C++. * Experience with one or more modern AI frameworks or serving systems, such as PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, or Megatron-LM. * Strong analytical and problem-solving skills. ## Description We are seeking a Senior LLM Systems Performance Engineer to build representative AI environments, characterize emerging workloads, and drive performance analysis for next-generation AI platforms. In this role, you will set up and operate realistic LLM serving and agentic AI environments, collect workload traces and performance data, and develop methodologies to characterize workload behavior. You will analyze system bottlenecks across compute, memory, communication, and scheduling resources, and evaluate how emerging workloads interact with AI accelerator architectures and system infrastructure. The ideal candidate combines hands-on experience building large-scale AI systems with strong performance engineering skills and a solid understanding of AI accelerator architecture. You should be comfortable working across the full stack-from application frameworks and serving systems to runtime software, networking, memory systems, and accelerator hardware.You will work closely with hardware architects, systems engineers, and software researchers to understand the performance implications of emerging workloads such as agentic AI, long-context reasoning, disaggregated inference, and Mixture-of-Experts models. Your analysis will help shape future hardware-software co-design decisions and guide the development of next-generation AI infrastructure. Location: Daily onsite presence at our San Jose, CA office / U.S. headquarters in alignment with our Flexible Work policy. What You'll Do * Build and operate representative AI environments, including agentic workflows, distributed inference systems, disaggregated serving architectures, and MoE deployments. * Collect workload traces, telemetry, and performance data from real-world AI applications; characterize workload behavior, develop representative benchmarks, and identify performance bottlenecks across compute, memory, communication, and scheduling resources. * Evaluate AI systems across the full hardware and software stack, and analyze the impact of runtime, memory hierarchy, interconnect, and accelerator architecture on application performance. * Collaborate with hardware and software teams to drive performance analysis, architecture exploration, and hardware-software co-design for next-generation AI platforms., At Samsung Semiconductor, we use Artificial Intelligence (AI) tools in the recruitment process to enhance efficiency. However, AI is used as a support tool, not a final decision-maker. All hiring decisions are made by our human recruiting team and hiring managers to ensure every candidate is evaluated fairly and holistically. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)