> Markdown version of [/jobs/ext/2035636-staff-software-engineer-ai-ml](https://www.wearedevelopers.com/jobs/ext/2035636-staff-software-engineer-ai-ml). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer AI/ML - **Company:** Samsung - **Location:** San Jose, CA, United States - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, C++ (Programming Language), Distributed Systems, Memory Management, Python (Programming Language), Performance Tuning, Tensorflow, Pytorch, Large Language Models, Multi-Agent Systems, Caching, Information Technology, Machine Learning Operations, TensorRT, Virtual Agents - **Published:** August 12, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/staff-software-engineer-ai-ml-san-jose-ca-usa-58913400 ## About the Role on accelerated hardware via quantization, batching, caching, and hardware-aware deployment * Identify and solve AI acceleration challenges, including memory bottlenecks and hardware architecture considerations * Collaborate with researchers and developers to enable cutting-edge ML work and optimize system performance Tasks * BS with 10+ years, MS with 8+ years, or PhD with 5+ years in Computer Science, Electrical Engineering, or related field * Proven deployment/operation of LLM- or vision-powered systems with strong evaluation and safe rollout practices (A/B testing, canary releases) * Hands-on experience with agentic AI frameworks (LangGraph, CrewAI, ADK) and multi-agent patterns * Experience with LLM inference optimization and serving (vLLM, TensorRT-LLM) including quantization and KV-cache management * Strong Python and C/C++, with deep PyTorch/TensorFlow experience and large-scale distributed systems * Preferred: MLops exposure, AI acceleration hardware experience, understanding of PPA ## Description Experteer Overview In this role you will design and optimize the AI-enabled design and compute infrastructure that powers Samsung's semiconductor R&D. You'll shape company-wide AI tooling, collaborating with researchers and engineers to deploy robust, scalable agentic AI solutions. Expect to tackle performance, cost, and latency challenges across multi-domain workloads. This is a hands-on leadership role that drives AI adoption and practical impact across the organization. Compensation / Benefits * Design, build, and productize agentic AI applications spanning prototype to production and operation * Architect multi-agent systems with frameworks like LangGraph, LangChain, AutoGen, or CrewAI, including orchestration and memory management * Establish evaluation frameworks, metrics, benchmarks, and regression harnesses for agentic/multimodal systems * Drive quality, cost, and latency trade-offs using data to meet product requirements within compute and memory budgets * Optimize inference on accelerated hardware via quantization, batching, caching, and hardware-aware deployment * Identify and solve AI acceleration challenges, including memory bottlenecks and hardware architecture considerations * Collaborate with researchers and developers to enable cutting-edge ML work and optimize system performance Tasks * BS with 10+ years, MS with 8+ years, or PhD with 5+ years in Computer Science, Electrical Engineering, or related field * Proven deployment/operation of LLM- or vision-powered systems with strong evaluation and safe rollout practices (A/B testing, canary releases) * Hands-on experience with agentic AI frameworks (LangGraph, CrewAI, ADK) and multi-agent patterns * Experience with LLM inference optimization and serving (vLLM, TensorRT-LLM) including quantization and KV-cache management * Strong Python and C/C++, with deep PyTorch/TensorFlow experience and large-scale distributed systems * Preferred: MLops exposure, AI acceleration hardware experience, understanding of PPA trade-offs Key requirements * 4+ weeks paid time off * medical/dental/vision/401k * charitable giving match * emotional wellness support and therapy * onsite cafe and gym * flexible environment ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)