> Markdown version of [/jobs/ext/2039083-staff-software-engineer-ai-ml](https://www.wearedevelopers.com/jobs/ext/2039083-staff-software-engineer-ai-ml). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer AI/ML - **Company:** Samsung - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $163,000.0 - $253,000.0 - **Contract:** Temporary contract - **Skills:** A/B Testing, Artificial Intelligence, Systems Engineering, Computer Vision, Distributed Systems, Memory Management, Python (Programming Language), Machine Learning, Language Modeling, Network Planning and Design, Tensorflow, Software Engineering, Pytorch, Large Language Models, Multi-Agent Systems, Information Technology, Low Latency, Machine Learning Operations, TensorRT, Software Version Control - **Published:** August 12, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17911441?backUrl=%2Fcareer%2F17911441%2FStaff-Software-Engineer-Ai-Ml-California-San-Jose ## About the Role * BS with 10+ years, MS with 8+ years, or PhD with 5+ years of experience in Computer Science, Electrical Engineering, or a related field * Demonstrated track record of deploying and operating LLM- or vision-powered systems in production, with proven experience in evaluation and safe rollout practices (A/B testing, canary releases, offline/online evaluation) - including defining quality metrics, benchmarking models, and driving quality/cost/latency trade-off decisions * Hands-on expertise with agentic AI frameworks (LangGraph, CrewAI, ADK) - multi-agent orchestration and patterns, tool routing, state management, and/or interoperability protocols (MCP, A2A) * Experience with LLM inference optimization and serving (e.g., vLLM, TensorRT-LLM), covering quantization, KV-cache management, and hardware-aware deployment * Strong Python and C/C++ skills with deep experience in PyTorch/TensorFlow, and experience designing, building, and securing large-scale distributed systems Preferred Qualification * PhD of software engineering experience * MLOps exposure: prompt/version management, monitoring, observability * Experience with ML, graphics or computer vision accelerator * Understanding of PPA (performance, power, and area) trade-offs, memory controller architecture and/or general computer architecture is beneficial * Familiarity with state-of-the-art AI workloads and their compute and memory requirements * Experience with performance modeling of heterogenous systems is beneficial * Ability to meet aggressive project deadlines in a team environment ## Description We are seeking a Senior Staff Engineer to build and optimize the EDA design environment and large-scale compute infrastructure that supports our semiconductor design organizations, and to establish company-wide standards and processes for EDA licensing and R&D software. In this role, you will design and standardize the shared design environments and infrastructure used across multiple design organizations, driving engineering productivity and cost efficiency at scale. AI/PI Group is an internal AI and Process Innovation organization within Samsung DSA, dedicated to transforming how our company works through artificial intelligence. We accelerate AI adoption across the entire organization - from reshaping day-to-day workflows and automating core internal processes to empowering our workforce with practical, intelligent tools. The Applied AI Engineering team is the driving force behind this vision, leading innovation at the intersection of machine learning and system engineering to develop and operate our next-generation AI frameworks. We aim to enable frontier AI models to autonomously plan, retrieve information, coordinate with tools, and execute multi-step workflows across our internal knowledge ecosystem. We are actively seeking talented Machine Learning Engineers specializing in building next-generation AI/ML solutions. What You'll Do * Design, build, and productize agentic AI applications that combine vision and language models - owning the architecture end-to-end from prototype through deployment and ongoing operation. * Architect multi-agent systems using frameworks such as LangGraph, LangChain, AutoGen, or CrewAI, including agent orchestration, tool and function routing, state and memory management, error recovery, and human-in-the-loop workflows. * Establish evaluation frameworks and quality bars for agentic and multimodal systems: define reference and non-reference metrics, build benchmark suites and regression harnesses, and instrument agent trajectories for offline and online evaluation. * Drive quality, cost, and latency trade-off decisions with data - selecting models, routing strategies, and inference configurations that meet product requirements within compute and memory budgets. * Optimize inference performance on accelerated hardware through quantization, batching, caching, kernel- and runtime-level tuning, and model/hardware co-design; profile workloads to identify and eliminate bottlenecks. * Identify and solve multi-discipline AI acceleration problems, especially with memory bottlenecks, involving algorithms, network design, hardware architecture * Work with researchers and application developers to enable the latest machine learning work to optimize performance., At Samsung Semiconductor, we use Artificial Intelligence (AI) tools in the recruitment process to enhance efficiency. However, AI is used as a support tool, not a final decision-maker. All hiring decisions are made by our human recruiting team and hiring managers to ensure every candidate is evaluated fairly and holistically. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)