Staff Software Engineer AI/ML

Samsung
San Jose, CA, United States
1 day ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

A/B Testing Artificial Intelligence C++ (Programming Language) Distributed Systems Memory Management Python (Programming Language) Performance Tuning Tensorflow Pytorch Large Language Models Multi-Agent Systems Caching
+4 more
Information Technology Machine Learning Operations TensorRT Virtual Agents

Job description

Experteer Overview In this role you will design and optimize the AI-enabled design and compute infrastructure that powers Samsung’s semiconductor R&D. You’ll shape company-wide AI tooling, collaborating with researchers and engineers to deploy robust, scalable agentic AI solutions. Expect to tackle performance, cost, and latency challenges across multi-domain workloads. This is a hands-on leadership role that drives AI adoption and practical impact across the organization. Compensation / Benefits * Design, build, and productize agentic AI applications spanning prototype to production and operation * Architect multi-agent systems with frameworks like LangGraph, LangChain, AutoGen, or CrewAI, including orchestration and memory management * Establish evaluation frameworks, metrics, benchmarks, and regression harnesses for agentic/multimodal systems * Drive quality, cost, and latency trade-offs using data to meet product requirements within compute and memory budgets * Optimize inference on accelerated hardware via quantization, batching, caching, and hardware-aware deployment * Identify and solve AI acceleration challenges, including memory bottlenecks and hardware architecture considerations * Collaborate with researchers and developers to enable cutting-edge ML work and optimize system performance Tasks * BS with 10+ years, MS with 8+ years, or PhD with 5+ years in Computer Science, Electrical Engineering, or related field * Proven deployment/operation of LLM- or vision-powered systems with strong evaluation and safe rollout practices (A/B testing, canary releases) * Hands-on experience with agentic AI frameworks (LangGraph, CrewAI, ADK) and multi-agent patterns * Experience with LLM inference optimization and serving (vLLM, TensorRT-LLM) including quantization and KV-cache management * Strong Python and C/C++, with deep PyTorch/TensorFlow experience and large-scale distributed systems * Preferred: MLops exposure, AI acceleration hardware experience, understanding of PPA trade-offs Key requirements * 4+ weeks paid time off * medical/dental/vision/401k * charitable giving match * emotional wellness support and therapy * onsite cafe and gym * flexible environment

Requirements

on accelerated hardware via quantization, batching, caching, and hardware-aware deployment * Identify and solve AI acceleration challenges, including memory bottlenecks and hardware architecture considerations * Collaborate with researchers and developers to enable cutting-edge ML work and optimize system performance Tasks * BS with 10+ years, MS with 8+ years, or PhD with 5+ years in Computer Science, Electrical Engineering, or related field * Proven deployment/operation of LLM- or vision-powered systems with strong evaluation and safe rollout practices (A/B testing, canary releases) * Hands-on experience with agentic AI frameworks (LangGraph, CrewAI, ADK) and multi-agent patterns * Experience with LLM inference optimization and serving (vLLM, TensorRT-LLM) including quantization and KV-cache management * Strong Python and C/C++, with deep PyTorch/TensorFlow experience and large-scale distributed systems * Preferred: MLops exposure, AI acceleration hardware experience, understanding of PPA

Benefits & conditions

trade-offs Key requirements * 4+ weeks paid time off * medical/dental/vision/401k * charitable giving match * emotional wellness support and therapy * onsite cafe and gym * flexible environment

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:47 min

Performance benchmarks and Android developer adoption

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

Videos

See all

Related articles

See all