> Markdown version of [/jobs/ext/2294322-senior-staff-engineer-ai-workloads-storage](https://www.wearedevelopers.com/jobs/ext/2294322-senior-staff-engineer-ai-workloads-storage). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Staff Engineer - AI Workloads & Storage - **Company:** Samsung - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $189,000.0 - $301,000.0 - **Contract:** Permanent contract - **Skills:** Adobe Flash, Artificial Intelligence, C++ (Programming Language), Computer Engineering, Extract Transform Load (ETL), Linux, Firmware, PCI Express, Remote Direct Memory Access, Software Engineering, SystemC, Weka, Ceph (Software), Large Language Models, Perf (Linux), Information Technology, Low Latency, TensorRT, Nvme - **Published:** August 29, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88183116/1 ## About the Role * Bachelor's degree 15+ years relevant industry experience or Master's degree 13+ years' experience or PhD with 10+ years relevant industry experience. * Extensive experience (typically10-15+ years) in systems, storage, or ML-systems software, with a track record of architecting systems that materially improved performance, reliability, or cost. * Demonstratedtechnical leadership and cross-team influence: setting direction, driving decisions across organizational boundaries, and mentoring senior engineers. * Working knowledge of modern AI inference, especially transformer architectures - attention, KV cache, batching, and the memory/compute trade-offs of serving large models. * Deep systems-level understandingof the Linux storage stack (block layer, I/O scheduling, NVMe) and ofNAND/SSD internals(flash-translation layer, garbage collection, endurance/write-amplification, latency behavior), plus hands-on performance analysis skill (e.g., perf, ftrace, eBPF, blktrace, fio). * Fluency inPython plus a systems language(C/C++, Rust, or Go). * MS or PhDin Computer Science, Electrical/Computer Engineering, or a related field preferred - or equivalent practical experience. Preferred Qualification * Hands-on experience with the moderninference stack: vLLM, SGLang, LMCache, NVIDIA Dynamo, TensorRT-LLM, or Triton. * Familiarity withGPU-adjacent data movement and memory frameworks: NIXL, DOCA / DOCA MemOps, GPUDirect Storage, RDMA, NVMe-oF, and BlueField / DPU offload. * Understanding of GPU and TPU architecture(memory hierarchy, interconnects, and how accelerator design shapes I/O and data-movement demands) is highly desired. * Experience withuser-mode storage access frameworks: SPDK, uNVMe, libvfn, or similar. * SSD firmware experience- flash-translation layer, wear-leveling and garbage-collection algorithms, and data-placement features such asFDP / streams / ZNS- ideally paired with the ability to co-design firmware and host-side placement policy from workload characterization. * AI-workload characterization and benchmarkingexperience, and familiarity withSNIA Storage.AIandMLCommons / MLPerf. * Transactional / discrete-event or system-level modelingexperience in frameworks such as SystemC, SimPy, or similar. * Experience withSSD architecture and interfaces- NVMe (including ZNS, Flexible Data Placement / FDP), open-channel SSDs, computational storage - and with PCIe Gen5, CXL, and large-scale GPU-cluster storage (VAST, WEKA, Lustre, Ceph). ## Description * Own AI workload characterization.Profile production and emerging LLM inference, RAG, and training workloads to quantify their I/O, bandwidth, latency, and capacity demands, and turn those findings into concrete storage and memory-hierarchy design decisions. * Identify optimal data placement.Analyze workload access patterns to determine how data should be placed and separated on flash, and map those insights onto SSD data-placement technologies such asNVMe Flexible Data Placement (FDP)andstreamsto reduce write amplification and improve endurance, latency, and QoS. * Collaborate with key customersto identify differentiating SSD capabilities for AI workloads, and develop proof-of-concept implementations as part of those customer engagements - turning workload insights into demonstrable data-path, tiering, and data-placement wins. * Lead deep-dive performance analysisspanning the inference runtime, the Linux storage and networking stack, and the underlying hardware, tuning for latency, throughput, cost, and GPU utilization. * Build and evaluate transactional and system-level modelsof proposed architectures to de-risk decisions before hardware exists, and validate them against measured behavior. * Engage with the standards and open ecosystem- SNIA (including Storage.AI), MLCommons/MLPerf, and the open inference stack - to align our work with where the industry is heading and to shape it where we can. * Set technical direction others build on.Make build-vs-buy and architectural calls, establish benchmarking methodology and best practices, and mentor engineers across the org. * Partner cross-functionallywith product, hardware, and research teams, and with external vendors and partners, to bring architectures from concept to deployment., At Samsung Semiconductor, we use Artificial Intelligence (AI) tools in the recruitment process to enhance efficiency. However, AI is used as a support tool, not a final decision-maker. All hiring decisions are made by our human recruiting team and hiring managers to ensure every candidate is evaluated fairly and holistically. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)