> Markdown version of [/jobs/ext/2846912-agi-artificial-general-intelligence-computing-lab](https://www.wearedevelopers.com/jobs/ext/2846912-agi-artificial-general-intelligence-computing-lab). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AGI (Artificial General Intelligence) Computing Lab - **Company:** Samsung - **Location:** San Jose, CA, United States - **Experience:** Experienced - **Salary:** $163,000.0 - $253,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Nvidia CUDA, Computer Programming, Extract Transform Load (ETL), Distributed Systems, Open Source Technology, Pytorch, Backend, Enterprise Integration, Free and Open-Source Software, Front End Software Development - **Published:** September 11, 2026 - **Apply:** https://www.thejobnetwork.com/job/7021d943-4392-404a-9884-a4e824bbce12/staff-engineer-compiler ## About the Role * Bachelor's with 10+ years, or Master's with 8+ years, or PhD's with 5+ years of industry experience. * 3-5+ years of industry experience in at least one of: Triton, Helion, MLIR, XLA, TVM, Inductor, IREE, CUTLASS, or a proprietary equivalent (More experienced candidates will also be considered at relevant levels). * Experience designing a kernel DSL or its IR from scratch, or making non-trivial language-level changes to an existing one. * Experience with MLIR - writing dialects, passes, or backend integration. * Experience building PyTorch backends for non-CUDA accelerators (XPU, ROCm, MPS, TPU, custom). * Experience with kernel autotuning, performance modeling, or cost-based compilation * Background in HPC, distributed systems, or NUMA-aware programming - anything that built intuition for non-flat memory * Open-source contributions to PyTorch, Triton, Helion, LLVM/MLIR, or similar projects is a big plus. ## Description * Adapting torch.compile to our backend: lowering Inductor's IR to our hardware, defining what gets fused, what gets specialized, and where the compiler should yield to hand-written kernels. * Building or extending kernel DSLs for our hardware: taking a tile-based programming model (Triton-style), a higher-level expression (Helion-style), or a custom DSL we design, and lowering it to our ISA, our memory hierarchy, and our collective primitives. Where existing DSLs' GPU assumptions break, deciding what to change in the frontend, the IR, or the backend. * Designing placement and scheduling passes: given a graph and our distributed memory model, deciding where tensors live, when to migrate them, and how to overlap compute with data movement. This is the layer where our hardware's differentiator shows up most directly. * Implementing parallelism-aware lowering: making tensor, pipeline, expert, and sequence parallelism first-class in the compiler IR rather than bolted on at the framework layer. * Fusion, tiling, and memory planning: the classical compiler problems, reframed for a non-uniform memory hierarchy where the right tile size and the right placement are coupled decisions. * Upstream contributions: where we use open-source DSLs, we want our work to land upstream rather than live in a private fork. You'll engage with upstream review processes for PyTorch, Triton, Helion, and adjacent projects., At Samsung Semiconductor, we use Artificial Intelligence (AI) tools in the recruitment process to enhance efficiency. However, AI is used as a support tool, not a final decision-maker. All hiring decisions are made by our human recruiting team and hiring managers to ensure every candidate is evaluated fairly and holistically. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)