> Markdown version of [/jobs/ext/2719855-research-engineer-ai-systems](https://www.wearedevelopers.com/jobs/ext/2719855-research-engineer-ai-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer - AI Systems - **Company:** Yotta Labs Inc - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, C++ (Programming Language), Profiling, Nvidia CUDA, Distributed Computing Environment, High-Level Architecture, Python (Programming Language), Linux Kernel, Open Source Technology, Performance Tuning, AI Infrastructure, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Information Technology, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/research-engineer-ai-systems-yotta-labs-8610906 ## About the Role * Proficiency in AI programming languages such as Python and C++. * Deep understanding of GPU architecture and performance optimization. * Experience with CUDA, Triton, ROCm/HIP, or AWS Neuron. * Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler). * Strong problem-solving skills and the ability to work in a collaborative, remote environment. * A background in computer science, engineering, or a related field is preferred. Preferred Experience * Contributions to open-source AI infra projects like vLLM, SGLang, PyTorch, or Triton. * Experience with with FlashAttention, PagedAttention, MoE, RLHF, or distributed AI systems. * Publications in top-tier conferences like MLSys, OSDI, SOSP, NSDI, SC, HPCA, or ISCA ## Description We are seeking a highly motivated AI Systems Research Engineer specializing in Trainium, GPU kernels, and LLM systems optimization. You will work at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, GPU/Accelerator Kernel Development, and Large Language Model Infrastructure. Your work will directly impact the scalability and performance of AI applications deployed on our platform., * Design and implement high-performance kernels for Attention, MoE, GEMM, collective communication, and quantization. * Optimize kernels for NVIDIA, AMD, and AWS Trainium. * Develop custom operators and graph optimizations using Neuron SDK, PyTorch/XLA, Torch Dynamo, and Neuron Compiler. * Improve performance of vLLM, SGLang, TensorRT-LLM, and custom inference runtimes. * Design scalable distributed training and inference solutions across thousands of accelerators. * Contribute to open-source projects, publish technical findings and engage with the developer community. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)