> Markdown version of [/jobs/ext/2704117-research-engineer-ai-ml-systems](https://www.wearedevelopers.com/jobs/ext/2704117-research-engineer-ai-ml-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer, AI/ML Systems - **Company:** Lightning AI - **Location:** San Francisco, CA, United States (Remote available) - **Salary:** $165,000.0 - $310,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Nvidia CUDA, Software Debugging, Programming Tools, Distributed Computing Environment, Distributed Systems, Machine Learning, Language Modeling, Open Source Technology, Performance Tuning, Software Construction, Software Engineering, AI Infrastructure, Pytorch, Deep Learning, Generative AI, Backend, Information Technology, HuggingFace, Free and Open-Source Software, Machine Learning Operations, GPT, Automation Anywhere - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/research-engineer-ai-ml-systems-lightning-ai-5647712 ## About the Role We're looking for a curious, adaptable Research Engineer who enjoys solving difficult technical problems and building across the AI stack to join our Research Engineering function here at Lightning. This role is intentionally broad, with a primary focus on post-training models and the systems that support it. You'll work across ML engineering, software engineering, and AI systems to improve how we develop, train, evaluate, and deploy models. As team priorities evolve, you'll have opportunities to contribute across developer tooling, infrastructure, and platform capabilities that help researchers and customers develop, train, and deploy AI more effectively. We're looking for someone who enjoys learning new technologies, working across multiple technical domains, and tackling whatever problems have the greatest impact. Strong software engineering fundamentals, curiosity, and a willingness to continuously learn are more important than already being an expert in every area of AI systems. If you've spent meaningful time building AI projects, experimenting with PyTorch, contributing to open source, reproducing research, or exploring new ideas because you're genuinely interested, we'd love to hear about it. This role is hybrid with a minimum of 2 in-office days per week in San Francisco, Seattle, NYC, or London, with fully remote work considered for candidates outside of our office hub locations. All employees participate in occasional team and company offsites., * Experience building, training, evaluating, or experimenting with deep learning models. * Hands-on experience with deep learning frameworks such as PyTorch. * Strong software engineering fundamentals building software and debugging and problem-solving skills, with the ability to investigate unfamiliar technical challenges. * Curiosity, initiative, and a demonstrated ability to quickly learn new technologies and technical domains. * Excellent communication and collaboration skills, including the ability to work effectively across research, product, infrastructure, and customer-facing engagements. * Comfortable working in fast-moving, ambiguous environments where priorities evolve over time. * Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience., * Experience with model training at scale, including distributed training, performance optimization, training stability, and/or large-scale experimentation. * Experience with transformer-based language models or modern generative AI systems. * Experience with distributed systems, cloud infrastructure, or large-scale machine learning workloads. * Familiarity with technologies such as CUDA, Hugging Face, DeepSpeed, FSDP, Triton, vLLM, SGLang, NVIDIA Molt, or related AI infrastructure tooling. * Experience contributing to open-source software or conducting research through academia, industry, or meaningful independent projects. * Startup experience or experience working on highly cross-functional engineering teams. * Master's degree or higher in Computer Science, Machine Learning, AI, or a related field. ## Description * Develop and post-train models, while building and improving the systems and workflows needed to run, evaluate, debug, and scale training workloads. * Build software, tooling, and platform capabilities that improve how researchers, developers, and customers develop, train, and deploy AI systems. * Contribute to Lightning's open-source projects by building new features, improving existing functionality, and collaborating with the broader developer community. * Work across deep learning systems, developer tooling, backend services, and platform infrastructure to solve a wide variety of engineering challenges. * Collaborate directly with customers to understand real-world AI workloads, investigate technical challenges, and translate those learnings into reusable product and platform improvements. * Prototype new ideas, evaluate approaches, and turn successful experiments into production-quality software. * Partner closely with research, product, and infrastructure engineering teams to improve developer experience, AI workflows, and platform capabilities. * Debug complex technical problems spanning machine learning, distributed systems, backend software, and developer tooling. * Learn new technologies quickly and contribute wherever your skills can have the greatest impact as team priorities evolve. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Localized Open Models in Production: What Builders Need to Know](https://www.wearedevelopers.com/videos/100270-localized-open-models-in-production-what-builders-need-to-know) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)