> Markdown version of [/jobs/ext/2565036-software-engineer-inference-runtime](https://www.wearedevelopers.com/jobs/ext/2565036-software-engineer-inference-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Inference Runtime - **Company:** LM STUDIO INC. - **Location:** New York, NY, United States (Remote available) - **Salary:** $150,000.0 - $350,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Nvidia CUDA, Computer Programming, Extract Transform Load (ETL), Python (Programming Language), Open Source Technology, Pytorch, Large Language Models, Machine Learning Operations, TensorRT - **Published:** August 7, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=68c66e2dd3b17485 ## About the Role * Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure * Strong programming ability in Python and C++ * Deep understanding of transformer architectures and the mechanics of model inference * Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement * Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM * Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution * Takes personal responsibility for the correctness and performance of their work Bonus Qualifications * Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM ## Description We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on., * Maintain and push forward our inference stack on-device and in the cloud * Bring up new model architectures and multimodal models * Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes * Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution * Benchmark and diagnose correctness and performance problems across the inference stack * Contribute upstream to open-source projects such as llama.cpp and MLX ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)