> Markdown version of [/jobs/ext/3651181-software-engineer-machine-learning-platform-gen-ai](https://www.wearedevelopers.com/jobs/ext/3651181-software-engineer-machine-learning-platform-gen-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Machine Learning Platform - Gen AI - **Company:** DOORDASH, INC. - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Cloud Computing, Software Code Optimization, Data Cleansing, Cursor, Software Debugging, Distributed Systems, Python (Programming Language), Node.Js, Performance Tuning, Software Engineering, Graphics Processing Unit (GPU), Retrieval-Augmented Generation, Large Language Models, Claude Code, AI Coding Agents, Agentic-AI, Discretization, Kubernetes, Information Technology, SGLang, Machine Learning Operations, TensorRT, VLLM, Model Inference, Data Pipelines, Serverless Computing - **Published:** October 9, 2026 - **Apply:** https://www.thejobnetwork.com/job/bda1b85b-b209-4890-a323-c88cf3c1d7b4/software-engineer-machine-learning-platform-gen-ai ## About the Role - B.S., M.S., or PhD. in Computer Science or equivalent - 3+ years of industry experience in software engineering - Strong backend engineering fundamentals, especially in Python and distributed systems. - Experience building production services, APIs, data pipelines, or ML infrastructure at scale. - Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization. - Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production - serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA). - Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities - Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software , - Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production - Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation - GPU performance work - multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning - Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems - Experience with LLM gateways, model routing, vendor abstraction, or cost attribution - Experience building developer platforms, internal platforms, or self-serve infrastructure - Experience building and deploying AI agents or MCP servers in production - Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases