> Markdown version of [/jobs/ext/2712441-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2712441-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** PLAUD INC. - **Location:** San Francisco, CA, United States - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Codecs, Distributed Systems, Machine Learning, Data Streaming, WebSocket, WebRTC, Graphics Processing Unit (GPU), Chatbots, Delivery Pipeline, Large Language Models, Parallel Computation, Audio Streaming, Backend, Kubernetes, Low Latency, Machine Learning Operations, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/machine-learning-engineer-inference-serving-speech-llm-san-francisco-plaud-8103777 ## About the Role * Have hands-on experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models. * Understand the intricate tradeoffs between latency, throughput, and Time-To-First-Token (or Time-To-First-Audio) in real-time streaming environments. * Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI. * Possess a deep understanding of GPU architectures (NVIDIA Ampere/Hopper) and the memory hierarchy, allowing you to identify and eliminate hardware bottlenecks. * Communicate clearly and collaborate effectively, as you will sit at the critical intersection between the core ML training team and the backend infrastructure team. * Thrive in fast-moving environments and genuinely enjoy the systems-engineering challenge of squeezing every last drop of performance out of a cluster of GPUs. * Are obsessed with building AI systems that natively understand and generate speech, ultimately creating a hardware-software AI companion that amplifies human productivity. Strong candidates may also have experience with: * Frontier Serving Frameworks: Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories). * Real-Time Audio Streaming: Experience handling continuous audio streams over WebSockets or WebRTC, deploying neural audio codecs, and managing chunked audio generation to minimize conversational latency. * Advanced Inference Techniques: Implementing cutting-edge generation algorithms such as speculative decoding, lookahead decoding, or chunked prefill. * Model Compression & Quantization: Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy. * Large-Scale Distributed Systems: Deploying multi-GPU (Tensor Parallelism) and multi-node inference pipelines, and managing autoscaling infrastructure using Kubernetes. ## Related Videos - [Performant Architecture for a Fast Gen AI User Experience](https://www.wearedevelopers.com/videos/1158-performant-architecture-for-a-fast-gen-ai-user-experience) - [Transforming Education: A Journey from interactive Markdown to Remote-Labs](https://www.wearedevelopers.com/videos/941-transforming-education-a-journey-from-interactive-markdown-to-remote-labs) - [Can You Touch the Internet? A Journey in Remote Touch With Edge Computing and Haptic Coding](https://www.wearedevelopers.com/videos/1463-can-you-touch-the-internet-a-journey-in-remote-touch-with-edge-computing-and-haptic-coding) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Raise your voice!](https://www.wearedevelopers.com/videos/10-raise-your-voice) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)