> Markdown version of [/jobs/ext/1908785-founding-speech-model-performance-engineer](https://www.wearedevelopers.com/jobs/ext/1908785-founding-speech-model-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Founding Speech Model Performance Engineer - **Company:** Slng Ltd - **Location:** Barcelona, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Microsoft Word, Nvidia CUDA, Pytorch, ONNX (Open Neural Network Exchange) Format - **Published:** August 4, 2026 - **Apply:** https://www.jobleads.com/es/job/e5eb448d696c91016b0a07cbfa27d3069 ## About the Role * Strong background in ASR/TTS. * PyTorch/ONNX experience. * Familiar with GPU profiling and optimisation. * Fluency in English. ## Description As Speech Model Performance Engineer, you'll work closely with Ismael and the founding tech team to shape the technical foundation of SLNG. From TTS voices to multilingual ASR, you'll benchmark, optimise, and productionize speech inference at scale., You'll Make Speech Models Fast. Example Initiatives * Quantise neural TTS and STT models to run with minimal latency on heterogeneous GPU hardware. * Benchmark ASR models across dialectal variations, measuring Word Error Rate (WER) and latency trade-offs. * Implement continuous batching and KV-cache reuse for streaming inference. * Profile GPU utilisation with CUDA kernels to identify bottlenecks in large-scale inference. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Developing the Rich Text Editor for DeepL.com](https://www.wearedevelopers.com/videos/1172-developing-the-rich-text-editor-for-deepl-com) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Making neural networks portable with ONNX](https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this)