> Markdown version of [/jobs/ext/166746-senior-ml-engineer](https://www.wearedevelopers.com/jobs/ext/166746-senior-ml-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Engineer - **Company:** Invoca, Inc. - **Location:** Austin, TX, United States (Remote available) - **Experience:** Expert - **Salary:** $152,000.0 - $228,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Monitoring of Systems, Python (Programming Language), Machine Learning, Language Modeling, Management of Software Versions, Pytorch, Large Language Models, Deep Learning, Model Validation, Data Lakes, Kubernetes, Information Technology, Low Latency, Deployment Automation, HuggingFace, Machine Learning Operations, Hardware Infrastructure, Spacy, GPT, Data Pipelines, Software Library - **Published:** May 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e73e8fd2c508a14c ## About the Role Do you have experience in Machine learning libraries?, Do you have a Bachelor's degree?, * 5+ years of ML Engineering experience with a strong production focus * Advanced Python and deep learning proficiency (PyTorch, HuggingFace Transformers, spaCy) * Demonstrated track record deploying and maintaining transformer-based NLP models in production * Hands-on experience fine-tuning SLMs/LLMs (LoRA, QLoRA, PEFT) and optimizing models via quantization, batching, and throughput tuning * Proficiency with inference infrastructure: Triton, Baseten, vLLM, TGI, SageMaker, Vertex AI, or similar * Experience building production-grade APIs that expose ML models to downstream consumers * Familiarity with MLOps tooling, model monitoring, and eval platforms (Braintrust, MLflow, or equivalent) * B.S. in Computer Science, Engineering, Statistics, or equivalent; advanced degree a plus * Familiarity with RLHF or preference training is a bonus, Candidates must be based within ~2 hour drive of these areas. Occasional business travel may be required. ## Description We're hiring a Senior ML Engineer to own the productionization layer of Invoca's ML stack - model serving, inference optimization, fine-tuning, and the APIs and pipelines that tie it all together. You'll be a primary driver of the infrastructure powering our Context Engine and agentic AI workflows, working closely with Data Scientists, Data Engineers, and Applied AI Engineers. Core Focus & Primary Ownership * Lead End-to-End MLOps and Productionization: Architect, implement, and maintain CI/CD pipelines for ML artifacts - including model evaluation, versioning, and automated deployment. Serve as the primary SME for operational excellence across the Invoca ML stack. * Design and Optimize SLM/LLM Deployment: Own the full inference infrastructure: model serving on Triton Inference Server, Baseten, and Kubernetes-based GPU infrastructure. Profile and tune for low latency and high throughput, and build robust, scalable APIs for internal and external model access. Broader Contributions * Fine-Tune Language Models: Apply parameter-efficient fine-tuning methods (LoRA, QLoRA, PEFT) to adapt transformer-based SLMs and LLMs for high-impact NLP applications in conversation intelligence. * Evolve ML Infrastructure: Contribute to model training infrastructure, data pipelines, and data lake foundations to keep the systems powering our models reliable and scalable. * Collaborate Across Teams: Partner closely with Data Scientists, Data Engineers, and Applied AI Engineers to build the foundational ML systems behind Invoca's agentic AI products. * Deliver Customer Value: Work with product and engineering to understand customer needs and ship ML solutions that make a measurable difference. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Speak, Code, Deploy: Transforming Developer Experience with Voice Commands](https://www.wearedevelopers.com/videos/1159-speak-code-deploy-transforming-developer-experience-with-voice-commands) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)