> Markdown version of [/jobs/ext/1469249-staff-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1469249-staff-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Engineer - **Company:** NATIVE AI LLC - **Location:** Santa Barbara, United States (Remote available) - **Salary:** $200,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Cloud Computing, Cloud Engineering, Continuous Integration, Python (Programming Language), Machine Learning, Language Modeling, Open Source Technology, Performance Tuning, Azure Machine Learning, Software Safety, Systems Integration, Unstructured Data, Alwayson, Graphics Processing Unit (GPU), Autoscaling, Large Language Models, Deep Learning, AWS ECS, Low Latency, Machine Learning Operations, TensorRT, Data Pipelines, Docker - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/2064ca23-8f94-4faf-ab88-7108ed831709 ## About the Role * Systems thinker: You think in terms of platforms and long-term leverage, not just features. * Production builder: You've built and scaled ML infrastructure in production with meaningful business impact. * Ambiguity: You operate effectively in high ambiguity, turning unclear infra problems into clear direction. * Owner-operator: You take ownership with a founder/owner-operator mindset, act with urgency, and focus on outcomes. * Pace: You have a strong desire to move fast and deliver impact, while maintaining sound engineering judgment. * Collaboration: You are humble, collaborative, and low-ego, and you elevate those around you. * Sustainability: You value work-life balance as a foundation for sustained high performance. * Reliability mindset: You treat ML infra like any other production system - SLOs, on-call, observability, postmortems. Must Have * ML infra at scale: Has built and operated production ML infrastructure on AWS - ECS, SageMaker, GPUs, autoscaling, and cost controls. * Inference platforms: Production experience with model serving for both LLMs and custom models; understands quantization, batching, and routing. * Provider breadth: Direct experience integrating with Google (Vertex / Gemini), OpenAI, and Anthropic APIs in production. * Training capability: Has trained or fine-tuned language models end-to-end; comfortable with deep learning, evaluation, and inference. * Cloud-native engineering: Strong Python, Docker, dependency management, and CI/CD for AI workloads. * RAG & agents: Working knowledge of LangChain / LangGraph and modern RAG patterns over structured and unstructured data. * Cost optimization: Demonstrated experience reducing unit cost of AI workloads without regressing quality or latency. * AI safety & authorization: Hands-on experience operating AI guardrails, scoped tool permissions, and authorization layers for production AI systems. Nice to Have * Experience training Small Language Models for production use. * GPU performance tuning (vLLM, TensorRT, Triton, or similar). * Prior Staff-level role at a company with a significant AI infra footprint. * Experience with ontology-driven systems or knowledge graphs supporting AI applications. * Contributions to open-source ML infrastructure or LLM tooling. ## Description We're hiring a Staff Machine Learning Engineer to help move forward the ML platform that every AI initiative at AppFolio depends on - training, fine-tuning, inference, RAG, evaluation, and cost. You'll keep our AI cloud always-on, observable, and economical, while staying close enough to applications to influence model and agent design. This role works at the intersection of ML infrastructure, applied AI, and cost discipline. You'll partner closely with our Voice & Agents and Research ML engineers to harden their prototypes into production systems, and help move forward the platform layer that lets Realm-X scale across AppFolio's entire customer base. Your Impact * ML Platform: Design and operate AppFolio's ML infrastructure on AWS - ECS, SageMaker, GPU fleets, model serving, autoscaling, and cost controls. * Drive AI Cost Discipline: Optimize cost across all AI applications - provider routing, caching, batch vs. real-time, model size selection, and inference economics. * Multi-Provider Reliability: Maintain reliable, multi-provider LLM access across Google, OpenAI, and Anthropic with sensible fallbacks and abstractions. * Training & Fine-Tuning Stack: Build the training and fine-tuning stack for Small Language Models, including data pipelines, GPU orchestration, and evaluation. * Productionize Research: Partner with Voice & Agents and Research ML engineers to harden their prototypes into production systems with SLOs, on-call rotations, and observability. * AI Safety & Guardrails: Operate AppFolio's AI safety and authorization layer - guardrails on AWS, scoped tool permissions, and human-in-the-loop gates for autonomous agent actions. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Celery on AWS ECS - the art of background tasks & continuous deployment](https://www.wearedevelopers.com/videos/561-celery-on-aws-ecs-the-art-of-background-tasks-continuous-deployment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tips, Techniques, and Common Pitfalls Debugging Kafka](https://www.wearedevelopers.com/videos/838-tips-techniques-and-common-pitfalls-debugging-kafka) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)