> Markdown version of [/jobs/ext/2548362-senior-ai-engineer](https://www.wearedevelopers.com/jobs/ext/2548362-senior-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Engineer - **Company:** TMTG, LLC - **Location:** Sarasota, FL, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, C Sharp (Programming Language), Data Centers, Failover, Python (Programming Language), PostgreSQL, RabbitMQ, Ruby on Rails, Redis, Search Technologies, TypeScript, ReactJS, Large Language Models, Kotlin, Api Design, Golang - **Published:** August 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=4c8161ae710e3702 ## About the Role * 5+ years of backend engineering experience in a statically typed language (Go, Java, Kotlin, C#) * You've shipped production LLM-backed features - retrieval-augmented generation, streaming responses, tool use - and lived with them after launch * Hands-on experience with the LLM serving stack: routing across multiple model providers, failover, token streaming, and cost/usage metering * Experience building retrieval systems: vector search, embedding pipelines, context assembly, and citation-backed answers * You think in failure modes: hallucination, retrieval misses, provider outages, cost blowouts - and you build the instrumentation to catch them * Pragmatic about evaluation - you know how to measure whether AI answers are actually good (relevance, safety, source quality) with simple, repeatable tests, not just academic benchmarks * Strong API design instincts; comfortable owning a service end to end, from schema to deploy to on-call * US-based and authorized to work in the United States Nice to Have * Go (strongly preferred) * Python * Experience with LLM gateways or serving infrastructure (LiteLLM, vLLM, TGI, or similar) * Vector databases (Qdrant, pgvector) and embedding pipelines * Fine-tuning open-weight models (LoRA or full fine-tunes) and the eval discipline that goes with it * Content moderation or safety tooling experience* Familiarity with Ruby on Rails or React/TypeScript (you'll integrate with both) Our Stack Go, Ruby on Rails, Python, React/TypeScript, PostgreSQL, Redis, RabbitMQ. Services deployed across multiple data centers. ## Description You'll work directly with the platform architect and other developers to bring deep, hands-on LLM-stack expertise: you've built these systems before, you know where they break, and you know what "good" looks like in production. The work is greenfield. You won't be maintaining someone else's pipeline - you'll be standing up the core AI infrastructure for a platform with millions of active users, and shaping the engineering practices around it as the team grows. ## Related Videos - [LLMs in the wild: Building an AI agent that survives production](https://www.wearedevelopers.com/videos/100319-llms-in-the-wild-building-an-ai-agent-that-survives-production) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)