> Markdown version of [/jobs/ext/3006422-machine-learning-operations-mlops-engineer](https://www.wearedevelopers.com/jobs/ext/3006422-machine-learning-operations-mlops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Operations (MLOps) Engineer - **Company:** Gallatin Inc - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $80,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Continuous Integration, Federal Information Processing Standards (FIPS), Python (Programming Language), PostgreSQL, Network Installation Services, Azure Machine Learning, Management of Software Versions, Large Language Models, Amazon Virtual Private Cloud (VPC), Containerization, Kubernetes, Machine Learning Operations, TensorRT, Hardware Infrastructure, Feature Extraction - **Published:** September 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=458971f80d7b329d ## About the Role * 5+ years in MLOps, ML platform, or infrastructure engineering, including meaningful time working on systems with real users. * Strong Python skills and comfort in a production codebase, not just notebooks. * Deep Kubernetes and containerization experience, plus infrastructure as code. * Production experience with AWS ML/Azure infrastructure (SageMaker, EKS, or equivalent). * Hands-on GPU infrastructure experience: scheduling, utilization, memory sizing, and cost. * Hands-on experience deploying and operating production software in IL5 or IL6 environments, including disconnected or restricted-network deployments. Production ML Judgment * You have shipped an LLM or ML system to production and then had to keep it working. You know what breaks. * You have built evaluation and monitoring for ML systems rather than adopting a vendor dashboard and hoping. * You can reason about where a pipeline's quality actually comes from and say so when a metric is measuring the wrong thing., * You are comfortable with ambiguity and owning a domain end to end. This is a small team; there is no one to hand the pager to. * You are willing to learn the mission domain. The engineers who do best here become genuinely interested in the logistics problem itself., * LLM serving and inference optimization (vLLM, TensorRT-LLM, quantization, prefix caching). * Retrieval-grounded systems in production: hybrid retrieval, re-ranking, index freshness, and citation quality. * Edge, on-premises, or disconnected deployment. * Experience supporting ATO, continuous authorization, or production operations in classified environments. * Defense, intelligence, or another accredited or regulated environment. * Palantir Foundry, PostgreSQL/pgvector, NATS/JetStream, or ArgoCD. * A degree in CS, engineering, or a related technical field, or the equivalent built the hard way. Mission and Identity We are building the system that enables faster, smarter logistics decisions in contested environments, and we're doing it with a team of seasoned entrepreneurs, operators, and technologists who have built and scaled solutions in this space before. We hold ourselves to an extremely high standard. We value clear thinking, direct communication, and the kind of ownership that doesn't stop until something actually works. ## Description In this role, you will build the systems that move a model out of a notebook and into the hands of a planner. Sometimes, that means deploying to an air-gapped rack in a tent instead of a VPC. You will own the path from training run to deployed capability: the infrastructure it runs on, the release process that ships it, the evaluation harness that proves it works, and the telemetry that tells us when it stops working. Our AI/ML team works across retrieval-grounded systems for doctrine and logistics data, document and feature extraction, military symbol recognition, optimization and movement models, and an LLM agent platform. This role underpins that work: building the infrastructure, release processes, and evaluation systems that make it shippable and keep it honest in production. You will have plenty of room to shape how we build it. Training & Serving Infrastructure * Own model training, fine-tuning, and batch inference infrastructure across AWS (SageMaker, EKS) and on-premises GPU hardware. * Stand up and tune LLM inference serving: vLLM-class stacks, quantization, continuous batching, KV-cache and throughput sizing. Make the hosted-vs-local call with numbers behind it. * Build for DDIL: local inference with configurable fallback, degraded-mode behavior, and sane resource envelopes on hardware we do not get to choose. Release & Reproducibility * Build CI/CD for models and pipelines: versioned datasets, a model registry, promotion gates, and rollback that actually works under pressure. * Own infrastructure as code, containerization, and GitOps deployment across environments ranging from a dev cluster to a disconnected enclave. * Make reproducibility a hard requirement. Any result we put in front of a customer or evaluator must be reproducible from a commit and dataset version., * Build and own the evaluation harness: regression suites for retrieval and extraction pipelines, LLM-as-judge pipelines with measured judge-human agreement, and adversarial and held-out sets. * Instrument production for drift, latency, cost, retrieval quality, and failure modes, including quiet ones such as a retrieval miss that produces a fluent but wrong answer. * Make our metrics defensible to external test and evaluation reviewers. "We think it's good" is not a deliverable., * Own ingestion, versioning, and lineage for logistics and doctrinal data drawn from a heterogeneous set of authoritative sources. * Build and operate embedding and feature pipelines, incremental indexing, and the unglamorous systems that keep a retrieval index fresh. * Build the human-in-the-loop infrastructure: confidence-scored routing, review queues, and feedback capture that improves the next model., * Deploy and operate ML systems in IL5 and IL6 environments, including air-gapped or restricted-network enclaves. Build the release, observability, artifact-management, and incident-response workflows those environments require. * Support ATO and continuous-authorization work with implementation evidence tied to applicable security controls (NIST SP 800-171, NIST SP 800-53 Rev. 5, CMMC Level 2, FIPS 140-3, and RMF/eMASS). * Handle CUI and classified data correctly without being asked twice. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [MLOps - What’s the deal behind it?](https://www.wearedevelopers.com/videos/392-mlops-what-s-the-deal-behind-it) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)