> Markdown version of [/jobs/ext/3053940-devopsengineer](https://www.wearedevelopers.com/jobs/ext/3053940-devopsengineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOpsEngineer - **Company:** LEO Inc - **Location:** Wellington, FL, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Computer-Aided Design, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Audit Trail, Microsoft Azure, Bash Shell, Cloud Computing, Computer Clusters, Nvidia CUDA, Computer Networks, Continuous Integration, Data as a Services, DevOps, Federal Information Processing Standards (FIPS), Github, Graph Database, Identity and Access Management, Python (Programming Language), Key Management, Neo4j, Network Segmentation, Octopus Deploy, OpenStack, Query Optimization, Regression Testing, Prometheus, Software Deployment, Management of Software Versions, Datadog, Pulumi, Graphics Processing Unit (GPU), Large Language Models, Grafana, HybridCloud, Gitlab-ci, Tanzu, Kubernetes, Low Latency, Machine Learning Operations, TensorRT, Hardware Infrastructure, Terraform, Jenkins, Vulnerability Analysis, Vmware - **Published:** September 24, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2984267373-2 ## About the Role * 5+ years in DevOps, SRE, or platform engineering, with at least 2 years operating ML or AI workloads in production. * Strong fluency with Kubernetes, container orchestration, and infrastructure-as-code (Terraform, Pulumi, or equivalent). * Hands-on experience deploying LLM inference at scale - you know the tradeoffs between vLLM, TGI, Triton, and managed APIs, and when to use which. * Solid Python skills for tooling, automation, and glue code; comfort with Bash and at least one systems language is a plus. * Experience operating GPU infrastructure (NVIDIA drivers, CUDA, MIG, GPU operator, scheduling) in either cloud (A10/A100/H100 instances) or on-prem environments. * Production experience with CI/CD (GitHub Actions, GitLab CI, Jenkins, ArgoCD) and GitOps patterns. * Strong security posture: IAM, secrets management, network segmentation, vulnerability scanning, supply-chain security (SBOMs, signed artifacts, SLSA). * Experience with observability stacks (Prometheus, Grafana, OpenTelemetry, Loki, Elastic, Datadog) and applying them to ML systems. * Demonstrated ability to work with sensitive data and operate within compliance frameworks. Nice to Have * Direct experience deploying systems in CJIS, FedRAMP, IL4/5, or equivalent regulated environments. * Experience with air-gapped or cross-domain deployments. * Familiarity with LLM-specific tooling: LangSmith, Langfuse, Helicone, Phoenix, Weights & Biases, MLflow. * Vector and graph database operations at scale - sharding, replication, backup, query tuning. * Experience with FedRAMP-authorized cloud regions (AWS GovCloud, Azure Government, GCC High) or on-prem cloud (OpenStack, VMware Tanzu). * Familiarity with model and prompt evaluation in CI - automated guardrails, regression tests against curated eval sets. * Experience with policy-as-code (OPA, Kyverno, Sentinel) and admission controllers. * Background supporting law enforcement, corrections, intelligence, or defense missions. * Active or recent security clearance. * Familiarity with 28 CFR Part 23, CJIS Security Policy, NIST 800-, or FISMA controls. ## Description We're hiring a DevOps / SRE to deploy, operate, and harden the AI systems that support corrections operations and intelligence analysis. Our data scientists build LLM-powered agents, RAG pipelines, and ontology-driven analytics - your job is to make sure those systems run reliably, securely, and auditably in environments where uptime, data segregation, and chain-of-custody actually matter. You'll own the path from a trained model or agent prototype to a production system that analysts depend on, in infrastructure that meets CJIS, FedRAMP, or equivalent standards. What You'll Do * Design and operate the deployment platform for LLM applications, agentic systems, RAG pipelines, and supporting data services across cloud, on-prem, and air-gapped environments. * Build CI/CD pipelines for model and application delivery - including model registries, prompt and config versioning, evaluation gates, and rollback paths. * Stand up and maintain inference infrastructure: GPU clusters, model serving (vLLM, TGI, Triton, Ollama, TensorRT-LLM), vector databases (pgvector, Weaviate, Qdrant, Milvus), and graph databases (Neo4j, Neptune). * Operate Kubernetes (EKS, AKS, GKE, or on-prem) as the backbone for AI workloads, with GPU scheduling, autoscaling, and workload isolation. * Implement observability for AI systems specifically - not just CPU and latency, but token throughput, model drift, agent trace logs, tool-call success rates, retrieval quality, and cost per request. * Harden environments to meet CJIS, FedRAMP Moderate/High, StateRAMP, or DoD IL4/5 controls as applicable - encryption at rest and in transit, key management, audit logging, FIPS-validated crypto, and boundary controls. * Enforce data segregation, classification boundaries, and need-to-know access through network policy, IAM, and secrets management (Vault, AWS Secrets Manager, KMS/HSM). * Build deployment patterns for air-gapped or classified enclaves - including offline model distribution, signed artifacts, and dependency mirroring. * Manage incident response for AI systems: runbooks, on-call rotations, blameless postmortems, and the special failure modes that come with LLMs (hallucination spikes, prompt injection, retrieval poisoning, runaway tool loops). * Partner with data scientists, security, and compliance teams to ship safely - and push back when a deploy would compromise security or reliability. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [WebAssembly: The Next Frontier of Cloud Computing](https://www.wearedevelopers.com/videos/972-webassembly-the-next-frontier-of-cloud-computing) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [The Private AI Platform: Why Agentic Apps Need a Private Application Platform](https://www.wearedevelopers.com/videos/100162-the-private-ai-platform-why-agentic-apps-need-a-private-application-platform) - [Generating code with Angular schematics](https://www.wearedevelopers.com/videos/129-generating-code-with-angular-schematics) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)