MLOps Engineer - AI/ML Systems Deployment (TS/SCI Preferred)

Rackner, Inc.
Cleveland, OH, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Computer Vision Cloud Computing Cloud Engineering Program Optimization Computer Programming Continuous Integration Monitoring of Systems Python (Programming Language) Machine Learning
+11 more
Prometheus Software Engineering Delivery Pipeline Grafana SC Clearance Build Management Containerization Kubernetes Build Tools Machine Learning Operations Docker

Job description

Rackner is hiring an MLOps Engineer to move AI/ML systems from prototype deployment operational use in a secure, mission-focused environment.

This is not a research role-this is where models become reliable, repeatable, auditable systems that run in real-world conditions.

This role is ideal for engineers who want to:

  • Work across AI/ML, Kubernetes, infrastructure, and mission systems
  • Own deployed systems, not just experiments
  • Build high-demand MLOps expertise in secure and constrained environments
  • Deliver technology that is used, trusted, and operational

You will help operationalize AI/ML capabilities where reliability, performance, and trust matter most.

What You’ll Do

Operationalize AI/ML Systems

  • Deploy AI/ML models and ML-enabled applications into secure, real-world environments
  • Move workflows from experimentation into containerized, repeatable deployment pipelines
  • Support batch and real-time inference architectures
  • Bridge model development, software engineering, and platform operations

Own the ML Lifecycle

  • Build and operate production-grade ML pipelines
  • Support model versioning, lineage, reproducibility, and lifecycle governance
  • Work with tools such as MLflow, Kubeflow, Airflow, Argo, ClearML, or similar platforms

Build Cloud-Native ML Infrastructure

  • Deploy and support Kubernetes-based ML workloads
  • Containerize models, pipelines, and services using Docker or similar tools
  • Support CI/CD, automation, and repeatable deployment patterns for AI/ML systems

Engineer for Reliability

  • Monitor model and system performance after deployment
  • Support observability using tools such as Prometheus, Grafana, OpenTelemetry, or similar
  • Detect and resolve issues related to latency, reliability, drift, degradation, or resource usage

Support Secure and Constrained Environments

  • Help deploy AI/ML systems in secure, CAC-enabled, or constrained environments
  • Support limited compute, restricted data, degraded connectivity, and other operational constraints
  • Optimize systems for reliability and usability beyond ideal lab conditions

Create Repeatable Systems

  • Develop runbooks, deployment documentation, and operational playbooks
  • Build systems that can be understood, maintained, and operated by others

Requirements

Work Arrangement: On-site preferred; remote may be considered for highly aligned, clearance-ready candidates able to support secure / CAC-enabled environments and travel as needed Clearance: Active TS/SCI strongly preferred; active Secret may be considered for upgrade Requirement: U.S. citizenship required

Build and Deploy Real-World AI Systems, * U.S. citizenship

  • Background in deploying ML systems, AI-enabled applications, or production software
  • Strong programming skills in Python
  • Hands-on work with Docker, containers, or containerized deployment
  • Familiarity with Kubernetes or cloud-native environments
  • Understanding of CI/CD, automation, or pipeline-based delivery
  • Clear communication of technical decisions, tradeoffs, and ownership
  • Ability to operate in a CAC-enabled or secure environment, * Active TS/SCI clearance
  • Active Secret clearance with eligibility for upgrade
  • Familiarity with ML lifecycle tools such as MLflow, Kubeflow, Airflow, Argo, ClearML, or similar
  • Background in model serving, inference APIs, or deploying ML systems in production
  • Exposure to LLMs, transformer-based models, computer vision, NLP, or applied AI solutions
  • Hands-on work with Kubernetes-based ML workloads
  • Knowledge of observability and monitoring tools such as Prometheus, Grafana, or OpenTelemetry
  • Experience in DoD, defense, intelligence, regulated, or mission-critical settings
  • Work in edge, offline, air-gapped, low-bandwidth, D-DIL, or limited-compute environments

Clearance Requirements

  • Active TS/SCI clearance strongly preferred
  • Candidates with an active Secret clearance may be considered and supported for upgrade
  • Candidates without an active clearance must be:
  • U.S. citizens
  • eligible to obtain and maintain a clearance
  • able to work in a CAC-enabled or secure environment

About the company

Rackner is a software consultancy that builds cloud-native solutions for startups, enterprises, and the public sector. We are an energetic, growing team focused on solving complex problems through:

  • Distributed systems
  • DevSecOps
  • AI/ML
  • Cloud-native architecture

Our approach is cloud-first, cost-effective, and outcome-driven, delivering systems that scale and perform in real-world environments.

Benefits & Perks

  • 100% covered certifications & training aligned to your role
  • 401(k) with 100% match up to 6%
  • Highly competitive PTO
  • Comprehensive Medical, Dental, Vision coverage
  • Life Insurance + Short & Long-Term Disability
  • Home office & equipment plan
  • Industry-leading weekly pay schedule

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:19 min

Introduction to DevOps for AI and MLOps

Aarno Aukia · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all