SRE (AI/ML)

Kforce Inc.
Maryland Heights, MO, United States
4 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Cloud Computing DevOps Identity and Access Management Python (Programming Language) Machine Learning Azure Machine Learning Large Language Models
+10 more
Snowflake Amazon Virtual Private Cloud (VPC) Cloudformation Amazon Relational Database Service Kubernetes Machine Learning Operations Cloudwatch Terraform Amazon Redshift Databricks

Job description

Kforce has a client in Maryland Heights, MO that is seeking a SRE (AI/ML) with hands-on AI/ML engineering experience to design, operate, and scale reliable cloud platforms that power production AI/ML and GenAI workloads., The SRE will own reliability, automation, and cost/performance of AWS infrastructure while partnering with data science and application teams to productionize models, ML pipelines, and LLM/RAG solutions. The role blends classic SRE and cloud administration with modern MLOps and LLMOps practices.

Requirements

  • 8+ years total experience with at least 3+ years as SRE/DevOps/Cloud Admin on AWS in production
  • Strong hands-on with core AWS services: EC2, VPC, IAM, S3, RDS, EKS/ECS, Lambda, CloudWatch, CloudTrail
  • Proficient in Python and one of Bash; comfortable writing reusable modules and operational tooling
  • Solid IaC experience with Terraform (preferred) and/or CloudFormation/CDK
  • Kubernetes (EKS) in production: workloads, autoscaling, networking (CNI), and secrets
  • Working knowledge of ML lifecycle: data prep, training, evaluation, deployment, and monitoring
  • Hands-on with at least one of: SageMaker, Bedrock, Vertex AI, or Azure ML in production
  • Experience operating LLM/GenAI or classical ML services with observability and guardrails
  • Strong incident management, on-call, and postmortem culture experience

Nice to have:

  • AWS certifications: Solutions Architect Professional, DevOps Engineer Professional, Security Specialty, or ML Specialty
  • Experience with LangChain, LlamaIndex, Strands, or agentic frameworks in regulated environments
  • Vector databases and retrieval tuning (hybrid search, re-ranking, chunking strategies)
  • Model governance, responsible AI, PII/PCI handling, and audit-ready ML pipelines
  • Exposure to Databricks, Snowflake, Redshift or Iceberg-based lakehouses
  • GPU/accelerator experience (A10/A100/H100, Inferentia/Trainium) and cost-aware inference design

Benefits & conditions

The pay range is the lowest to highest compensation we reasonably in good faith believe we would pay at posting for this role. We may ultimately pay more or less than this range. Employee pay is based on factors like relevant education, qualifications, certifications, experience, skills, seniority, location, performance, union contract and business needs. This range may be modified in the future.

We offer comprehensive benefits including medical/dental/vision insurance, HSA, FSA, 401(k), and life, disability & ADD insurance to eligible employees. Salaried personnel receive paid time off. Hourly employees are not eligible for paid time off unless required by law. Hourly employees on a Service Contract Act project are eligible for paid sick leave.

Note: Pay is not considered compensation until it is earned, vested and determinable. The amount and availability of any compensation remains in Kforce’s sole discretion unless and until paid and may be modified in its discretion consistent with the law.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all