SRE (AI/ML)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Kforce has a client in Maryland Heights, MO that is seeking a SRE (AI/ML) with hands-on AI/ML engineering experience to design, operate, and scale reliable cloud platforms that power production AI/ML and GenAI workloads., The SRE will own reliability, automation, and cost/performance of AWS infrastructure while partnering with data science and application teams to productionize models, ML pipelines, and LLM/RAG solutions. The role blends classic SRE and cloud administration with modern MLOps and LLMOps practices.
Requirements
- 8+ years total experience with at least 3+ years as SRE/DevOps/Cloud Admin on AWS in production
- Strong hands-on with core AWS services: EC2, VPC, IAM, S3, RDS, EKS/ECS, Lambda, CloudWatch, CloudTrail
- Proficient in Python and one of Bash; comfortable writing reusable modules and operational tooling
- Solid IaC experience with Terraform (preferred) and/or CloudFormation/CDK
- Kubernetes (EKS) in production: workloads, autoscaling, networking (CNI), and secrets
- Working knowledge of ML lifecycle: data prep, training, evaluation, deployment, and monitoring
- Hands-on with at least one of: SageMaker, Bedrock, Vertex AI, or Azure ML in production
- Experience operating LLM/GenAI or classical ML services with observability and guardrails
- Strong incident management, on-call, and postmortem culture experience
Nice to have:
- AWS certifications: Solutions Architect Professional, DevOps Engineer Professional, Security Specialty, or ML Specialty
- Experience with LangChain, LlamaIndex, Strands, or agentic frameworks in regulated environments
- Vector databases and retrieval tuning (hybrid search, re-ranking, chunking strategies)
- Model governance, responsible AI, PII/PCI handling, and audit-ready ML pipelines
- Exposure to Databricks, Snowflake, Redshift or Iceberg-based lakehouses
- GPU/accelerator experience (A10/A100/H100, Inferentia/Trainium) and cost-aware inference design
Benefits & conditions
The pay range is the lowest to highest compensation we reasonably in good faith believe we would pay at posting for this role. We may ultimately pay more or less than this range. Employee pay is based on factors like relevant education, qualifications, certifications, experience, skills, seniority, location, performance, union contract and business needs. This range may be modified in the future.
We offer comprehensive benefits including medical/dental/vision insurance, HSA, FSA, 401(k), and life, disability & ADD insurance to eligible employees. Salaried personnel receive paid time off. Hourly employees are not eligible for paid time off unless required by law. Hourly employees on a Service Contract Act project are eligible for paid sick leave.
Note: Pay is not considered compensation until it is earned, vested and determinable. The amount and availability of any compensation remains in Kforce’s sole discretion unless and until paid and may be modified in its discretion consistent with the law.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Dev Digest 120 - Apple and peers
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production