AWS Cloud SRE - AI/ML Platform

Brooksource
Maryland Heights, United States
7 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$145,600.0 - $166,400.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Cloud Computing Cloud Engineering Continuous Integration DevOps Fraud Prevention and Detection Graph Database Identity and Access Management
+15 more
Python (Programming Language) Key Management Machine Learning Reliability Engineering Azure Machine Learning Scripting Large Language Models Amazon Virtual Private Cloud (VPC) Amazon Relational Database Service Containerization Kubernetes Machine Learning Operations Cloudwatch Terraform Serverless Computing

Job description

Brooksource is seeking a Senior AWS Cloud Site Reliability Engineer (SRE) - AI/ML Platform to join a high-impact Fraud Prevention AI/ML Engineering Team within a Fortune 500 telecommunications client. In this contract role, you will own and enhance AWS infrastructure, automation, reliability, and production environments supporting modern AI/ML and GenAI workloads. You will operate at the intersection of traditional SRE/cloud engineering and modern MLOps/LLMOps, enabling teams to move AI applications and models into reliable, scalable production environments. The broader program includes agentic AI and anomaly-detection capabilities, leveraging technologies such as SageMaker, Bedrock, graph-based solutions, RAG, and other machine-learning services. This engagement is expected through the end of 2026, with potential for extension based on performance and business needs. You will have the opportunity to contribute to a high-visibility, new-build AI/ML initiative and work alongside talented engineers in a dynamic, relationship-driven environment.

WHAT YOU’LL DO

  • Design, operate, and support scalable AWS infrastructure powering production AI/ML, GenAI, and anomaly-detection workloads
  • Administer and automate AWS services across networking, compute, identity, storage, containers, serverless, monitoring, and security
  • Collaborate with AI/ML and application engineers to productionize models, ML pipelines, LLM/RAG solutions, and other AI-driven services
  • Build and maintain infrastructure automation using Terraform and CI/CD, supporting containerized workloads on Kubernetes/EKS and related services
  • Enhance platform reliability through monitoring, observability, incident response, troubleshooting, scaling, operational automation, and post-deployment validation
  • Support AI/ML platform capabilities involving SageMaker, Bedrock, graph databases, vector databases, and retrieval-based architectures
  • Ensure operational controls and best practices for production AI/ML, GenAI, and classical ML services
  • Take ownership of technical requirements from implementation through integration, deployment, and ongoing support

Requirements

  • 8+ years of overall technical experience, including 3+ years in SRE, DevOps, Cloud Administration, or Cloud Engineering roles within production AWS environments
  • Deep hands-on AWS experience with services such as EC2, VPC, IAM, S3, RDS, EKS/ECS, Lambda, CloudWatch, and CloudTrail
  • Strong Infrastructure-as-Code and automation skills using Terraform and scripting (Python, Bash, or similar)
  • Hands-on experience operating Kubernetes/EKS workloads, including scaling, networking, deployment, and secrets management
  • Working knowledge of the machine-learning lifecycle and production experience with SageMaker, Bedrock, or comparable AI/ML platforms
  • Experience supporting production AI/ML, GenAI, LLM/RAG, or classical ML services with observability, reliability, and operational controls
  • Demonstrated ownership mindset, taking technical requirements through implementation, integration, deployment, and ongoing support

Benefits & conditions

  • 401(k) Matching Plan
  • Medical, Dental, & Vision Plans
  • Relationship-driven process to find your best fit
  • 6 Paid Holidays
  • Regular meetings to ensure quality in your engagement
  • Employee Assistance Program (EAP) with virtual counseling, financial and legal services, life coaching, and more
  • Flexible, supportive work environment within a nationally recognized IT and Engineering services provider

About the company

At Brooksource, relationships are the foundation of everything we do. Since 2000, we’ve built lasting partnerships with clients, consultants, and internal teams to deliver an exceptional experience across every engagement. As a trusted IT and Engineering services provider, Brooksource supports Fortune 500 organizations through Experience-Driven Staffing, Professional Services, and Elevate, our proprietary Workforce Transformation program. We offer flexible hiring models, including contract, contract-to-hire, and direct placement to meet your evolving business needs. Brooksource is a certified partner of leading platforms, including Salesforce, AWS, Microsoft, and Google Cloud, enabling us to deliver scalable, end-to-end technology solutions. With a growing national footprint, Brooksource is redefining expectations in IT consulting, engineering services, and technology workforce solutions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

4:21 min

Scaling operations using Azure AI Foundry tools

Maxim Salnikov Maxim Salnikov · World Congress 2025

Videos

See all

Related articles

See all