Senior AWS AgentCore Platform Engineer

Halian
Reading, PA, United States
2 days ago
Apply on www.disabledperson.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Continuous Integration DevOps Monitoring of Systems Identity and Access Management Reliability Engineering Data Logging Large Language Models Generative AI AI Platforms Infrastructure Automation Frameworks
+7 more
Deployment Automation Machine Learning Operations Virtual Agents Cloud Optimization Cloudwatch Terraform Dynatrace

Job description

Are you passionate about building enterprise-scale AI platforms on AWS? We are seeking a Senior AWS AgentCore Platform Engineer to help strengthen and scale a next-generation AI Agent platform built on Amazon Bedrock AgentCore.

In this role, you will be responsible for designing and implementing the infrastructure, observability, security, monitoring, and cost governance capabilities that enable multiple teams to develop, deploy, and operate AI agents securely and efficiently.

What You’ll Be Doing

AI Platform Observability & Reliability

  • Design and implement observability solutions across AWS AgentCore environments.
  • Build distributed tracing and structured logging for AI agent workflows.
  • Monitor LLM interactions, tool usage, MCP integrations, and agent execution paths.
  • Evaluate and implement monitoring solutions including CloudWatch, Dynatrace, and AWS-native services.
  • Develop deployment validation pipelines and operational health checks.

Monitoring & Incident Management

  • Create and manage alerting strategies for platform and agent workloads.
  • Define operational runbooks and troubleshooting procedures.
  • Integrate monitoring alerts with Microsoft Teams and email notification systems.
  • Improve platform reliability using SRE best practices.

Cloud Cost Governance

  • Build cost visibility and reporting frameworks for AI workloads.
  • Monitor Bedrock token consumption and usage trends.
  • Implement tagging strategies, AWS Budgets, and anomaly detection.
  • Create dashboards showing AI platform spend by team and department.

Security & Access Control

  • Design scalable IAM and ABAC access control frameworks.
  • Implement secure multi-team isolation models.
  • Evaluate and deploy Cedar-based authorization strategies.
  • Develop reusable Terraform modules for security and infrastructure automation.

Infrastructure Automation

  • Build Infrastructure-as-Code solutions using Terraform.
  • Automate provisioning and management of AgentCore, Bedrock, and AWS platform resources.
  • Support CI/CD pipelines and platform engineering initiatives.

Requirements

  • AWS Cloud Platform
  • Amazon Bedrock
  • AgentCore (or AI/Agentic Platform Experience)
  • Terraform
  • DevOps / Site Reliability Engineering (SRE)
  • AWS IAM
  • CloudWatch
  • MLOps
  • SageMaker
  • Observability & Monitoring
  • Infrastructure as Code

Preferred Skills

  • Dynatrace
  • LangFuse
  • LiteLLM
  • MCP (Model Context Protocol)
  • Vector Databases
  • FinOps / Cloud Cost Optimization
  • CI/CD Automation
  • Enterprise AI Platforms

Required Experience

  • 7+ years of AWS Cloud Engineering experience
  • 5+ years working in DevOps, Platform Engineering, or SRE environments
  • Experience supporting production cloud platforms at enterprise scale
  • Hands-on experience with Terraform and AWS IAM
  • Experience supporting MLOps, Generative AI, or AI platform initiatives

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.disabledperson.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:32 min

Overview of Terraform and Terraform Cloud features

Devlin Duldulao · LIVE

Videos

See all

Related articles

See all