AWS AgentCore Platform Engineer - 67417
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
As a Senior AWS AgentCore Platform Engineer, you’ll work alongside Cloud Architects, AI Engineers, Platform Engineers, and Security specialists to establish scalable, secure, and observable AI platforms. You’ll play a critical role in defining the operational foundation that enables enterprise teams to deploy AI agents with confidence, governance, and efficiency.
This is an exciting opportunity to shape enterprise AI infrastructure, drive innovation in LLMOps, and influence platform standards across multiple business units., * Design and implement enterprise-grade observability solutions for AI agent ecosystems built on AWS Bedrock, AgentCore, and MCP servers.
- Assess and optimize CloudWatch, X-Ray, Bedrock logging, and AgentCore tracing capabilities against agentic workflow requirements.
- Conduct gap analyses and implement observability solutions using Dynatrace and other monitoring platforms.
- Develop distributed tracing frameworks for AI workloads, including:
- LLM decision paths
- Tool invocations
- Sub-agent interactions
- MCP server communications
- Build structured logging frameworks to support troubleshooting, governance, and performance optimization.
- Design post-deployment validation pipelines for AI agents and MCP servers, including deployment health monitoring and registration verification.
Cost Governance & Optimization
- Architect cost visibility and governance frameworks across AI workloads.
- Extend cloud tagging strategies to include agent runtimes, vector databases, MCP services, and Bedrock token consumption.
- Develop cost allocation models to provide spending transparency by team, department, and application.
- Build dashboards and reporting solutions for AI platform cost tracking and forecasting.
- Configure AWS Budgets, automated alerts, anomaly detection, and optimization recommendations.
- Deliver automated cost reporting through Microsoft Teams and email channels.
Monitoring & Incident Management
- Define enterprise monitoring standards and alerting frameworks across AI platform services.
- Create and manage P1-P4 alerting strategies covering:
- Deployment failures
- Runtime exceptions
- Tool invocation errors
- MCP connectivity issues
- Integrate monitoring and notification workflows with Microsoft Teams and email.
- Develop operational runbooks and self-service documentation within Confluence.
- Evaluate AWS-native and third-party monitoring solutions and recommend target-state architectures.
Security & Platform Governance
- Assess IAM architectures and multi-team access models for enterprise-scale AI environments.
- Design Attribute-Based Access Control (ABAC) frameworks to support secure multi-team isolation.
- Evaluate Cedar policy engine capabilities within AgentCore for fine-grained authorization models.
- Develop reusable Terraform modules to enforce governance, security, and compliance standards.
- Identify scalability risks and implement secure platform design patterns for enterprise AI adoption.
Platform Engineering & Automation
- Build and maintain Infrastructure-as-Code solutions using Terraform.
- Design and enhance CI/CD pipelines supporting AI platform deployments.
- Collaborate with engineering, security, architecture, and business stakeholders in Agile environments.
- Drive platform standardization, automation, and operational excellence initiatives., We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We’re also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We’re always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you’ll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.
Requirements
- 8+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud services including:
- IAM
- CloudWatch
- AWS Lambda
- AWS Bedrock
- Cloud-native monitoring and governance services
- Hands-on experience implementing observability and distributed tracing solutions using tools such as:
- Dynatrace
- Jaeger
- Honeycomb
- OpenTelemetry
- Experience designing and managing Infrastructure-as-Code using Terraform.
- Strong background building and maintaining CI/CD pipelines in enterprise environments.
- Experience working in Agile teams utilizing Microsoft Teams, Confluence, and modern collaboration tools.
Preferred Qualifications
- Experience supporting AI, Generative AI, or LLM-based platforms.
- Familiarity with AgentCore, LangChain, LangFuse, LiteLLM, MCP servers, or similar AI orchestration frameworks.
- Understanding of LLM lifecycle management, prompt execution flows, token consumption tracking, and AI workload optimization.
- Knowledge of cloud cost management, FinOps practices, and governance frameworks.
- Experience designing enterprise-scale security architectures using ABAC and policy-based authorization models.
- Strong analytical and problem-solving skills with the ability to translate complex technical challenges into scalable platform solutions.
Success Factors
- Passion for emerging AI technologies and cloud-native engineering.
- Ability to balance reliability, security, performance, and cost optimization.
- Strong communication skills with the ability to influence technical and business stakeholders.
- Proven ability to lead platform modernization initiatives and establish engineering best practices.
About the company
We’re Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world’s potential. We’re people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what’s now to what’s next. We make it happen through the power of acceleration.
Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don’t expect you to ‘fit’ every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us., We’re a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you., Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.
We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What is Agentic Programming and Why Should Developers Care?
How to Become an AI Engineer
Stephan Gillich - Bringing AI Everywhere
The Overflow: AI and Agentic Coding