Cloud Platform Engineer (Agentic Ai)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
The project is for one of the world’s famous science and technology companies in pharmaceutical industry, supporting initiatives in AWS, AI and data engineering, with plans to launch over 20 additional initiatives in the future.We are seeking a highly skilled Cloud Engineer to lead the infrastructure design, deployment, and operations of the AI agent orchestration platform on AWS.This role is responsible for building and managing a Kubernetes-native, enterprise-grade platform that supports scalable AI agent workloads across development, QA, and production environments.ResponsibilitiesDesign, provision, and manage AWS infrastructure using Terraform, aligned with the AWS Well-Architected Framework.Core services include:Amazon EKSVPCIAMRoute 53Own and operate EKS clusters end-to?endManaged node group lifecycle management:Karpenter-based autoscalingCluster add?on lifecycle upgradesIRSA (IAM Roles for Service Accounts) configurationMulti-AZ high availability and resilienceCI/CD & GitOpsBuild and maintain automated deployment pipelines using:GitHub ActionsArgoCD (GitOps)Implement release strategies:Canary releasesSecurity & ComplianceIntegrate AWS-native security and governance controls:AWS WAFGuardDutySecurity HubKMS (encryption)Secrets ManagerExternal Secrets OperatorEnforce policy controls using:Observability & MonitoringImplement and manage observability stack:Amazon Managed PrometheusAmazon Managed GrafanaCloudWatch Container InsightsAWS X?Ray (distributed tracing)AI/ML IntegrationLeverage AWS AI/ML services to support agent orchestration:Cost Optimization (FinOps)Spot InstancesSavings PlansKarpenter bin?packing strategiesScheduled scale?to?zero for non?production environmentsPlatform & Engineering CollaborationPartner with platform and ML teams to:Integrate MCP servers and execution frameworksSupport extensibility of the agent ecosystemSkillsMust have4+ years of hands?on AWS experienceAWS Certifications:Required: AWS Solutions Architect (Associate or Professional)Preferred: DevOps Engineer, Security Specialty Kubernetes & EKS ExpertiseStrong hands?on experience with:EKS cluster provisioning and operationsManaged node groups and KarpenterKubernetes RBAC and network policies Infrastructure as Code (Terraform)Advanced Terraform capabilities:Remote state management (S3 + DynamoDB)Security scanning (Checkov, tfsec) AWS Services ProficiencyDeep knowledge of:EKS, ECR, ALB, Route 53, ACMIAM, KMS, Secrets ManagerIAM Identity CenterCloudTrail, AWS ConfigGuardDuty, Security Hub, AWS WAF AI/ML ExposurePractical experience with:SageMaker (model deployment and endpoints)Comprehend (NLP and PII detection) DevOps & IdentityExperience with:GitOps tools (ArgoCD or Flux)CI/CD pipelines for container workloadsGitHub Actions ? AWSEKS OIDC provider integration Observability & DebuggingFamiliarity with:OpenTelemetryAWS X?RayStrong understanding of:Pod Security StandardsAdmission webhooksService account least?privilege principlesNice to haveExperience with AI agent frameworks:LangChain, Claude Agent SDK, or similarKnowledge of emerging protocols:A2A (Agent?to?Agent)Familiarity with:Amazon Bedrock Agents, Knowledge Bases, GuardrailsNamespace isolationProgramming/debugging skills:Python, Go, or Node.jsAWS Cost ExplorerLanguagesThe project is for one of the world’s famous science and technology companies in pharmaceutical industry, supporting initiatives in AWS, AI and data engineering, with plans to launch over 20 additional initiatives in the future.We are seeking a highly skilled Cloud Engineer to lead the infrastructure design, deployment, and operations of the AI agent orchestration platform on AWS.This role is responsible for building and managing a Kubernetes-native, enterprise-grade platform that supports scalable AI agent workloads across development, QA, and production environments.ResponsibilitiesDesign, provision, and manage AWS infrastructure using Terraform, aligned with the AWS Well-Architected Framework.Core services include:Amazon EKSVPCIAMRoute 53Own and operate EKS clusters end?to?endManaged node group lifecycle management:Karpenter-based autoscalingCluster add?on lifecycle upgradesIRSA (IAM Roles for Service Accounts) configurationMulti-AZ high availability and resilienceCI/CD & GitOpsBuild and maintain automated deployment pipelines using:GitHub ActionsArgoCD (GitOps)Implement release strategies:Canary releasesSecurity & ComplianceIntegrate AWS-native security and governance controls:AWS WAFGuardDutySecurity HubKMS (encryption)Secrets ManagerExternal Secrets OperatorEnforce policy controls using:Observability & MonitoringImplement and manage observability stack:Amazon Managed PrometheusAmazon Managed GrafanaCloudWatch Container InsightsAWS X?Ray (distributed tracing)AI/ML IntegrationLeverage AWS AI/ML services to support agent orchestration:Cost Optimization (FinOps)Spot InstancesSavings PlansKarpenter bin?packing strategiesScheduled scale?to?zero for non?production environmentsPlatform & Engineering CollaborationPartner with platform and ML teams to:Integrate MCP servers and execution frameworksSupport extensibility of the agent ecosystemSkillsMust have4+ years of hands?on AWS experienceAWS Certifications:Required: AWS Solutions Architect (Associate or Professional)Preferred: DevOps Engineer, Security Specialty Kubernetes & EKS ExpertiseStrong hands?on experience with:EKS cluster provisioning and operationsManaged node groups and KarpenterKubernetes RBAC and network policies Infrastructure as Code (Terraform)Advanced Terraform capabilities:Remote state management (S3 + DynamoDB)Security scanning (Checkov, tfsec) AWS Services ProficiencyDeep knowledge of:EKS, ECR, ALB, Route 53, ACMIAM, KMS, Secrets ManagerIAM Identity CenterCloudTrail, AWS ConfigGuardDuty, Security Hub, AWS WAF AI/ML ExposurePractical experience with:SageMaker (model deployment and endpoints)Comprehend (NLP and PII detection) DevOps & IdentityExperience with:GitOps tools (ArgoCD or Flux)CI/CD pipelines for container workloadsGitHub Actions ? AWSEKS OIDC provider integration Observability & DebuggingFamiliarity with:OpenTelemetryAWS X?RayStrong understanding of:Pod Security StandardsAdmission webhooksService account least?privilege principlesNice to haveExperience with AI agent frameworks:LangChain, Claude Agent SDK, or similarKnowledge of emerging protocols:A2A (Agent?to?Agent)Familiarity with:Amazon Bedrock Agents, Knowledge Bases, GuardrailsNamespace isolationProgramming/debugging skills:Python, Go, or Node.jsAWS Cost ExplorerLanguagesHoyGKEAlloyDBFirstIgnite makes software for university tech transfer offices.Those are the people who take researchcoming out of a university lab and get it patented, licensed, or spun out into a company.The roleWe’re hiring a Senior AI Agent Engineer.You’ll build the agents in our product, and you’ll build theevals that tell us whether each change made them better or worse.The work is document-heavy rather than chat.The agents run multi-step, call tools, read long andinconsistently formatted source material, check it against existing records, and produce output that aperson reviews before anything happens with it.Accuracy matters more here than speed or novelty.Most of the engineering effort goes into precision,traceability, and getting the agent to hand off to a human at the right moment.You’ll report to the Head of Engineering and work with product and the full-stack team.If you’veshipped agents before, you’ve probably had the experience of changing a prompt and having no ideawhether you improved anything.That problem is most of this job.What you’ll doDesign and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling,included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK,LangGraph, MCP, and Temporal Cloud.Wrap our APIs and our partners’ APIs as tools an agent can call over MCP.Some of those systemsare old, single-tenant, and outside our control, so a fair amount of the work is translation.Get structured data out of long documents and match it against records that already exist.Expectentity resolution and fuzzy matching, and expect much of it to run in batch.Stand up eval suites using various evaluation frameworks and tooling, included but not limited toPromptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses.Measure tool-use correctness, trajectory quality, and whether the agent finished the task.Every agent here produces a draft that a person signs off on.Build the citations and confidencesignals that make that review fast, and give the agent a clear way to elevate.Sit with product and domain experts and turn vague quality goals into something measurable.Sometimes the only dataset available for that is tiny, or confidential, or both.Instrument production traffic, turn real customer interactions into golden datasets, and run them asregression tests.Compare models against each other (OpenAI, Anthropic, open-weight), along with promptstrategies and agent designs, and know what each option costs in latency and quality.Bootstrap quality signal for features that have no production traffic yet.That usually meansgenerating synthetic documents and test cases, including the ugly edge cases real customers willeventually send us, and knowing where synthetic data stops being a good proxy.Write the templates, docs, and tooling the rest of the team needs to run evals without coming toyou.Required Qualifications:3+ years of engineering experience, including hands-on work on LLM or agent systems that realusers touched.You’ve evaluated agents, not only models, and you know why single-turn accuracy says little abouta multi-step run.You’ve integrated against APIs you don’t own, including old ones with bad documentation, andturned them into something an agent can call reliably.You’re comfortable with document pipelines: pulling data out, normalizing it, and checking it againsta structured source of truth.You’ve used at least one LLM evaluation framework, in-house tooling included.You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) andwhat to do about it.You can tell a real regression from noise, and design an experiment that answers the questionbeing asked instead of a nearby one.You can read a customer call transcript, work out which failures matter, and ship a fix and an evalfor them.You write clearly.Engineers won’t act on eval results they don’t read or don’t trust.You’re based somewhere between New York time (ET) and Western European time.Italy is thefurthest east we can go.You might currently be titledTitles are all over the place in this space.If the work above matches what you already do, apply.We’llgo on what you’ve shipped.Preferred Qualifications:You’ve evaluated retrieval systems: RAG, hybrid search, reranking.You’ve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAIAgents SDK, and you know how long-running tool use goes wrong.You have a background in information retrieval or search relevance.You’ve worked somewhere an agent’s output carried financial or compliance consequences.You’ve built internal tooling that non-engineers used on their own to label and review model output.This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.The roleWe’re hiring a Senior AI Agent Engineer.You’ll build the agents in our product, and you’ll build theevals that tell us whether each change made them better or worse.The work is document-heavy rather than chat.The agents run multi-step, call tools, read long andinconsistently formatted source material, check it against existing records, and produce output that aperson reviews before anything happens with it.Accuracy matters more here than speed or novelty.Most of the engineering effort goes into precision,traceability, and getting the agent to hand off to a human at the right moment.You’ll report to the Head of Engineering and work with product and the full-stack team.If you’veshipped agents before, you’ve probably had the experience of changing a prompt and having no ideawhether you improved anything.That problem is most of this job.What you’ll doDesign and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling,included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK,LangGraph, MCP, and Temporal Cloud.Wrap our APIs and our partners’ APIs as tools an agent can call over MCP.Some of those systemsare old, single-tenant, and outside our control, so a fair amount of the work is translation.Get structured data out of long documents and match it against records that already exist.Expectentity resolution and fuzzy matching, and expect much of it to run in batch.Stand up eval suites using various evaluation frameworks and tooling, included but not limited toPromptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses.Measure tool-use correctness, trajectory quality, and whether the agent finished the task.Every agent here produces a draft that a person signs off on.Build the citations and confidencesignals that make that review fast, and give the agent a clear way to elevate.Sit with product and domain experts and turn vague quality goals into something measurable.Sometimes the only dataset available for that is tiny, or confidential, or both.Instrument production traffic, turn real customer interactions into golden datasets, and run them asregression tests.Compare models against each other (OpenAI, Anthropic, open-weight), along with promptstrategies and agent designs, and know what each option costs in latency and quality.Bootstrap quality signal for features that have no production traffic yet.That usually meansgenerating synthetic documents and test cases, including the ugly edge cases real customers willeventually send us, and knowing where synthetic data stops being a good proxy.Write the templates, docs, and tooling the rest of the team needs to run evals without coming toyou.Required Qualifications:3+ years of engineering experience, including hands-on work on LLM or agent systems that realusers touched.You’ve evaluated agents, not only models, and you know why single-turn accuracy says little abouta multi-step run.You’ve integrated against APIs you don’t own, including old ones with bad documentation, andturned them into something an agent can call reliably.You’re comfortable with document pipelines: pulling data out, normalizing it, and checking it againsta structured source of truth.You’ve used at least one LLM evaluation framework, in-house tooling included.You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) andwhat to do about it.You can tell a real regression from noise, and design an experiment that answers the questionbeing asked instead of a nearby one.You can read a customer call transcript, work out which failures matter, and ship a fix and an evalfor them.You write clearly.Engineers won’t act on eval results they don’t read or don’t trust.You’re based somewhere between New York time (ET) and Western European time.Italy is thefurthest east we can go.You might currently be titledTitles are all over the place in this space.If the work above matches what you already do, apply.We’llgo on what you’ve shipped.Preferred Qualifications:You’ve evaluated retrieval systems: RAG, hybrid search, reranking.You’ve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAIAgents SDK, and you know how long-running tool use goes wrong.You have a background in information retrieval or search relevance.You’ve worked somewhere an agent’s output carried financial or compliance consequences.You’ve built internal tooling that non-engineers used on their own to label and review model output.This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.#J-*****-Ljbffr
Requirements
stack:Amazon Managed PrometheusAmazon Managed GrafanaCloudWatch Container InsightsAWS X?Ray (distributed tracing)AI/ML IntegrationLeverage AWS AI/ML services to support agent orchestration:Cost Optimization (FinOps)Spot InstancesSavings PlansKarpenter bin?packing strategiesScheduled scale?to?zero for non?production environmentsPlatform & Engineering CollaborationPartner with platform and ML teams to:Integrate MCP servers and execution frameworksSupport extensibility of the agent ecosystemSkillsMust have4+ years of hands?on AWS experienceAWS Certifications:Required: AWS Solutions Architect (Associate or Professional)Preferred: DevOps Engineer, Security Specialty Kubernetes & EKS ExpertiseStrong hands?on experience with:EKS cluster provisioning and operationsManaged node groups and KarpenterKubernetes RBAC and network policies Infrastructure as Code (Terraform)Advanced Terraform capabilities:Remote state management (S3 + DynamoDB)Security scanning (Checkov, tfsec) AWS Services ProficiencyDeep knowledge of:EKS, ECR, ALB, Route 53, ACMIAM, KMS, Secrets ManagerIAM Identity CenterCloudTrail, AWS ConfigGuardDuty, Security Hub, AWS WAF AI/ML ExposurePractical experience with:SageMaker (model deployment and endpoints)Comprehend (NLP and PII detection) DevOps & IdentityExperience with:GitOps tools (ArgoCD or Flux)CI/CD pipelines for container workloadsGitHub Actions ? AWSEKS OIDC provider integration Observability & DebuggingFamiliarity with:OpenTelemetryAWS X?RayStrong understanding of:Pod Security StandardsAdmission webhooksService account least?privilege principlesNice to haveExperience with AI agent frameworks:LangChain, Claude Agent SDK, or similarKnowledge of emerging protocols:A2A (Agent?to?Agent)Familiarity with:Amazon Bedrock Agents, Knowledge Bases, GuardrailsNamespace isolationProgramming/debugging skills:Python, Go, or Node.jsAWS Cost ExplorerLanguagesThe project is for one of the world’s famous science and technology companies in pharmaceutical industry, supporting initiatives in AWS, AI and data engineering, with plans to launch over 20 additional initiatives in the future.We are seeking a highly skilled Cloud Engineer to lead the infrastructure design, deployment, and operations of the AI agent orchestration platform on AWS.
About the company
The project is for one of the world’s famous science and technology companies in pharmaceutical industry, supporting initiatives in AWS, AI and data engineering, with plans to launch over 20 additional initiatives in the future.We are seeking a highly skilled Cloud Engineer to lead the infrastructure design, deployment, and operations of the AI agent orchestration platform on AWS.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Dev Digest 121 - AI goes offline
What is Agentic Programming and Why Should Developers Care?
I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.