Lead Platform Reliability Engineer, Global AI Platform & Solutions
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
Experteer Overview In this role you will lead reliability for a shared platform that powers internal AI solution development. You’ll partner with global teams to define and meet SLOs, scale platforms, and improve observability and incident response. You’ll build self-service and automation to reduce toil, while ensuring security, governance, and data residency compliance. This is a hands-on, cross-functional role shaping a platform-as-a-product that enables high-traffic AI workflows at scale. Compensation / Benefits * Define SLOs/SLIs, manage operations budgets, reduce MTTR, capacity planning, and autoscale tuning * Build and maintain logging, metrics, tracing, alerting; instrument platform components; create runbooks and dashboards * On-call incident response: triage, mitigate, root-cause, postmortems and corrective actions * Develop self-service capabilities and automation: AIOps/MLOps/GitOps/CI/CD pipelines and operational automations * Manage infrastructure via Terraform/Ansible; prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor’s in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off
Requirements
_ prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor’s in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
Stephan Gillich - Bringing AI Everywhere
How to Become an AI Engineer
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence