Lead Platform Reliability Engineer, Global AI Platform & Solutions
Role details
Job location
Tech stack
Job description
Experteer Overview In this role you will lead reliability for a shared platform that powers internal AI solution development. You'll partner with global teams to define and meet SLOs, scale platforms, and improve observability and incident response. You'll build self-service and automation to reduce toil, while ensuring security, governance, and data residency compliance. This is a hands-on, cross-functional role shaping a platform-as-a-product that enables high-traffic AI workflows at scale. Compensation / Benefits * Define SLOs/SLIs, manage operations budgets, reduce MTTR, capacity planning, and autoscale tuning * Build and maintain logging, metrics, tracing, alerting; instrument platform components; create runbooks and dashboards * On-call incident response: triage, mitigate, root-cause, postmortems and corrective actions * Develop self-service capabilities and automation: AIOps/MLOps/GitOps/CI/CD pipelines and operational automations * Manage infrastructure via Terraform/Ansible; prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor's in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off
Requirements
_ prevent configuration drift * Enforce identity/RBAC, secrets management, supply chain security, regulatory controls; collaborate with risk and audit * Optimize resource usage and cost: capacity planning, rightsizing, reservations/spot * Safe change management: progressive delivery, policy-as-code guardrails * Treat platform as a product: define operational SLAs aligned to roadmap, service catalog, and developer experience * Collaborate with global engineering, security, and AI governance teams for cross-geo regulations and data residency * Operate scalable backend services for high-traffic agent interactions, retrieval operations, and real-time execution * Maintain AI services runbooks and playbooks; enable GOCC Tasks * Bachelor's in Computer Science/Engineering or equivalent (not strictly required) * 5-8 years in DevOps/Platform Engineering or Production Operations * Proven track record with large-scale distributed systems and on-call experience * Cloud-native development experience: Azure, Kubernetes, containers, CI/CD, observability stacks * Proficiency in Python and/or Java/Scala/TypeScript for backend services and automation * Knowledge of AI solutions, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflows * Understanding of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning * Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance) * Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs Key requirements * health and dental insurance * mental health support * vision coverage * disability and life insurance * retirement savings plans * paid time off