> Markdown version of [/jobs/ext/1254486-site-reliability-engineer-gcp-sre-ai](https://www.wearedevelopers.com/jobs/ext/1254486-site-reliability-engineer-gcp-sre-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (GCP SRE) - AI - **Company:** Purple Drive Technologies LLC - **Location:** Alpharetta, GA, United States - **Experience:** Expert - **Salary:** $135,200.0 - $145,600.0 - **Contract:** Temporary to permanent - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, BigQuery, Cloud Computing, Computer Programming, Continuous Integration, DevOps, Disaster Recovery, Data Flow Control, Github, Monitoring of Systems, Python (Programming Language), Machine Learning, Reliability Engineering, Prometheus, Azure Machine Learning, YAML, Data Logging, Scripting, Google Cloud, Cloud Monitoring, Spring Cloud, Istio, System Availability, Large Language Models, Grafana, Infrastructure as Code (IaC), Git, Kubernetes, Information Technology, Performance Monitor, Machine Learning Operations, Virtual Agents, Functional Programming, Terraform, Devsecops, Docker, Jenkins - **Published:** July 13, 2026 - **Apply:** https://www.careerjet.com/jobad/us83622f2ab07d083188e831f3749ece44 ## About the Role * Bachelor's degree in Computer Science, Information Technology, Engineering, or related field. * 6-8 years of experience in Site Reliability Engineering or Cloud Operations. * Strong hands-on experience with Google Cloud Platform. * Excellent troubleshooting, communication, and production support skills. ## Description We are seeking an experienced Google Cloud Site Reliability Engineer (GCP SRE) with expertise in Google Cloud Platform (GCP), Site Reliability Engineering, Kubernetes, and AI/ML platforms. The ideal candidate will be responsible for ensuring the reliability, scalability, availability, and operational excellence of cloud-native applications while supporting Vertex AI, Agentic AI solutions, and large-scale production environments. The candidate should have strong experience in incident management, production support, cloud monitoring, and Kubernetes-based deployments, with the ability to reduce operational noise through proactive observability and automation. Key Responsibilities * Design, deploy, and support highly available applications on Google Cloud Platform (GCP). * Maintain production reliability and uptime following SRE principles. * Monitor cloud infrastructure and applications using observability tools. * Lead P1/P2 incident management, root cause analysis (RCA), and post-incident reviews. * Implement automation to reduce operational overhead and monitoring noise. * Deploy and manage Kubernetes workloads on Google Kubernetes Engine (GKE). * Build and maintain CI/CD pipelines for cloud-native applications. * Support Vertex AI model training, deployment, and inference pipelines. * Collaborate with AI/ML engineers to operationalize machine learning solutions. * Work with Agentic AI assistants and AI-driven automation workflows. * Optimize cloud infrastructure performance, scalability, and cost. * Implement security, governance, and best practices across GCP environments. * Participate in on-call support and production incident rotations. Required Technical Skills Google Cloud Platform * Google Cloud Platform (GCP) * Google Kubernetes Engine (GKE) * Pub/Sub * BigQuery * Cloud Spanner * Dataflow * Firestore * Vertex AI Site Reliability Engineering * Site Reliability Engineering (SRE) * Production Support * Incident Management * Root Cause Analysis (RCA) * High Availability * Disaster Recovery * Performance Monitoring * Reliability Engineering Containers & DevOps * Kubernetes * Docker * Helm * CI/CD * Git * Jenkins / GitHub Actions (Preferred) AI / ML * Vertex AI * Machine Learning Pipelines * Model Deployment * Agentic AI * AI Assistants * LLM Operations (Preferred) Monitoring & Observability * Prometheus * Grafana * Cloud Monitoring * Cloud Logging * Alerting * Monitoring Optimization Programming * Python * Bash/Shell Scripting * YAML Preferred Skills * AWS (EC2, S3, Lambda) * Terraform * Infrastructure as Code (IaC) * Anthos * Service Mesh (Istio) * GitOps * ArgoCD * AI/ML Operations (MLOps) * DevSecOps, * Google Cloud Platform (GCP) * Kubernetes (GKE) * Docker * Helm * Pub/Sub * BigQuery * Cloud Spanner * Dataflow * Firestore * Vertex AI * Incident Management * Site Reliability Engineering (SRE) * Monitoring & Observability * Root Cause Analysis (RCA) * Agentic AI * AI Assistants ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [SEO in an AI world - Google vs. ChatGPT and survival tips for content creators](https://www.wearedevelopers.com/magazine/534-seo-in-an-ai-world-google-vs-chatgpt-and-survival-tips-for-content-creators)