> Markdown version of [/jobs/ext/1186790-site-reliability-engineer-sre-generative-ai-platform](https://www.wearedevelopers.com/jobs/ext/1186790-site-reliability-engineer-sre-generative-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer SRE - Generative AI Platform - **Company:** Htc Inc. - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Salary:** $101,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Backup Devices, Bash Shell, Cloud Computing, Information Systems, Continuous Integration, Data Infrastructure, Linux, DevOps, Distributed Systems, Github, Python (Programming Language), Key Management, PostgreSQL, Machine Learning, MongoDB, Redis, Reliability Engineering, Ansible, Prometheus, Runbook, Service Pack, Vault (Revision Control System), YAML, Zabbix, Datadog, Data Logging, Scripting, Cloud Platform System, Chatbots, System Availability, Delivery Pipeline, Large Language Models, Grafana, Kubernetes Helm Charts, Multi-Cloud, Generative AI, Backend, Gitlab, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Apache Kafka, Nintex, Terraform, Splunk, Appdynamics, Jenkins - **Published:** July 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=09afe3e820846fc5 ## About the Role Build and support the cloud infrastructure behind enterprise Generative AI platforms. We are looking for a hands-on SRE with strong Kubernetes, Terraform, automation, and observability experience who is ready to grow into advanced AI platform reliability work., * 5+ years of experience in SRE, DevOps, platform engineering, cloud infrastructure, Linux systems, or related technical roles. * Hands-on experience operating or supporting Kubernetes in a production or high-availability environment. * Experience with Terraform or similar infrastructure-as-code tools. * Scripting or automation experience using Python, Bash, Ansible, YAML, or similar. * Experience with monitoring, logging, alerting, or observability tools such as ELK, Zabbix, Prometheus, Grafana, Splunk, Datadog, AppDynamics, or similar. * Experience troubleshooting incidents, performing root cause analysis, and improving reliability after production issues. * Familiarity with CI/CD concepts and deployment automation. * Strong Linux, networking, infrastructure, or cloud troubleshooting skills. * Ability to learn new tools quickly and work effectively in a fast-paced engineering environment. * Strong communication skills and willingness to collaborate with senior engineers, developers, architects, and business stakeholders., * Experience with GCP, AWS, or Azure cloud environments. * Exposure to GCP/GKE, AWS/EKS, or Azure/AKS. * Experience with Helm charts. * Exposure to Harness, GitHub Actions, GitLab, Jenkins, or Azure DevOps. * Experience with OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar enterprise observability tools. * Exposure to backend systems such as Kafka, PostgreSQL, Redis, MongoDB, or Vault. * Experience supporting AI, ML, LLM, chatbot, conversational AI, or data platform environments. * Experience with security patching, access controls, secrets management, and production governance. * Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent relevant experience., * SRE, DevOps, Infrastructure, Platform, or Cloud Operations: 5 years (Required) * expert level Kubernetes managing deployments at scale: 3 years (Required) ## Description Site Reliability Engineer to support cloud-native infrastructure and platform reliability for enterprise Generative AI and conversational experience platforms. This role is ideal for a strong hands-on SRE / DevOps / platform engineer who has solid production experience with Kubernetes, Terraform, automation, observability, incident response, and CI/CD, and is ready to grow deeper into multi-cloud and Generative AI platform environments. You do not need to be the lead architect on day one, but you should be bright, adaptable, technically curious, and able to learn quickly in a fast-moving platform engineering environment. What You'll Do * Support Kubernetes-based platform infrastructure for AI and data service workloads. * Build and maintain infrastructure using Terraform and infrastructure-as-code practices. * Assist with Helm chart updates, Kubernetes configuration, deployment automation, and environment management. * Help implement monitoring, alerting, logging, and observability across platform services. * Troubleshoot production issues across Kubernetes clusters, infrastructure, CI/CD pipelines, and distributed systems. * Automate repetitive operational tasks using Python, Bash, Ansible, YAML, or similar tools. * Support deployment pipelines and release processes using CI/CD tools such as Harness, GitHub Actions, GitLab, Jenkins, or Azure DevOps. * Assist with rollout patterns such as canary deployments, blue/green releases, and rollback processes. * Support operational processes including patching, upgrades, backups, capacity planning, and incident response. * Work with senior SREs, architects, developers, DevOps, and security teams to improve reliability and scalability. * Contribute to documentation, runbooks, troubleshooting guides, and platform best practices. * Learn and support backend platform components such as Kafka, PostgreSQL, Redis, Vault, MongoDB, and n8n. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)