> Markdown version of [/jobs/ext/1316307-cloud-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1316307-cloud-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Site Reliability Engineer - **Company:** The Eeo - **Location:** San Jose, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Computer Programming, Databases, Continuous Integration, Data Centers, DevOps, Distributed Systems, Github, Python (Programming Language), Memcached, NoSQL, Octopus Deploy, Redis, Reliability Engineering, Ansible, Prometheus, Software Engineering, SQL Databases, AI Infrastructure, Datadog, Scripting, Graphics Processing Unit (GPU), Google Cloud, Grafana, Caching, Infrastructure as Code (IaC), Cloudformation, Kubernetes, Information Technology, Machine Learning Operations, Terraform, Docker, Elk Stack, Jenkins - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/ba3ba4b8-8ba4-4039-8500-330a2204a0f9 ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. * 3-5+ years of experience in a Site Reliability Engineer, DevOps, or related role supporting a large-scale, customer-facing service in a public cloud environment (AWS, Google Cloud Platform, Azure). * Strong programming/scripting skills in languages like Python, Go, or Java. * Proven experience with containerization and orchestration technologies (Docker, Kubernetes). * Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog). * Solid experience with Infrastructure as Code (e.g., Terraform, CloudFormation). * Familiarity with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD). * Excellent problem-solving skills and a systematic approach to troubleshooting complex distributed systems. What Will Make You Stand Out (Nice-to-Haves) * Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure. * Direct experience supporting ML/AI inferencing services in production. * Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs. * Knowledge of model serving frameworks like vLLM, SGLang or Ray. * Understanding of MLOps principles and practices. * Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached). * Strong Linux/Unix system administration fundamentals. ## Description As a Cloud Site Reliability Engineer (SRE) specializing in our AI Inferencing Service, you will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability. What You'll Do Service Ownership & On-Call: Take shared ownership of the production inferencing service, including its availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning across multiple regions. This includes implementing and supporting AI infrastructure in new regions, such as Asia, Europe, and Latin America, to support the growth of our business. Participate in a balanced on-call rotation to provide 24/7 support for the service. On-Call & Work-Life Balance We believe a sustainable on-call schedule is critical for long-term success and team health. Our on-call philosophy is built on the following principles: * Balanced Rotation: The on-call rotation is shared equally across the team, typically following a primary/secondary (follow-the-sun) model to ensure no single person bears a disproportionate burden. * Focus on Prevention: We invest heavily in automation, robust testing, and system design to prevent pages before they happen. The goal of on-call is not to heroically fight fires, but to manage rare, complex failures and use those learnings to make the system more resilient. * Actionable Alerts: We have a strict policy against alert fatigue. Alerts must be actionable and require immediate human intervention. * Incident Management: Lead the response to incidents affecting the inferencing service, driving blameless post-mortems and implementing corrective actions to prevent recurrence. * Monitoring & Alerting: Develop and maintain advanced monitoring, alerting, and dashboarding (using tools like Prometheus, Grafana, Datadog) to gain deep insights into service health, model performance (e.g., latency, throughput, error rates), and accelerator utilization. A key responsibility is ensuring alerts are actionable and have a low false-positive rate, minimizing on-call fatigue. * Performance & Scalability: Proactively identify and eliminate performance bottlenecks. Design and implement auto-scaling policies to handle variable inference loads cost-effectively. Use insights from on-call incidents to drive improvements that enhance system stability and scalability. * Infrastructure as Code (IaC): Manage and evolve our cloud infrastructure (on AWS, Google Cloud Platform, and/or Azure along with on-prem) using tools like Terraform and Ansible, ensuring it is secure, repeatable, and scalable. * CI/CD & Automation: Champion automation by building and improving CI/CD pipelines for the seamless and safe deployment of new model versions and service updates. A core goal is to automate manual toil identified during on-call shifts, reducing future operational overhead. * Capacity Planning: Forecast infrastructure needs based on product roadmaps and usage trends. Work with finance and engineering teams to manage cloud costs and optimize spending. * SLOs & SLIs: Define, measure, and report on Service Level Objectives (SLOs) and Indicators (SLIs) for the inferencing platform, using data to drive prioritization and reliability investments. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe)