> Markdown version of [/jobs/ext/1975609-ai-enabled-platform-sre-engineer](https://www.wearedevelopers.com/jobs/ext/1975609-ai-enabled-platform-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI-Enabled Platform/SRE Engineer - **Company:** TECHCAFEHUB LLC - **Location:** Richardson, TX, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Build Automation, Computer Programming, Continuous Integration, Disaster Recovery, Github, Monitoring of Systems, Python (Programming Language), Node.Js, Performance Tuning, Reliability Engineering, Runbook, Software Engineering, Systems Integration, Datadog, Google Cloud, System Availability, Large Language Models, Grafana, Apigee, Kubernetes, Infrastructure Automation Frameworks, Rancher, Graphql, Restful APIs, Terraform, Splunk, Appdynamics, Microservices - **Published:** August 7, 2026 - **Apply:** https://www.dice.com/job-detail/3f2edf81-44fe-410b-98ec-129bf5dd63e8 ## About the Role * Site Reliability Engineering (SRE) - Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence. * Kubernetes Platform Engineering - 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization. * Cloud & Infrastructure Automation - Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation. * Software Development - 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows). * Observability & Monitoring - Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring. * API & Microservices Engineering - Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies. * AI-Driven Operations (AIOps) - Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows. ## Description We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale., * Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations. * Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution. * Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies. * Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting. * Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments. * Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives. * Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Navigating the AI Wave in DevOps](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)