AI-Enabled Platform/SRE Engineer

TECHCAFEHUB LLC
Richardson, TX, United States
21 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Build Automation Computer Programming Continuous Integration Disaster Recovery Github Monitoring of Systems Python (Programming Language) Node.Js Performance Tuning
+19 more
Reliability Engineering Runbook Software Engineering Systems Integration Datadog Google Cloud System Availability Large Language Models Grafana Apigee Kubernetes Infrastructure Automation Frameworks Rancher Graphql Restful APIs Terraform Splunk Appdynamics Microservices

Job description

We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale., * Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.

  • Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.

  • Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.

  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.

  • Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.

  • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.

  • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Requirements

  • Site Reliability Engineering (SRE) - Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.

  • Kubernetes Platform Engineering - 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.

  • Cloud & Infrastructure Automation - Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.

  • Software Development - 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).

  • Observability & Monitoring - Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.

  • API & Microservices Engineering - Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.

  • AI-Driven Operations (AIOps) - Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all