Senior AI-Enabled Platform/SRE Engineer
zuven Technologies
Dallas, TX, United States
19 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$101,700.0 - $173,200.0
Working hours
Regular working hours
Job source
Tech stack
Java (Programming Language)
Application Programming Interfaces (APIs)
Artificial Intelligence
Build Automation
Computer Programming
Continuous Integration
Disaster Recovery
Github
Monitoring of Systems
Python (Programming Language)
Node.Js
Performance Tuning
+18 more
Reliability Engineering
Runbook
Software Engineering
Systems Integration
Datadog
System Availability
Large Language Models
Grafana
Apigee
Kubernetes
Infrastructure Automation Frameworks
Rancher
Graphql
Restful APIs
Terraform
Splunk
Appdynamics
Microservices
Job description
We are seeking a Senior Kubernetes-focused SRE with strong cloud automation and software engineering skills who can leverage AI/LLMs to automate operations and improve platform reliability at scale., * Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
- Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
- Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
- Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
- Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
- Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
- Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Requirements
- Site Reliability Engineering (SRE) Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
- Kubernetes Platform Engineering 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
- Cloud & Infrastructure Automation Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
- Software Development 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
- Observability & Monitoring Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
- API & Microservices Engineering Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
- AI-Driven Operations (AIOps) Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
Benefits & conditions
- $101,700-173,200 per year
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DC
Daniel Cranney
11 months ago
BR
Benjamin Ruschin
Navigating the AI Shift
about 1 year ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 1 month ago