SRE

OpenKyber LLC
Atlanta, GA, United States
1 day ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours

Tech stack

.NET Framework Applications Architecture Microsoft Azure Bash Shell Continuous Integration DevOps Domain Name System (DNS) Python (Programming Language) Microsoft SQL Server Windows PowerShell RabbitMQ Role-Based Access Control
+17 more
Prometheus SAP (Applications) Software Deployment VMware VSphere Datadog Scripting Cloud Monitoring System Availability Grafana Kubernetes Rancher Data Management Api Gateway Splunk Dynatrace Jenkins Microservices

Job description

We are seeking a Senior Platform Reliability Engineer / Lead SRE with strong experience in Kubernetes and Rancher to support mission-critical production environments. The ideal candidate will improve platform reliability, optimize deployments, enhance observability, and collaborate with development and infrastructure teams to ensure stable, scalable application operations. Responsibilities

  • Improve Kubernetes/Rancher platform reliability and stability
  • Optimize Helm deployments and CI/CD pipelines
  • Troubleshoot pod crashes, OOMKills, deployment failures, and storage issues
  • Implement monitoring, dashboards, alerts, and operational runbooks
  • Review Kubernetes manifests and recommend best practices
  • Partner with development and infrastructure teams to improve application reliability
  • Support production deployments, go-live activities, and steady-state operations

Requirements

  • 8+ years of Platform Engineering / DevOps / SRE / Production Operations experience
  • Strong hands-on experience with Kubernetes (Production)
  • Hands-on experience with Rancher managed Kubernetes
  • Expertise in Helm, RBAC, Namespaces, PVC/PV, ConfigMaps, Secrets, Ingress
  • Troubleshooting OOMKills, Pod Failures, Storage, Networking, DNS & Resource Issues
  • CI/CD experience with Azure DevOps and/or Jenkins
  • Observability tools: Grafana, Prometheus, Splunk, ELK, Datadog, Dynatrace, Azure Monitor
  • Scripting using Bash, Python, or PowerShell
  • Experience reviewing application architecture for reliability and scalability
  • Excellent troubleshooting, documentation, and communication skills

Preferred Skills

  • .NET Microservices
  • SQL Server
  • RabbitMQ on Kubernetes
  • vSphere
  • CSI / Enterprise Storage
  • API Gateway / Ingress Controllers
  • SAP Integration
  • Manufacturing / MES / GMP environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Loading talks and stories from around this role…