SRE
OpenKyber LLC
Atlanta, GA, United States
1 day ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source
Tech stack
.NET Framework
Applications Architecture
Microsoft Azure
Bash Shell
Continuous Integration
DevOps
Domain Name System (DNS)
Python (Programming Language)
Microsoft SQL Server
Windows PowerShell
RabbitMQ
Role-Based Access Control
+17 more
Prometheus
SAP (Applications)
Software Deployment
VMware VSphere
Datadog
Scripting
Cloud Monitoring
System Availability
Grafana
Kubernetes
Rancher
Data Management
Api Gateway
Splunk
Dynatrace
Jenkins
Microservices
Job description
We are seeking a Senior Platform Reliability Engineer / Lead SRE with strong experience in Kubernetes and Rancher to support mission-critical production environments. The ideal candidate will improve platform reliability, optimize deployments, enhance observability, and collaborate with development and infrastructure teams to ensure stable, scalable application operations. Responsibilities
- Improve Kubernetes/Rancher platform reliability and stability
- Optimize Helm deployments and CI/CD pipelines
- Troubleshoot pod crashes, OOMKills, deployment failures, and storage issues
- Implement monitoring, dashboards, alerts, and operational runbooks
- Review Kubernetes manifests and recommend best practices
- Partner with development and infrastructure teams to improve application reliability
- Support production deployments, go-live activities, and steady-state operations
Requirements
- 8+ years of Platform Engineering / DevOps / SRE / Production Operations experience
- Strong hands-on experience with Kubernetes (Production)
- Hands-on experience with Rancher managed Kubernetes
- Expertise in Helm, RBAC, Namespaces, PVC/PV, ConfigMaps, Secrets, Ingress
- Troubleshooting OOMKills, Pod Failures, Storage, Networking, DNS & Resource Issues
- CI/CD experience with Azure DevOps and/or Jenkins
- Observability tools: Grafana, Prometheus, Splunk, ELK, Datadog, Dynatrace, Azure Monitor
- Scripting using Bash, Python, or PowerShell
- Experience reviewing application architecture for reliability and scalability
- Excellent troubleshooting, documentation, and communication skills
Preferred Skills
- .NET Microservices
- SQL Server
- RabbitMQ on Kubernetes
- vSphere
- CSI / Enterprise Storage
- API Gateway / Ingress Controllers
- SAP Integration
- Manufacturing / MES / GMP environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleβ¦