Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
We are looking for a passionate and experienced Site Reliability Engineer (SRE) to join our Cloud Platform team. The ideal candidate will have hands-on experience managing large-scale Kubernetes clusters on public cloud environments (AKS, EKS, or GKE) and a strong understanding of modern SRE and DevOps practices. You will be responsible for ensuring high availability, reliability, scalability, and performance of our cloud-native infrastructure and CI/CD systems., * Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure AKS, AWS EKS, or Google GKE).
-
Implement and maintain infrastructure as code using tools such as Terraform.
-
Collaborate with development and operations teams to improve system reliability and deployment automation.
-
Build and maintain CI/CD pipelines using Jenkins or similar tools.
-
Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.
-
Automate operational tasks using Python or other scripting languages.
-
Contribute to observability and monitoring improvements using modern tools and best practices.
-
Participate in on-call rotations and incident response processes.
Requirements
-
5-9 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
-
Strong hands-on experience managing Kubernetes clusters in production (AKS/EKS/GKE).
-
Proficiency with Terraform and cloud infrastructure automation.
-
Practical experience with Jenkins and CI/CD pipeline management.
-
Sound understanding of SRE principles (incident management, blameless postmortems, capacity planning, error budgets, etc.).
-
Good programming or scripting skills in Python (preferred) or similar languages.
-
Strong analytical, troubleshooting, and problem-solving abilities.
-
Excellent written and verbal communication skills.
-
Experience with Prometheus, Grafana, or OpenTelemetry for observability.
-
Exposure to GitOps practices and tools (e.g. Flux). Skills
- Jenkins CI/CD
- Kubernetes
- Python
- Site Reliability Engineer (Job Titles)
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Find a Developer Job: 12 Best Job Sites For Developers
Fully Remote Software Engineer Jobs
Where To Find Software Engineering Jobs
The 12 Best Jobs for Software Engineers