Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+23 more
Job description
The SRE will be responsible for designing, deploying, automating, and supporting highly available, scalable, and secure containerized applications in cloud-native environments. You will work closely with development, operations, and security teams to ensure the reliability, performance, and efficiency of our production systems., * Design, deploy, and manage Kubernetes clusters on-premises and/or cloud-managed such as EKS, AKS, GKE to support scalable microservices architectures.
- Automate infrastructure provisioning and application deployment using Infrastructure as Code (IaC) tools such as Terraform, Helm, or CloudFormation.
- Monitor, troubleshoot, and optimize system performance using observability tools.
- Implement and manage CI/CD pipelines to ensure rapid, repeatable, and reliable software delivery.
- Ensure system reliability, availability, and security through proactive monitoring, incident response, and root cause analysis.
- Develop and maintain runbooks, dashboards, and documentation for operational procedures and system architectures.
- Participate in on-call rotations and respond to production incidents, ensuring minimal downtime and fast recovery.
- Collaborate with development and operations teams to drive DevOps and SRE best practices including capacity planning, scaling, and cost optimization.
- Continuously improve automation tooling and processes to reduce manual work and increase system reliability.
Requirements
We are seeking a highly skilled Site Reliability Engineer (SRE) with deep expertise in Kubernetes and cloud technologies AWS, Azure, or GCP., * 3 years experience as an SRE, DevOps Engineer, or similar role supporting large-scale systems.
- Expertise in Kubernetes deployment, scaling, upgrades, troubleshooting, and networking.
- Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP).
- Proficiency in scripting/programming (Python, Bash, Go, etc.).
- Experience with IaC tools (Terraform, Helm, CloudFormation, ARM, etc.).
- Strong knowledge of Linux systems administration and networking concepts.
- Familiarity with monitoring, logging, and ingestion tools (Prometheus, Grafana, ELK, EFK).
- Experience with CI/CD tools (Jenkins, GitLab CI, ArgoCD, etc.).
- Understanding of security best practices in cloud and containerized environments.
- Excellent troubleshooting and problem-solving skills.
- Strong communication and collaboration skills., * Certified Kubernetes Administrator (CKA) or similar certification.
- Experience with service mesh (Istio, Linkerd), ingress controllers, and API gateways.
- Experience in a multicloud or hybrid cloud environment.
- Familiarity with GitOps practices and tools (ArgoCD, Flux).
- Experience with disaster recovery, backup, and business continuity planning., Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
This role is ideal for engineers who are passionate about automation, reliability, and modern cloud-native architectures and who thrive in fast-paced, collaborative environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Find a Developer Job: 12 Best Job Sites For Developers
Where To Find Software Engineering Jobs
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again