Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+38 more
Job description
We are looking for a highly skilled Site Reliability Engineer (SRE) to join our engineering team in New York City. You will be responsible for ensuring the reliability, scalability, performance, and security of our production systems and cloud infrastructure. You will work at the intersection of software engineering, cloud infrastructure, DevOps, and operations. The ideal candidate is passionate about automation, observability, distributed systems, and building highly reliable platforms that can scale with business growth., * Design, build, and maintain highly available and scalable production infrastructure.
-
Improve system reliability, performance, scalability, and operational efficiency.
-
Develop automation to eliminate repetitive operational tasks and reduce manual intervention.
-
Manage and optimize cloud infrastructure across AWS, GCP, or Azure.
-
Build and maintain CI/CD pipelines for reliable and automated deployments.
-
Manage containerized workloads using Docker and Kubernetes.
-
Implement infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code tools.
-
Establish monitoring, logging, alerting, and observability across production systems.
-
Work with tools such as Prometheus, Grafana, Datadog, CloudWatch, ELK, or OpenTelemetry.
-
Troubleshoot production incidents and participate in incident response and root-cause analysis.
-
Define and monitor SLIs, SLOs, and SLAs.
-
Identify system bottlenecks and implement performance and reliability improvements.
-
Develop disaster recovery, backup, and business continuity strategies.
-
Collaborate with software engineers to build reliable and operationally efficient applications.
-
Participate in on-call rotations and help improve incident-management processes.
-
Document infrastructure, operational procedures, architecture, and incident learnings.
Requirements
-
2-8 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.
-
Strong experience with Linux/Unix systems administration.
-
Hands-on experience with AWS, GCP, or Azure.
-
Strong knowledge of Docker and Kubernetes.
-
Experience with Terraform or other Infrastructure-as-Code tools.
-
Proficiency in at least one programming or scripting language such as Python, Go, Bash, or Java.
-
Experience building and managing CI/CD pipelines.
-
Strong understanding of networking, DNS, TCP/IP, load balancing, and security fundamentals.
-
Experience with monitoring, logging, alerting, and observability platforms.
-
Understanding of distributed systems, scalability, high availability, and fault tolerance.
-
Strong troubleshooting and problem-solving skills.
-
Experience with incident response and root-cause analysis.
Nice-to-Have Skills
-
Experience with AWS EKS, ECS, EC2, Lambda, RDS, S3, CloudFront, and CloudWatch.
-
Experience with Prometheus, Grafana, Datadog, New Relic, or OpenTelemetry.
-
Knowledge of service meshes such as Istio or Linkerd.
-
Experience with Kafka, RabbitMQ, or other distributed messaging systems.
-
Experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
-
Knowledge of security engineering and cloud security best practices.
-
Experience implementing SLOs, error budgets, and reliability metrics.
-
Familiarity with FinOps and cloud-cost optimization.
-
Experience operating large-scale distributed systems.
-
AWS/GCP/Azure or Kubernetes certifications.
-
Experience in fintech, SaaS, high-growth startups, or other high-availability environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleβ¦