Site Reliability Engineer (SRE)

WEEKDAY, INC.
New York, NY, United States
1 day ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$145,000.0 - $225,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Amazon Web Services Amazon Cloudfront Amazon Elastic Compute Cloud Amazon S3 Microsoft Azure Backup Devices Bash Shell Software as a Service Cloud Computing Cloud Computing Security Cloud Engineering
+38 more
Computer Programming DevOps Disaster Recovery Distributed Systems Domain Name System (DNS) Fault Tolerance Github Python (Programming Language) Enterprise Messaging Systems RabbitMQ Reliability Engineering Prometheus Software Engineering TCP/IP Datadog Data Logging Scripting Load Balancing Istio System Availability Delivery Pipeline Grafana Reliability of Systems Cloudformation Amazon Relational Database Service Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Apache Kafka Linkerd (Service Mesh) Cloudwatch Terraform New Relic (SaaS) Docker Jenkins Golang

Job description

We are looking for a highly skilled Site Reliability Engineer (SRE) to join our engineering team in New York City. You will be responsible for ensuring the reliability, scalability, performance, and security of our production systems and cloud infrastructure. You will work at the intersection of software engineering, cloud infrastructure, DevOps, and operations. The ideal candidate is passionate about automation, observability, distributed systems, and building highly reliable platforms that can scale with business growth., * Design, build, and maintain highly available and scalable production infrastructure.

  • Improve system reliability, performance, scalability, and operational efficiency.

  • Develop automation to eliminate repetitive operational tasks and reduce manual intervention.

  • Manage and optimize cloud infrastructure across AWS, GCP, or Azure.

  • Build and maintain CI/CD pipelines for reliable and automated deployments.

  • Manage containerized workloads using Docker and Kubernetes.

  • Implement infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code tools.

  • Establish monitoring, logging, alerting, and observability across production systems.

  • Work with tools such as Prometheus, Grafana, Datadog, CloudWatch, ELK, or OpenTelemetry.

  • Troubleshoot production incidents and participate in incident response and root-cause analysis.

  • Define and monitor SLIs, SLOs, and SLAs.

  • Identify system bottlenecks and implement performance and reliability improvements.

  • Develop disaster recovery, backup, and business continuity strategies.

  • Collaborate with software engineers to build reliable and operationally efficient applications.

  • Participate in on-call rotations and help improve incident-management processes.

  • Document infrastructure, operational procedures, architecture, and incident learnings.

Requirements

  • 2-8 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.

  • Strong experience with Linux/Unix systems administration.

  • Hands-on experience with AWS, GCP, or Azure.

  • Strong knowledge of Docker and Kubernetes.

  • Experience with Terraform or other Infrastructure-as-Code tools.

  • Proficiency in at least one programming or scripting language such as Python, Go, Bash, or Java.

  • Experience building and managing CI/CD pipelines.

  • Strong understanding of networking, DNS, TCP/IP, load balancing, and security fundamentals.

  • Experience with monitoring, logging, alerting, and observability platforms.

  • Understanding of distributed systems, scalability, high availability, and fault tolerance.

  • Strong troubleshooting and problem-solving skills.

  • Experience with incident response and root-cause analysis.

Nice-to-Have Skills

  • Experience with AWS EKS, ECS, EC2, Lambda, RDS, S3, CloudFront, and CloudWatch.

  • Experience with Prometheus, Grafana, Datadog, New Relic, or OpenTelemetry.

  • Knowledge of service meshes such as Istio or Linkerd.

  • Experience with Kafka, RabbitMQ, or other distributed messaging systems.

  • Experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.

  • Knowledge of security engineering and cloud security best practices.

  • Experience implementing SLOs, error budgets, and reliability metrics.

  • Familiarity with FinOps and cloud-cost optimization.

  • Experience operating large-scale distributed systems.

  • AWS/GCP/Azure or Kubernetes certifications.

  • Experience in fintech, SaaS, high-growth startups, or other high-availability environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Loading talks and stories from around this role…