> Markdown version of [/jobs/ext/3659950-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/3659950-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** WEEKDAY, INC. - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $145,000.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Amazon Cloudfront, Amazon Elastic Compute Cloud, Amazon S3, Microsoft Azure, Backup Devices, Bash Shell, Software as a Service, Cloud Computing, Cloud Computing Security, Cloud Engineering, Computer Programming, DevOps, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Github, Python (Programming Language), Enterprise Messaging Systems, RabbitMQ, Reliability Engineering, Prometheus, Software Engineering, TCP/IP, Datadog, Data Logging, Scripting, Load Balancing, Istio, System Availability, Delivery Pipeline, Grafana, Reliability of Systems, Cloudformation, Amazon Relational Database Service, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Apache Kafka, Linkerd (Service Mesh), Cloudwatch, Terraform, New Relic (SaaS), Docker, Jenkins, Golang - **Published:** October 9, 2026 - **Apply:** https://www.careerjet.com/jobad/us18bc078ff2db90db06312dd5c33c8b7f ## About the Role * 2-8 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field. * Strong experience with Linux/Unix systems administration. * Hands-on experience with AWS, GCP, or Azure. * Strong knowledge of Docker and Kubernetes. * Experience with Terraform or other Infrastructure-as-Code tools. * Proficiency in at least one programming or scripting language such as Python, Go, Bash, or Java. * Experience building and managing CI/CD pipelines. * Strong understanding of networking, DNS, TCP/IP, load balancing, and security fundamentals. * Experience with monitoring, logging, alerting, and observability platforms. * Understanding of distributed systems, scalability, high availability, and fault tolerance. * Strong troubleshooting and problem-solving skills. * Experience with incident response and root-cause analysis. Nice-to-Have Skills * Experience with AWS EKS, ECS, EC2, Lambda, RDS, S3, CloudFront, and CloudWatch. * Experience with Prometheus, Grafana, Datadog, New Relic, or OpenTelemetry. * Knowledge of service meshes such as Istio or Linkerd. * Experience with Kafka, RabbitMQ, or other distributed messaging systems. * Experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD. * Knowledge of security engineering and cloud security best practices. * Experience implementing SLOs, error budgets, and reliability metrics. * Familiarity with FinOps and cloud-cost optimization. * Experience operating large-scale distributed systems. * AWS/GCP/Azure or Kubernetes certifications. * Experience in fintech, SaaS, high-growth startups, or other high-availability environments. ## Description We are looking for a highly skilled Site Reliability Engineer (SRE) to join our engineering team in New York City. You will be responsible for ensuring the reliability, scalability, performance, and security of our production systems and cloud infrastructure. You will work at the intersection of software engineering, cloud infrastructure, DevOps, and operations. The ideal candidate is passionate about automation, observability, distributed systems, and building highly reliable platforms that can scale with business growth., * Design, build, and maintain highly available and scalable production infrastructure. * Improve system reliability, performance, scalability, and operational efficiency. * Develop automation to eliminate repetitive operational tasks and reduce manual intervention. * Manage and optimize cloud infrastructure across AWS, GCP, or Azure. * Build and maintain CI/CD pipelines for reliable and automated deployments. * Manage containerized workloads using Docker and Kubernetes. * Implement infrastructure using Terraform, CloudFormation, or similar Infrastructure-as-Code tools. * Establish monitoring, logging, alerting, and observability across production systems. * Work with tools such as Prometheus, Grafana, Datadog, CloudWatch, ELK, or OpenTelemetry. * Troubleshoot production incidents and participate in incident response and root-cause analysis. * Define and monitor SLIs, SLOs, and SLAs. * Identify system bottlenecks and implement performance and reliability improvements. * Develop disaster recovery, backup, and business continuity strategies. * Collaborate with software engineers to build reliable and operationally efficient applications. * Participate in on-call rotations and help improve incident-management processes. * Document infrastructure, operational procedures, architecture, and incident learnings.