Site Reliability Engineer - AWS

Centillion Infotech
Buffalo, NY, United States
4 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$114,400.0 - $124,800.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon S3 Backup Devices Bash Shell Cloud Computing Cloud Engineering Computer Programming Continuous Delivery Continuous Integration Linux DevOps Disaster Recovery
+29 more
Distributed Systems Amazon DynamoDB Failover Github Identity and Access Management Python (Programming Language) Reliability Engineering Prometheus Runbook Software Vulnerability Management Datadog Data Logging Scripting Performance Testing Grafana Gitlab Cloudformation Amazon Relational Database Service Kubernetes Infrastructure Automation Frameworks Deployment Automation AWS Fargate Route53 Cloudwatch Api Gateway Terraform Dynatrace Devsecops Jenkins

Job description

We are seeking an experienced Site Reliability Engineer with strong Amazon Web Services expertise to build, automate, monitor, and support highly available cloud platforms. The ideal candidate should have hands-on experience with Amazon Web Services, Kubernetes, ECS/Fargate, Infrastructure as Code, observability, incident management, and production reliability. Key Responsibilities Design, build, and operate scalable and highly available cloud infrastructure on Amazon Web Services. Deploy and manage containerized applications using Kubernetes, ECS, and Fargate. Automate infrastructure provisioning using Terraform and CloudFormation. Define and monitor Service-Level Indicators, Service-Level Objectives, and Service-Level Agreements. Implement monitoring, logging, tracing, alerting, and observability solutions. Develop dashboards using Prometheus, Grafana, ELK, CloudWatch, or Datadog. Participate in incident response, troubleshooting, and production-support activities. Conduct root-cause analysis and document corrective and preventive actions. Automate repetitive operational activities using Python, Bash, or Go. Build and maintain Continuous Integration and Continuous Deployment pipelines. Improve application availability, performance, scalability, security, and cost efficiency. Perform capacity planning, performance testing, and reliability assessments. Support disaster recovery, backup, business continuity, and failover processes. Collaborate with developers, cloud architects, security teams, and business stakeholders. Prepare operational documentation, architecture diagrams, and production runbooks. Required Skills, Job Title: Lead Site Reliability Engineer Location: Buffalo, NY Duration:6+Months(with Possible Extension) Role Responsibilities: Job Description: Overview Responsible …

  • 1 day ago
  • Apply easily, Lead Site Reliability Engineer Location: Buffalo, New York Pay Rate: $ 55 TO $ 60/hr Contract: Long-Term Schedule: Onsite 100% Onsite Experience: 11+ Years Required A…
  • 1 day ago
  • Apply easily

Requirements

6+ years of experience in Site Reliability Engineering, DevOps, cloud engineering, or production operations. Strong hands-on experience with Amazon Web Services. Experience with Kubernetes, Amazon ECS, and Fargate. Strong Infrastructure-as-Code experience using Terraform and CloudFormation. Experience with Prometheus, Grafana, ELK, CloudWatch, Datadog, or similar observability tools. Strong knowledge of Linux, networking, security, and distributed systems. Scripting or programming experience with Python, Bash, or Go. Experience with Jenkins, GitHub Actions, GitLab Continuous Integration, or similar deployment tools. Strong understanding of incident management, root-cause analysis, and problem management. Knowledge of Service-Level Indicators, Service-Level Objectives, error budgets, and availability metrics. Excellent analytical, troubleshooting, communication, and collaboration skills. Preferred Skills Experience with Amazon EKS, Lambda, Route 53, RDS, DynamoDB, S3, IAM, and API Gateway. Knowledge of OpenTelemetry and distributed tracing. Experience with infrastructure security, vulnerability remediation, and DevSecOps. Experience supporting enterprise production environments. Relevant Amazon Web Services, Kubernetes, Terraform, or Site Reliability Engineering certifications.

Benefits & conditions

  • $139,700-232,900 per year Overview: Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a sub…

  • 1 day ago +

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:23 min

Reviewing AWS infrastructure deployment configuration and planning

Devlin Duldulao · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:38 min

Extreme engineering culture for massive web data operations

Ariel Shulman Ariel Shulman +1 · WWC Europe 2026

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all