Site Reliability Engineer

Tata Consultancy Services Limited
Deerfield, IL, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$120,000.0 - $140,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Performance Management Microsoft Azure Cloud Computing Computer Networks Continuous Integration DevOps Disaster Recovery Domain Name System (DNS) Python (Programming Language) Log Analysis Windows PowerShell Role-Based Access Control
+16 more
Reliability Engineering Azure DevOps Pipelines Runbook System Programming Azure Automation Performance Testing Microsoft Power Automate Cloud Monitoring Grafana Firewalls (Computer Science) Infrastructure Automation Frameworks Bicep CIS Benchmarks Terraform Pagerduty Servicenow

Job description

  • Define, own, and enforce enterprise-wide SLOs, SLIs, and Error Budgets across all Tier-0 and Tier-1 Azure-hosted services; report SLA compliance to executive stakeholders monthly.
  • Lead architectural reviews for new services and ensure reliability non-functionals (availability targets, RTO/RPO) are embedded from design through to production.
  • Champion and implement chaos engineering practices using Azure Chaos Studio and custom fault injection frameworks to proactively surface reliability risks.
  • Drive Disaster Recovery (DR) design and conduct quarterly DR drills across Azure paired regions. Incident Management & On-Call
  • Serve as Incident Commander for P1/P2 major incidents, own end-to-end incident lifecycle from detection through resolution and Post-Incident Review (PIR).
  • Participate in a structured On-Call rotation with follow-the-sun global coverage; maintain response SLAs of <5 minutes for Tier-0 services.
  • Drive blameless post-mortem culture and ensure all action items from PIRs are tracked and delivered within agreed SLA.

Observability & Platform Engineering

  • Design and operate the enterprise observability stack: Azure Monitor, Log Analytics Workspaces, Application Insights, and Azure Managed Grafana; ensure full MELT (Metrics, Events, Logs, Traces) coverage.
  • Build and maintain alerting frameworks usi ng Azure Monitor Alert Rules and Azure Action Groups integrated with PagerDuty and ServiceNow.
  • Develop and operate platform automation, runbooks, and self-healing capabilities using Azure Automation, Logic Apps, and Python/PowerShell scripting.

CI/CD & Infrastructure Reliability

  • Collaborate with DevOps and development teams to embed reliability gates into Azure DevOps pipelines ; automated performance testing, synthetic monitoring, and progressive deployment (canary/blue-green) strategies.
  • Manage reliability of AKS clusters across multiple Azure regions, own node pool scaling, upgrade strategy and cluster hardening in alignment with CIS Benchmarks.
  • Contribute to infrastructure-as-code reliability reviews using Terraform/Bicep to enforce standards across Azure Landing Zones.

Requirements

Do you have experience in WAN?, * 7+ years of experience in SRE, platform engineering, or cloud infrastructure engineering in large-scale enterprise environments (10,000+ employees or equivalent complexity).

  • Deep, hands-on expertise with Microsoft Azure - minimum 4 years in a primary Azure cloud engineering role.
  • Expert-level proficiency with AKS: cluster lifecycle management, RBAC, network policies, pod security standards, cluster autoscaler, and Workload Identity.
  • Strong infrastructure-as-code skills: Terraform (required) and/or Bicep; experience managing Azure Landing Zones or Enterprise-Scale architecture.
  • Proficiency in at least one systems programming/scripting language: Python (preferred), Go, or PowerShell.
  • Experience designing and operating enterprise observability platforms using Azure Monitor, Log Analytics and Application Insights at scale.
  • Demonstrable track record of owning SLOs/SLIs and delivering measurable reliability improvements in production.
  • Strong knowledge of enterprise networking in Azure: Hub-and-Spoke/Virtual WAN, ExpressRoute, Azure Firewall, NSGs, Private Endpoints, and DNS Private Zones.

Required/Preferred Certifications:

  • AZ-104 AZ-305 (Preferred) AZ-400 (Preferred) CKA ITIL v4 Foundation

Benefits & conditions

Pulled from the full job description

  • Pet insurance
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Commuter assistance, * Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
  • Family Support: Maternal & Parental Leaves.
  • Insurance Options: Auto & Home Insurance, Identity Theft Protection.
  • Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
  • Time Off: Vacation, Time Off, Sick Leave & Holidays.
  • Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all