DevOps & Site Reliability Engineer

Tata Consultancy Services Limited
Deerfield, IL, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Continuous Integration DevOps Key Management Log Analysis Windows PowerShell Role-Based Access Control Reliability Engineering
+17 more
Web Platforms Scripting Google Cloud Load Balancing Cloud Monitoring Multi-Cloud Event Driven Architecture Git Flow Infrastructure Automation Frameworks Performance Monitor Cloud Integration Terraform Dynatrace Api Management Docker Key Vault Microservices

Job description

  • Act as Lead SRE for client’’s Digital platforms, owning reliability and stability outcomes
  • Define and enforce SRE standards, best practices, and operating models
  • Architect and govern highly available, scalable cloud platforms
  • Lead the design and implementation of CI/CD and IaC strategies
  • Establish proactive monitoring, alerting, and incident prevention mechanisms
  • Own major incident leadership, RCA execution, and corrective action tracking
  • Partner with application, security, and architecture teams to build reliability by design
  • Drive automation to reduce toil and improve operational efficiency
  • Mentor and coach SRE and DevOps engineers across teams
  • Influence roadmap decisions with a reliability, scalability, and cost lens

Requirements

Must Have Technical/Functional Skills

  • Cloud & Platform Engineering (Expert Level)
  • Deep expertise in Microsoft Azure, including:
  • Compute (VMs, App Services, Azure Container Apps)
  • Containers & Orchestration (AKS, Docker)
  • Networking (VNETs, Private Endpoints, Application Gateway, Load Balancers)
  • Storage, Azure Key Vault, Azure Monitor, Log Analytics
  • Proven experience designing enterprise grade, highly available cloud platforms
  • Strong understanding of hybrid and multi cloud architectures (AWS / Google Cloud Platform exposure preferred)

DevOps & Engineering Excellence

  • Advanced experience with Azure DevOps and CI/CD pipeline architecture
  • Infrastructure automation using Terraform (modules, state management, governance)
  • Strong scripting skills (PowerShell, Bash)
  • GitOps concepts, branching strategies, release orchestration

  • Site Reliability Engineering (Leadership Level)
  • Ownership of platform reliability, resiliency, and performance

Definition and governance of:

  • SLIs, SLOs, SLAs
  • Error budgets and reliability metrics
  • Advanced observability strategy:
  • Metrics, logs, traces, alerts, dashboards using Dynatrace
  • Incident response leadership, RCA facilitation, and long term remediation planning
  • Experience operating 99.9%-99.99% availability systems

Containers, APIs & Integration

  • Leadership-level experience with AKS-based platforms, ingress, and scaling strategies
  • Understanding of microservices, API-led and event-driven architectures
  • Familiarity with Azure Integration Services (Service Bus, Event Hub, API Management)

Security, Compliance & Cost

  • Secure cloud design using Key Vault, managed identities, RBAC
  • Cost optimization (FinOps mindset) across cloud infrastructure

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:53 min

Introduction to git flow and clean feature branches

Johannes Haux · WWC 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all