Site Reliability Engineer (SRE) - Azure Platform Engineering

Tech Mahindra Limited
Redmond, WA, United States
5 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Application Performance Management Application Services Automation of Tests Microsoft Azure C Sharp (Programming Language) Cloud Engineering Code Generation Continuous Integration Disaster Recovery Distributed Systems
+30 more
Intrusion Detection and Prevention Log Analysis SQL Azure Network Diagrams Network Planning and Design Windows PowerShell Release Management Reliability Engineering Azure Active Directory Site Reliability Engineering Practices Cloud Services Kusto Query Language Azure Machine Learning Software Deployment Software Engineering Data Streaming Systems Architecture Web Applications Data Logging Azure Data Factory GitHub Copilot System Availability Infrastructure as Code (IaC) Git Flow Production Code Bicep Azure AKS Software Version Control Dynatrace Microservices

Job description

· Design, build, deploy, and operate highly available, secure, and scalable Azure platforms.

· Architect production-grade cloud solutions leveraging Azure Kubernetes Service (AKS), App Services, Azure SQL, Azure Batch, Azure Storage, and Azure Data Explorer (Kusto).

· Create end-to-end system architecture, deployment topology, application flow, and network diagrams.

· Own platform reliability, scalability, security, disaster recovery, and operational excellence.

· Develop Infrastructure as Code (IaC) using Bicep and automate platform provisioning.

· Implement comprehensive observability solutions covering monitoring, logging, distributed tracing, alerting, dashboards, and incident response.

· Drive platform engineering best practices including CI/CD, GitOps, automated testing, release management, and operational readiness.

· Develop production-quality tools, services, and automation using C#, PowerShell, and Bicep.

· Partner with development teams to improve platform reliability, performance, security, and developer productivity.

· Leverage AI-assisted engineering practices to accelerate solution development from rapid proof-of-concepts to enterprise-scale production deployments., We are looking for a high-energy engineer who can:

  • Think architecturally.
  • Build pragmatically.
  • Automate relentlessly.
  • Operate confidently in production.
  • Use AI effectively to accelerate outcomes.
  • Influence without authority and drive platform excellence across teams.

“Tech Mahindra is an Equal Employment Opportunity employer. We promote and support a diverse workforce at all levels of the company. All qualified applicants will receive consideration for employment without regard to race, religion, color, sex, age, national origin, or disability. All applicants will be evaluated solely on the basis of their ability, competence, and performance of the essential functions of their positions with or without reasonable accommodations. Reasonable accommodations also are available in the hiring process for applicants with disabilities. Candidates can request a reasonable accommodation by contacting the company ADA Coordinator at .”

Requirements

We are seeking a highly skilled and hands-on Senior Site Reliability Engineer (SRE) to design, build, automate, and operate mission-critical Azure-based production platforms. This role requires a strong combination of cloud architecture, software engineering, platform operations, automation, observability, and AI-driven engineering practices.

The ideal candidate is a smart, highly motivated engineer with 7-8 years of experience who has successfully taken cloud platforms from concept through production deployment and ongoing operations at enterprise scale., Strong hands-on experience designing, deploying, and operating:

  • Azure Kubernetes Service (AKS/ACA)
  • Azure App Services / Web Apps
  • Azure SQL Database
  • Azure Data Explorer (Kusto)
  • Azure Batch
  • Azure Storage Services
  • Azure Networking
  • Azure Entra ID
  • Azure Monitor & Log Analytics

Infrastructure & Architecture

· End-to-end system design

· Distributed systems architecture

· High availability and disaster recovery

· Network design and connectivity patterns

· Application and data flow modeling

· Production readiness reviews

Infrastructure as Code

· Expert-level Bicep experience

· ARM templates (preferred)

· GitOps principles

Observability

· Azure Monitor

· Application Insights

· Log Analytics

· Kusto Query Language (KQL)

· Distributed tracing

· Incident detection and response

· Reliability engineering metrics

Software Engineering

Ability to write production-ready code in:

  • C# or PowerShell
  • Bicep

Experience with:

  • APIs and microservices
  • CI/CD pipelines
  • Source control best practices
  • Automated testing

AI-Native Expectations

The successful candidate must be a native AI adopter who routinely leverages:

  • GitHub Copilot
  • Azure AI services
  • Agent-based workflows
  • AI-assisted code generation
  • AI-powered operational analysis

Must demonstrate the ability to:

· Rapidly develop POCs using AI tooling.

· Mature POCs into secure, scalable production services.

· Use AI to improve engineering productivity and operational efficiency.

Preferred Qualifications

· Experience supporting Tier-0 or Tier-1 business-critical services.

· Experience with SRE practices including SLIs, SLOs, error budgets, and incident management.

· Exposure to enterprise platform engineering organizations.

· Experience leading technical execution across multiple engineering teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

4:09 min

Selecting infrastructure tools and determining proper abstraction layers

Alayshia Knighten Alayshia Knighten · World Congress 2024

4:21 min

Scaling operations using Azure AI Foundry tools

Maxim Salnikov Maxim Salnikov · World Congress 2025

Videos

See all

Related articles

See all