Senior Site Reliability Engineer

Cognizant Technology Solutions Corporation
Arizona City, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Arizona City, United States of America

Tech stack

API
Artificial Intelligence
Configuration Management
Code Review
Linux
Distributed Systems
Java Platform Enterprise Edition (J2EE)
Github
Monitoring of Systems
Python
Linux System Administration
Machine Learning
Networking Basics
Node.js
Performance Tuning
Release Management
Reliability Engineering
Ansible
Prometheus
Systems Integration
Web Platforms
Google Cloud Platform
System Availability
Large Language Models
Grafana
Multi-Agent Systems
Spring-boot
HybridCloud
Infrastructure as Code (IaC)
GIT
Containerization
AI Platforms
Kubernetes
Patch Management
Machine Learning Operations
Terraform
Software Version Control
Data Pipelines
Microservices

Job description

As a Senior Site Reliability Engineer you will make an impact by leading the operational excellence, reliability, and continuous improvement of AI-powered digital platforms supporting healthcare payer operations. You will oversee hybrid-cloud services leveraging Large Language Models (LLMs), MLOps practices, cloud-native technologies, and modern engineering frameworks to deliver secure, scalable, and compliant solutions that enhance member and provider experiences.

You will be a valued member of the Technology & Engineering team, collaborating closely with business stakeholders, product teams, platform engineers, AI specialists, and operations teams to ensure service stability, innovation, and regulatory compliance.

In This Role, You Will:

  • Own end-to-end service accountability for payer-focused AI, LLM, and ML-enabled platforms, ensuring high availability, performance, and compliance across hybrid environments.
  • Lead operational governance and continuous improvement initiatives aligned with enterprise service management best practices.
  • Oversee MLOps processes supporting LLM-powered applications, including model deployment, monitoring, retraining, validation, and rollback strategies.
  • Coordinate the deployment and lifecycle management of containerized services utilizing Kubernetes to deliver scalable, resilient, and highly available solutions.
  • Implement and govern Infrastructure as Code (IaC) practices using Terraform to provision and manage cloud and on-premises resources consistently and securely.
  • Standardize configuration management through Ansible automation to improve operational efficiency and reduce service disruptions.
  • Support the development, deployment, and operational management of Node.js-based services and APIs that integrate AI and machine learning capabilities.
  • Establish best practices for Git-based version control, release management, code reviews, and repository governance across application and infrastructure teams.
  • Drive operational excellence across Linux environments, including security hardening, patch management, performance optimization, and system monitoring.
  • Guide Python-based development supporting data pipelines, AI orchestration, automation frameworks, and analytics workloads.
  • Collaborate with business stakeholders, product owners, and healthcare domain experts to translate complex payer requirements into reliable technology services.
  • Lead incident management, root cause analysis, problem management, and service restoration activities to minimize business impact.
  • Monitor platform health through metrics, logs, traces, and observability tools while continuously improving service reliability and resilience.
  • Foster a culture of knowledge sharing, operational excellence, automation, and continuous improvement across distributed teams.

Work Model

We believe hybrid work is the way forward as we strive to provide flexibility wherever possible. Based on this role's business requirements, this is a hybrid position requiring attendance at a client or Cognizant office based on project needs.

The working arrangements for this role are accurate as of the date of posting and may change according to client and business requirements.

Requirements

  • Strong experience with Kubernetes and GCP (GKE)
  • Strong experience in IaC (Terraform), Helm, and GitHub Actions
  • Proficiency in Python, Ansible, Node.js
  • Strong experience with Prometheus and Grafana observability stack
  • Solid understanding of Linux systems and networking fundamentals
  • Experience in incident management, on-call support, and production triage
  • Hands-on experience with automation and CI/CD pipelines
  • Strong understanding of AI/ML concepts and AIOps practices (model lifecycle, monitoring, or AI-driven alerting)

These Will Help You Stand Out

  • Google Cloud Architect Certification
  • Certified Kubernetes Administrator (CKA)
  • Experience in Java/J2EE, Spring Boot
  • Experience supporting or operating ML/AI platforms or pipelines (MLOps)
  • Exposure to AIOps tools, anomaly detection, or predictive analytics systems
  • Experience with large-scale distributed systems and microservices architecture
  • Experience with GPU-based workloads or ML infrastructure on GCP
  • Knowledge of Kubeflow, Vertex AI, or ML pipelines
  • Experience integrating AI-driven automation into monitoring and incident response

Benefits & conditions

Cognizant offers the following benefits for this position, subject to applicable eligibility requirements:

  • Medical/Dental/Vision/Life Insurance.
  • Paid holidays plus Paid Time Off.
  • 401(k) plan and contributions.
  • Long-term/Short-term Disability.
  • Paid Parental Leave.
  • Employee Stock Purchase Plan.

About the company

At Cognizant, we're engineering modern businesses through innovation, technology, and human-centered solutions. You'll work with talented professionals, cutting-edge technologies, and industry-leading clients while helping shape the future of AI-driven healthcare solutions. We're excited to meet people who share our mission and can make an impact in a variety of ways. Don't hesitate to apply, even if you don't meet every requirement. We value diverse experiences, transferable skills, and a passion for innovation. Cognizant is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to sex, gender identity, sexual orientation, race, color, religion, national origin, disability, protected Veteran status, age, or any other characteristic protected by law.

Apply for this position