Technology Consultant - Site Reliability Engineer (SRE)

NTT DATA, Inc.
Atlanta, GA, United States
about 1 month ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Application Performance Management Microsoft Azure Unix Cloud Computing Linux DevOps Disaster Recovery Fault Tolerance Github Monitoring of Systems
+26 more
IBM WebSphere MQ Enterprise Messaging Systems Performance Tuning Reliability Engineering Ansible Prometheus Shell Script Software Deployment Datadog Diagnostic Tools Google Cloud Enterprise Software Applications Grafana Spring-boot Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Apache Kafka Restful APIs Terraform Splunk Dynatrace Docker Jenkins Microservices

Job description

NTT DATA’s Client is currently seeking an experienced Technology Consultant - Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability, Java, and production reliability. The ideal candidate will have experience supporting highly available and distributed enterprise applications, troubleshooting complex production issues, and driving automation and reliability improvements.

The role requires close collaboration with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, and operational efficiency.

Day to Day Job Duties

  • Manage and support business-critical applications running on Kubernetes and containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.

Requirements

  • 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years of hands-on experience with Kubernetes, Docker, and containerized application environments.
  • 4+ years of experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.

Nice to Have

  • Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
  • Experience with Kubernetes deployment and troubleshooting tools such as Helm.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Knowledge of Linux/Unix and Shell scripting.
  • Experience with Kafka, IBM MQ, or other messaging technologies.
  • Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
  • Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
  • Experience implementing distributed tracing and application performance monitoring.
  • Knowledge of incident management and ITIL processes.
  • Experience supporting high-volume, highly available, distributed enterprise applications.
  • Strong analytical, troubleshooting, communication, and problem-solving skills.

About the company

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann Chris Heilmann +2 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:04 min

Defining timestamps and the international standard format

Denny Biasiolli Denny Biasiolli · Europe 2026 Virtual

Videos

See all

Related articles

See all