Solution Architect - Agentic AI and Observability

Tata Consultancy Services Limited
Milford, OH, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$134,400.0 - $181,900.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Cloud Computing DevOps Information Technology Operations Machine Learning Reliability Engineering Wide Area Networks Datadog Scripting
+7 more
Large Language Models Grafana SolarWinds (Software) Virtual Agents Splunk Dynatrace Microservices

Job description

TCS Cloud Unit is looking for an experienced Observability and AI Solution Architect to design and implement enterprise-grade observability and AI solutions that provide deep visibility into infrastructure, applications, networks and transform IT Operations., * Provide observability strategies for infrastructure (servers, storage, cloud), applications (microservices, APIs), and networks (LAN/WAN, SD-WAN). Collaborate with DevOps, SRE, and IT operations teams to ensure end-to-end visibility and reliability.

  • Design and architect to deliver end-to-end AIOps and observability solutions, covering telemetry collection, ingestion, correlation, analytics, dashboards, and operational workflows.
  • Design and architect AIOps solutions using industry-leading platforms like OpenAI, AWS Bedrock, Google Gemini, Anthropic, and similar technologies. Develop predictive analytics and anomaly detection models to proactively identify and resolve operational issues.
  • Guidance to onshore and offshore solution teams through requirement understanding, solution creation, estimation, and defining operating model; support RFP solutioning with architecture, roadmap, and estimates.
  • Design and recommend integrations between monitoring/observability platforms and ITSM tools using APIs and service interfaces.
  • Integrate observability tools with ITSM platforms and automation workflows. Enable automated root cause analysis and remediation using AI/ML models. Define self-healing and runbook automation
  • Collaborate with business and IT teams to identify key metrics and integrate them into dashboards and alerting systems.
  • Establish observability standards, KPIs, and SLAs for performance and availability. Ensure compliance with security and regulatory requirements in monitoring solutions.

Requirements

This role requires expertise in leading observability platforms and hands-on experience in IT operations, combined with the ability to integrate AI-driven solutions for IT Operations (AIOps) using cutting-edge technologies such as LLMs, agentic frameworks, and industry-leading platforms like Anthropic, OpenAI, Bedrock, Gemini, and others. This role requires strong customer-facing capabilities, including architecture defense, RFP solutioning, and leadership of onshore and offshore solution and delivery teams., * 10+ years of experience in IT operations, managed services, infrastructure, cloud and application support or transformation roles with significant architecture responsibility.

  • Strong Hands-on experience implementing AIOps and observability solutions across multiple observability platforms (e.g., Grafana, Datadog, Splunk, Dynatrace, ScienceLogic, Solarwinds).
  • Strong experience in integration of monitoring, event management, and automation into ITSM platforms and dashboard developments
  • Strong experience in automation and orchestration driven through AI, including solutioning with scripting techniques
  • Experience in solution design using GenAI and agentic AI use cases in IT operations ex: automated triage, knowledge generation, and runbook generation, assisted AI.
  • Communication, offshore management and stakeholder management skills, present, influence, and defend technical solutions with customers and executive audiences.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all